Get in Touch

Course Outline

Tencent Hunyuan Production Fundamentals

  • Overview of Tencent Hunyuan model serving scenarios relevant to government applications
  • Production characteristics of large and MoE models
  • Common latency, throughput, and cost bottlenecks
  • Defining service-level objectives for inference workloads in a federal context

Deployment Architecture and Serving Flow

  • Core components of a production inference stack suitable for government use cases
  • Choosing between containerized, on-premise, and cloud deployment models that align with agency security policies
  • Model loading, request routing, and GPU allocation basics
  • Designing for reliability and operational simplicity to support continuous service delivery

Latency Optimization in Practice

  • Using optimized inference engines such as TensorRT where applicable to enhance performance
  • KV-cache concepts and practical cache tuning
  • Reducing startup, warmup, and response overhead to meet time-sensitive operational needs
  • Measuring time to first token and token generation speed for effective workload evaluation

Throughput, Batching, and GPU Efficiency

  • Continuous batching and request batching strategies to maximize resource utilization
  • Managing concurrency and queue behavior to ensure equitable service access
  • Improving GPU utilization without harming user experience for internal stakeholders
  • Handling long-context and mixed-workload requests typical of complex government data processing

Quantization and Cost Control

  • Why quantization matters for production serving in resource-constrained environments
  • Practical trade-offs of FP16, INT8, and other common precision options for compliance and efficiency
  • Balancing model quality, latency, and infrastructure cost to ensure fiscal responsibility
  • Building a simple cost optimization checklist for sustainable technology management

Operations, Monitoring, and Readiness Review

  • Autoscaling triggers for inference services to maintain service level agreements
  • Monitoring latency, throughput, cache usage, and GPU health to ensure system integrity
  • Logging, alerting, and incident response basics aligned with federal cybersecurity standards
  • Reviewing a reference deployment and creating an improvement plan to enhance operational readiness for government missions

Requirements

  • Fundamental comprehension of deployment and inference procedures associated with large language models
  • Practical experience utilizing containerization, cloud or on-premises infrastructure, and API-driven services within a government context for government applications
  • Proficiency in Python programming or general systems engineering responsibilities

Audience

  • Mechanical engineers tasked with operationalizing large language models in production environments
  • Platform engineers managing GPU-accelerated inference services
  • Solution architects developing scalable AI serving architectures
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories