Get in Touch

Course Outline

Overview of CANN Optimization Capabilities

  • Strategies for managing inference performance within the CANN environment
  • Optimization priorities for edge-based and embedded artificial intelligence systems
  • Governance of AI Core utilization and memory allocation protocols

Using Graph Engine for Analysis

  • Fundamentals of the Graph Engine and its execution pipeline
  • Visualization of operator graphs and runtime operational metrics
  • Modifying computational graphs to enhance optimization outcomes

Profiling Tools and Performance Metrics

  • Employment of the CANN Profiling Tool for comprehensive workload analysis
  • Assessment of kernel execution duration and identification of performance bottlenecks
  • Memory access profiling and implementation of tiling strategies

Custom Operator Development with TIK

  • Summary of TIK capabilities and the operator programming framework
  • Development of custom operators using the TIK DSL
  • Testing protocols and performance benchmarking for custom operators

Advanced Operator Optimization with TVM

  • Introduction to TVM integration within the CANN architecture
  • Auto-tuning methodologies for computational graphs
  • Criteria and procedures for selecting between TVM and TIK implementations

Memory Optimization Techniques

  • Management of memory layout and buffer placement efficiency
  • Methods for reducing on-chip memory consumption
  • Best practices for asynchronous execution and resource reuse

Real-World Deployment and Case Studies

  • Case study: performance tuning for smart city camera pipeline applications
  • Case study: optimization of the autonomous vehicle inference stack
  • Guidelines for iterative profiling and continuous process improvement

Summary and Next Steps

Requirements

  • Demonstrated proficiency in deep learning model architectures and associated training processes.
  • Practical experience deploying models via CANN, TensorFlow, or PyTorch environments.
  • Competency in Linux command-line interfaces, shell scripting, and Python development.

Target Audience

  • AI performance engineers seeking standardized solutions for government applications.
  • Inference optimization specialists responsible for system efficiency.
  • Developers specializing in edge AI or real-time computing systems for government use cases.
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories