Calculate MFU (Machine FLOP Utilization) for operators like matmul/GEMM/FlashAttention on Ascend NPU, providing clear formulas and derivation process Use thi...
编程
huawei-cloud-ascendc-operator-performance-optim
试用Develop and optimize custom operators using AscendC programming language. Analyze operator performance bottlenecks and conduct optimization validation. Based...
它能做什么
Develop and optimize custom operators using AscendC programming language. Analyze operator performance bottlenecks and conduct optimization validation. Based on AscendC and CANN toolkit Use this skill when the user wants to: (1) optimize performance-critical operators on Ascend NPU, (2) develop custom operators for specific workloads, (3) improve model inference performance through operator optimization Trigger: user mentions "AscendC", "operator optimization", "custom operator", "performance", "NPU optimization", "Ascend operator", "算子优化", "自定义算子", "算子开发", "AscendC算子", "性能优化"
技能文档
Huawei Cloud AscendC Operator Performance Optimization
Overview
This skill provides guidance for developing and optimizing custom operators using AscendC programming language.
Architecture: Performance Analysis → Bottleneck Identification → Operator Development → Optimization → Validation
Related Skills:
huawei-cloud-ascend-profiler-db-explorer- Performance data analysis and bottleneck identificationhuawei-cloud-ascend-small-model-migrate- Migration workflow that may require operator optimization
Architecture Components
This skill involves the following cloud services and components:
- AscendC: Programming language for custom operator development
- CANN: Huawei Cloud AI Computing Platform for NPU
- Ascend 910B: Target NPU hardware for operator deployment
- Ascend Profiler: Performance analysis tool for validation
Architecture Diagram:
┌─────────────────────────────────────────────────────────────┐
│ AscendC Operator Optimization Skill │
├─────────────────────────────────────────────────────────────┤
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Performance │───▶│ Bottleneck │───▶│ Operator │ │
│ │ Analysis │ │ Identification│ │ Development │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Profiling │ │ Optimization│ │ Validation │ │
│ │ Data │ │ Techniques │ │ & Testing │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└─────────────────────────────────────────────────────────────┘
Use Cases
Typical Problem Scenarios:
- Optimizing performance-critical operators on Ascend NPU
- Developing custom operators for specific workloads
- Improving model inference performance through operator optimization
- Fixing operator bottlenecks identified during profiling
- Implementing missing operators for NPU deployment
Typical User Phrases:
- "Optimize my custom operator for Ascend"
- "Develop AscendC operator for GEMM"
- "Improve inference performance on NPU"
- "Fix bottleneck operator"
- "Implement custom operator using AscendC"
- "AscendCOperator"
- "OptimizationAscendOperatorPerformance"
- "OperatorPerformance"
Scope
Supported:
- Custom operator development in AscendC
- Performance optimization for existing operators
- Operator validation and testing
Not supported:
- Non-AscendC operator development
- Framework-level optimizations
Core Workflow
1. Performance Analysis
- Use profiling tools to identify performance bottlenecks
- Analyze operator execution time and resource utilization
2. Bottleneck Identification
- Identify operators with high execution time
- Determine optimization opportunities
3. Operator Development
- Implement custom operators using AscendC
- Follow AscendC best practices
4. Optimization Techniques
- Memory optimization
- Compute optimization
- Data layout optimization
5. Validation
- Verify functional correctness
- Validate performance improvement
Reference Documents
| Document | Description |
|---|---|
| Acceptance Criteria | Functional acceptance criteria |
| Verification Method | Verification approach |
| Troubleshooting | Common issues and solutions |
Prerequisites
- CANN >= 7.0.0 installed
- AscendC >= 1.0.0 installed
- Ascend NPU driver installed and working properly
- Operator code or performance data to be optimized
Core Commands
# Analyze operator performance bottlenecks
msprof --output=/path/to/output ./my_operator
# Optimize operator using AscendC
# Refer to CANN development guide for operator development
Parameter Confirmation
| Parameter | Description | Required |
|---|---|---|
| Operator code path | Operator source code to be optimized | Yes |
| Output directory | Performance analysis result output path | Yes |
| Optimization strategy | Performance optimization scheme selection | No |
Output Format
Performance analysis results are saved in the specified output directory:
output/
├── summary.json # Performance summary
├── operator_stats.csv # Operator execution statistics
├── timeline.json # Execution timeline data
└── recommendations.md # Optimization recommendations
Summary JSON Structure:
{
"total_time_ms": 1234.56,
"operator_count": 42,
"top_operators": [
{"name": "CustomGEMM", "time_ms": 456.78, "percentage": 37.0},
{"name": "VectorAdd", "time_ms": 123.45, "percentage": 10.0}
],
"optimization_candidates": ["CustomGEMM", "DataTransfer"]
}
Validation Method
Functional Validation
- Run operator with test inputs
- Compare outputs with reference implementation
- Verify numerical accuracy (tolerance: 1e-5 for FP32, 1e-3 for FP16)
Performance Validation
- Benchmark operator before optimization
- Apply optimization changes
- Benchmark operator after optimization
- Calculate speedup ratio:
speedup = time_before / time_after
Acceptance Criteria
- Functional correctness: Output matches reference within tolerance
- Performance improvement: Speedup >= 1.2x (20% improvement)
- No regression: Other operators not affected
Best Practices
Memory Optimization
- Use GM (Global Memory) for large tensors
- Use L1/L0A/L0B for intermediate results in matrix operations
- Align memory access to 32-byte boundaries
- Reuse memory buffers when possible
Compute Optimization
- Vectorize operations using AscendC intrinsics
- Use MMA (Matrix Multiply Accumulate) for matrix operations
- Parallelize independent operations
- Minimize synchronization points
Data Layout Optimization
- Use NZ format for matrix operations
- Use ND format for vector operations
- Avoid unnecessary format conversions
- Consider memory coalescing for data access
Code Structure
- Separate compute logic from memory operations
- Use template metaprogramming for flexibility
- Document optimization assumptions
- Profile before and after each optimization
Notes
Common Pitfalls
- Memory bank conflicts: Ensure data is distributed across memory banks
- Unaligned access: Check 32-byte alignment for all buffers
- Excessive synchronization: Minimize barrier usage between kernels
- Wrong data format: Match format to operation type (NZ for matmul, ND for vector)
Performance Tips
- Profile first to identify real bottlenecks
- Focus on hot paths (operators with >10% total time)
- Consider algorithmic changes before micro-optimizations
- Test with realistic input sizes
- Validate correctness after each optimization
Debugging Tips
- Use
ASCENDC_DEBUG=1for verbose logging - Check CANN log files in
/var/log/npu/ - Compare with CPU reference implementation
- Use
msproffor detailed performance breakdown
Limitations
- AscendC operators are hardware-specific (910B)
- Not all PyTorch operators have AscendC equivalents
- Custom operators require CANN recompilation for deployment
相关技能
Collect operator-level performance data on Ascend NPU using msopprof tool. Supports both device mode and simulator mode, generates performance analysis repor...
用自然语言控制华为昇腾 NPU,本地或 SSH 远程执行 npu-smi 命令。
在华为云昇腾 910B DevServer 上按单机或双机(16 卡)拓扑部署并测试 LLM、VL、Embedding、Rerank 模型。
SSH 远程连接华为云昇腾设备,支持 NPU 监控、磁盘与容器运维,凭据仅驻留内存。
Migrate vision/detection/segmentation small models to Ascend NPU, covering the full workflow: model structure analysis, migration verification, performance p...
huaweicloud-skills-team 的更多技能
浏览全部技能用自然语言控制华为昇腾 NPU,本地或 SSH 远程执行 npu-smi 命令。
在华为云昇腾 910B DevServer 上按单机或双机(16 卡)拓扑部署并测试 LLM、VL、Embedding、Rerank 模型。
面向华为云资源的只读查询能力,用于资源清点、核对与参数发现。
通过本地 Python SDK 只读查询华为云 IAM 资源(用户、用户组、策略、委托、AK/SK、MFA、安全设置)。
在华为云 Flexus L 实例上一键部署 OpenClaw AI Agent 平台,并完成模型与通道配置。
在华为云 Flexus L 实例上一键部署 Hermes AI Agent 平台,并完成大模型与机器人通道配置。