编程

Huawei Cloud Cce Cluster Management

试用

通过 Python SDK 管理华为云 CCE 集群、节点池、节点和插件的全生命周期,危险操作需 confirm=true。

它能做什么

基于华为云 Python SDK,提供 CCE 集群的创建、休眠、唤醒与删除,以及节点池扩缩容、节点调度控制(cordon/uncordon/drain)、插件的安装更新与卸载,并支持绑定/解绑集群 EIP 以及拉取 kubeconfig。删除集群、休眠、节点驱逐等高风险操作必须两步确认:不带 confirm=true 时仅返回预览和风险提示,显式传入 confirm=true 才会真正执行。AK/SK 通过 HW_ACCESS_KEY 和 HW_SECRET_KEY 环境变量读取,仅在调用期间使用,不会落盘或写入日志。

什么时候用它

  • 创建 Turbo 集群并初始化首个节点池
  • 在维护窗口对节点池进行扩容或缩容
  • 下线节点前执行 cordon、drain 后删除
  • 安装或更新 coredns、metrics-server、everest 等核心插件

技能文档

Huawei Cloud CCE Cluster Management

Overview

Manage CCE (Cloud Container Engine) cluster lifecycle, including cluster creation/deletion/hibernation/awakening, node pool management, node scheduling control, and addon management.

⛔ Security Constraints

Dangerous Operation Confirmation Mechanism

This skill strictly enforces a two-step confirmation mechanism for all dangerous operations to prevent accidental service disruption or data loss.

All dangerous operations require confirm=true parameter to execute. Otherwise, they return a preview and confirmation prompt.

Operations Requiring Confirmation

ToolOperation TypeRisk LevelDescription
huawei_delete_cce_clusterDelete🔴 CriticalDeletes entire CCE cluster, irreversible
huawei_hibernate_cce_clusterHibernate🟠 HighStops all workloads, pauses control plane billing
huawei_awake_cce_clusterAwake🟠 HighResumes cluster from hibernation
huawei_resize_cce_nodepoolScale🟡 MediumAdjusts node pool size, affects capacity
huawei_delete_cce_nodepoolDelete🟠 HighDeletes node pool, affects business capacity
huawei_delete_cce_nodeDelete🟠 HighRemoves node from cluster, affects scheduling
huawei_uninstall_cce_addonUninstall🟠 HighRemoves addon, may affect cluster functionality
huawei_cce_node_cordonCordon🟡 MediumMarks node unschedulable, new pods won't be assigned
huawei_cce_node_uncordonUncordon🟡 MediumMarks node schedulable, new pods may be assigned immediately
huawei_cce_node_drainDrain🟠 HighEvicts all pods from node, affects running workloads

Workflow

Step 1: Preview Operation - Call without confirm parameter

# Example: Preview cluster deletion
python3 scripts/huawei-cloud.py huawei_delete_cce_cluster \
  region=cn-north-4 \
  cluster_id=xxx

Returns: operation preview, risk warning, confirmation example

Step 2: Confirm Execution - Call with confirm=true

# Example: Confirm and execute deletion
python3 scripts/huawei-cloud.py huawei_delete_cce_cluster \
  region=cn-north-4 \
  cluster_id=xxx \
  confirm=true

Credential Security

This skill strictly follows these security rules:

  1. No persistent credential storage - Never saves AK/SK, tokens, or certificates to disk
  2. No long-term memory cache - AK/SK exists only during API call, released afterward
  3. Only project ID memory cache - Non-sensitive project ID cached in process memory
  4. No credential leakage - Never includes AK/SK in logs, responses, or errors
  5. Temporary file cleanup - If temporary cert files are created, they are deleted immediately after use

AK/SK usage methods:

  • Environment variables HW_ACCESS_KEY / HW_SECRET_KEY / HW_REGION_NAME (process-level, not saved)
  • Per-call parameter (valid only for that call)

Prerequisites

Python Environment

  • Python 3.8+
  • Install SDKs: pip install huaweicloudsdkcce huaweicloudsdkcore
  • Optional for node operations: pip install kubernetes
export HW_ACCESS_KEY="your-access-key-id"
export HW_SECRET_KEY="your-secret-access-key"
export HW_REGION_NAME="cn-north-4"

IAM Permission Policies

Ensure the IAM user has the minimum required permissions:

PermissionDescription
cce:cluster:listList clusters
cce:cluster:getGet cluster details
cce:cluster:createCreate clusters
cce:cluster:deleteDelete clusters
cce:cluster:updateUpdate clusters (hibernate/awake/bind EIP)
cce:node:listList nodes
cce:node:getGet node details
cce:node:createCreate nodes
cce:node:deleteDelete nodes
cce:node:updateUpdate nodes (cordon/uncordon/drain)
cce:nodepool:listList node pools
cce:nodepool:createCreate node pools
cce:nodepool:deleteDelete node pools
cce:nodepool:updateUpdate node pools (resize)
cce:addon:listList addons
cce:addon:getGet addon details
cce:addon:createInstall addons
cce:addon:updateUpdate addons
cce:addon:deleteUninstall addons

Core Commands

Cluster Query

ToolFunctionParameters
huawei_list_cce_clustersList all CCE clusters in regionregion
huawei_get_cce_nodesGet detailed node informationregion, cluster_id, node_id
huawei_get_cce_kubeconfigGet cluster kubeconfigregion, cluster_id, duration

Cluster Management

ToolFunctionRisk LevelRequires Confirmation
huawei_create_cce_clusterCreate CCE cluster🟢 LowNo
huawei_delete_cce_clusterDelete CCE cluster🔴 CriticalYes
huawei_hibernate_cce_clusterHibernate cluster🟠 HighYes
huawei_awake_cce_clusterAwake cluster🟠 HighYes
huawei_bind_cce_cluster_eipBind cluster EIP🟢 LowNo
huawei_unbind_cce_cluster_eipUnbind cluster EIP🟡 MediumNo

Recommended defaults:

  • Cluster type: Turbo (best performance with ENI network)
  • Container network: eni for Turbo clusters
  • Naming format: --cluster (e.g., prod-web-cluster)

Node Pool Management

ToolFunctionRisk LevelRequires Confirmation
huawei_list_cce_nodepoolsList node pools🟢 LowNo
huawei_create_cce_nodepoolCreate node pool🟢 LowNo
huawei_delete_cce_nodepoolDelete node pool🟠 HighYes
huawei_resize_cce_nodepoolResize node pool🟡 MediumYes

Recommended defaults:

  • Naming format: --pool (e.g., prod-worker-pool)
  • Initial node count: 2 for HA, or 0 with autoscaling
  • Enable autoscaling for dynamic scaling

Node Management

ToolFunctionRisk LevelRequires Confirmation
huawei_list_cce_nodesList cluster nodes🟢 LowNo
huawei_create_cce_nodeCreate nodes directly🟢 LowNo
huawei_delete_cce_nodeDelete node🟠 HighYes
huawei_cce_node_cordonMark node unschedulable🟡 MediumYes
huawei_cce_node_uncordonMark node schedulable🟡 MediumYes
huawei_cce_node_drainEvict all pods from node🟠 HighYes
huawei_cce_node_statusQuery node scheduling status🟢 LowNo

Note: Prefer node pools for managed scaling. Direct node creation is for special cases.

Addon Management

ToolFunctionRisk LevelRequires Confirmation
huawei_list_cce_addonsList cluster addons🟢 LowNo
huawei_get_cce_addon_detailGet addon details🟢 LowNo
huawei_install_cce_addonInstall addon🟢 LowNo
huawei_uninstall_cce_addonUninstall addon🟠 HighYes
huawei_update_cce_addonUpdate addon🟡 MediumNo

Common addons:

  • coredns - DNS service
  • metrics-server - Monitoring metrics
  • everest - Storage driver

Network Prerequisites

ToolFunctionParameters
huawei_list_vpcList VPCs with CIDR inforegion
huawei_list_vpc_subnetsList subnets with AZ inforegion, vpc_id

Use these tools to find VPC/subnet IDs before cluster creation.


Supported Regions

Region CodeRegion Name
cn-north-4North China-Beijing 4
cn-north-1North China-Beijing 1
cn-north-2North China-Beijing 2
cn-east-3East China-Shanghai 1
cn-south-1South China-Guangzhou
cn-south-2South China-Guangzhou Friendly
cn-east-4East China II
cn-southwest-2Guiyang 1
ap-southeast-1Asia-Pacific-Hong Kong
ap-southeast-2Asia-Pacific-Bangkok
ap-southeast-3Asia-Pacific-Singapore

Output Format

All tools return JSON-formatted results containing:

  • status: operation result (success / error)
  • data: operation-specific response (cluster info, node list, addon details, etc.)
  • message: human-readable description of the result
  • warning: risk warning for dangerous operations (preview mode only)

Verification

See verification-method.md for detailed verification steps. Quick checklist:

  1. Verify AK/SK credentials are configured via environment variables
  2. Run huawei_list_cce_clusters to confirm API connectivity
  3. Test dangerous operation preview (call without confirm=true)
  4. Verify Turbo cluster ENI network configuration

Best Practices

  • Use environment variables (HW_ACCESS_KEY / HW_SECRET_KEY) for credentials — avoid hardcoding
  • Always preview dangerous operations before confirming with confirm=true
  • Use Turbo clusters (container_network_type=eni) for high-performance workloads
  • Resize node pools during low-traffic periods to minimize business impact
  • Keep node pools at ≥2 nodes for production workloads to ensure redundancy
  • Regularly check cluster health via huawei_list_cce_clusters and huawei_show_cce_cluster

References

DocumentDescription
task-cluster-management.mdCluster lifecycle operations
task-nodepool-management.mdNode pool operations
task-node-management.mdNode scheduling operations
iam-policies.mdIAM permission policies
verification-method.mdVerification steps
troubleshooting.mdTroubleshooting guide
cce-api-guide.mdCCE Python SDK API reference
cce-cluster-parameters.mdCluster/nodepool creation parameters

Notes

  • Ensure AK/SK has correct IAM permissions
  • Different regions may have different resource availability
  • All dangerous operations require confirmation
  • Deletion operations are irreversible
  • Hibernate cluster stops all workloads - use during non-business hours
  • Node drain evicts all pods - ensure sufficient replicas
  • Turbo clusters recommended for best performance with ENI network

常见问题

危险操作如何防止误执行?
删除集群、休眠集群、节点驱逐等高风险操作在未传 confirm=true 时只返回操作预览和风险说明,只有显式传入 confirm=true 后才会真正执行,便于提前核对影响范围。
凭据存放在哪里,会不会泄露?
AK/SK 通过环境变量 HW_ACCESS_KEY 与 HW_SECRET_KEY 传入,仅在当次 API 调用期间使用,既不落盘也不缓存,不会出现在日志或返回结果中。
使用前需要准备什么?
需要 Python 3.8+ 并安装 huaweicloudsdkcce 和 huaweicloudsdkcore(节点操作可选 kubernetes),同时 IAM 用户需具备 cce:cluster、cce:node、cce:nodepool、cce:addon 的 list、get、create、update、delete 等权限,详见技能文档。

相关技能

Huawei Cloud CCE/UCS workload lifecycle management skill using hcloud CLI for kubeconfig acquisition and kubectl for Kubernetes resource operations. Use this...

3 次安装1 星标

Query Huawei Cloud CCE (Cloud Container Engine) clusters and report their names, IDs, statuses, versions, and node information across a project. Use when listing CCE clusters, looking up a cluster name, showing cluster detail, or inspecting cluster status and nodes. Provides read-only inspection for daily operations, inventory reporting, and troubleshooting. Triggers include: CCE query, list CCE clusters, query CCE cluster names, CCE cluster inventory, show CCE cluster, list CCE nodes, check cluster status, CCE集群查询, 查询CCE集群, CCE集群名称, CCE集群列表, 查看CCE集群.

通过 hcloud CLI 全生命周期管理华为云 CCI 容器实例:命名空间、网络、工作负载、日志查询,并内置安全确认机制。

作者 shijingcheng3 次安装1 星标

Huawei Cloud CCE Node failure diagnosis skill using Python SDK dispatcher. Use this skill when the user wants to: (1) diagnose CCE node NotReady, node resour...

4 次安装1 星标

Huawei Cloud CCE auto-remediation runner skill that converts remediation intent into preview-first, confirm-required, post-verify execution plans. Use this s...

5 次安装1 星标

查询华为云 CCE 集群 Pod/Node 指标及 ECS、ELB、EIP、NAT 资源指标,支持基于阈值的异常检测。

作者 shijingcheng4 次安装1 星标

shijingcheng 的更多技能

浏览全部技能

查询华为云 CCE 集群 Pod/Node 指标及 ECS、ELB、EIP、NAT 资源指标,支持基于阈值的异常检测。

作者 shijingcheng4 次安装1 星标

在华为云 SWR 上配置跨区域镜像同步和触发器,让镜像推送自动变成 CCE/CCI 部署更新。

作者 shijingcheng3 次安装1 星标

通过 hcloud CLI 管理华为云 SWR 命名空间、镜像仓库、版本标签、登录凭证与配额。

作者 shijingcheng3 次安装1 星标

通过 hcloud CLI 管理华为云 SWR 镜像权限、保留规则、共享下载域名与委托关系。

作者 shijingcheng3 次安装1 星标

通过 hcloud CLI 全生命周期管理华为云 CCI 容器实例:命名空间、网络、工作负载、日志查询,并内置安全确认机制。

作者 shijingcheng3 次安装1 星标