编程

vmware-vks

Use this skill whenever the user needs to manage vSphere Kubernetes Service (VKS) — Supervisor clusters, vSphere Namespaces, and TKC cluster lifecycle. Directly handles: check VKS compatibility, create/delete namespaces, create/scale/upgrade/delete TKC clusters, get kubeconfig, check Harbor registry. Always use this skill for "create Kubernetes cluster", "scale workers", "upgrade K8s version", "create namespace", "get kubeconfig", or any VKS/TKC task. Do NOT use for vanilla VM operations (use vmware-aiops), non-vSphere Kubernetes (e.g., kubeadm, EKS, AKS), or AVI/AKO load balancing (use vmware-avi). For networking use vmware-nsx.

它能做什么

Use this skill whenever the user needs to manage vSphere Kubernetes Service (VKS) — Supervisor clusters, vSphere Namespaces, and TKC cluster lifecycle. Directly handles: check VKS compatibility, create/delete namespaces, create/scale/upgrade/delete TKC clusters, get kubeconfig, check Harbor registry. Always use this skill for "create Kubernetes cluster", "scale workers", "upgrade K8s version", "create namespace", "get kubeconfig", or any VKS/TKC task. Do NOT use for vanilla VM operations (use vmware-aiops), non-vSphere Kubernetes (e.g., kubeadm, EKS, AKS), or AVI/AKO load balancing (use vmware-avi). For networking use vmware-nsx.

技能文档

VMware VKS

Disclaimer: This is a community-maintained open-source project and is not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc. "VMware" and "vSphere" are trademarks of Broadcom. Source code is publicly auditable at github.com/vmware-skills/VMware-VKS under the MIT license.

AI-powered VMware vSphere Kubernetes Service (VKS) management — 23 MCP tools.

Requires vSphere 8.x+ with Workload Management enabled. Companion skills: vmware-aiops (VM lifecycle), vmware-monitor (monitoring), vmware-storage (storage), vmware-nsx (NSX networking), vmware-nsx-security (DFW/firewall), vmware-aria (metrics/alerts/capacity), vmware-avi (AVI/ALB/AKO), vmware-harden (compliance baselines). | vmware-pilot (workflow orchestration) | vmware-policy (audit/policy)

What This Skill Does

CategoryCapabilitiesCount
SupervisorCompatibility check, status, storage policies3
NamespaceList, get, create with quotas, update, delete with TKC guard, VM classes6
TKC ClustersList, get, versions, create, scale, upgrade, delete with workload guard7
VM ServiceVM snapshots, VM groups + bootOrder, VM multi-NIC readout (vm-operator CRDs, read-only)3
AccessSupervisor kubeconfig, TKC kubeconfig, Harbor registry, storage usage4

Quick Install

uv tool install vmware-vks
vmware-vks check

When to Use This Skill

  • Check if vSphere environment supports VKS
  • Create, update, or delete Supervisor Namespaces with resource quotas
  • Deploy, scale, upgrade, or delete TKC (TanzuKubernetesCluster) clusters
  • Get kubeconfig for Supervisor or TKC clusters
  • Check Harbor registry info or storage usage

Use companion skills for:

  • VM lifecycle, deployment → vmware-aiops
  • Inventory, health, alarms → vmware-monitor
  • iSCSI, vSAN, datastore → vmware-storage
  • Load balancing, AVI/ALB, AKO, Ingress → vmware-avi
User IntentRecommended Skill
Read-only monitoringvmware-monitor
Storage: iSCSI, vSANvmware-storage
VM lifecycle, deploymentvmware-aiops
vSphere Kubernetes Service (vSphere 8.x+)vmware-vks ← this skill
NSX networking: segments, gateways, NATvmware-nsx
NSX security: DFW rules, security groupsvmware-nsx-security
Aria Ops: metrics, alerts, capacity planningvmware-aria
Multi-step workflows with approvalvmware-pilot
Compliance baselines (CIS / 等保 / PCI-DSS), drift detection, LLM remediation advisorvmware-harden (uv tool install vmware-harden)
Load balancer, AVI, ALB, AKO, Ingressvmware-avi (uv tool install vmware-avi)
Audit log queryvmware-policy (vmware-audit CLI)

Common Workflows

Deploy a New TKC Cluster

Pre-flight (judgment):

  • Supervisor must be vSphere 8.x+ with WCP enabled — supervisor check returns pass/fail. If fail, no amount of TKC commands will work; resolve at vSphere/WCP layer first.
  • K8s version: pick a TKR version that's still supported by VMware (not EOL). New clusters on EOL versions look fine until you need a CVE patch and there isn't one.
  • VM class sizing: best-effort-* for dev, guaranteed-* for prod. A best-effort worker can be evicted under host pressure — production workloads need guaranteed.
  • Storage policy: must already exist in vCenter. list_supervisor_storage_policies first and pass the returned policy ID (not the display name); creating a TKC against a missing policy fails after CP boot, leaving partial state.
  • Control-plane count: 1 for dev, 3 for prod (HA). Cannot upgrade from 1→3 without recreating; choose right the first time.
  • Namespace quota: TKC consumes CP + worker × (cpu, memory) from namespace quota. If quota is too tight, workers fail to schedule with no obvious error.
  • TKC API version: auto-detected at runtime via the K8s discovery API (prefers cluster.x-k8s.io/v1 when the Supervisor serves it, falls back to v1beta1 on vSphere 8.0). No manual selection needed; advanced callers can override via the api_version parameter on generate_tkc_yaml().

Steps:

  1. vmware-vks supervisor check --target prod → must pass
  2. vmware-vks tkc versions -n → pick a non-EOL TKR
  3. (If new namespace) vmware-vks namespace create dev --storage-policy --cpu --apply --dry-run then real
  4. vmware-vks tkc create dev-cluster -n dev --version --control-plane 1 --workers 3 --vm-class best-effort-large --apply --dry-run then real
  5. Wait for phase=running (typically 10-15 min); do not assume success on apply return
  6. vmware-vks kubeconfig get dev-cluster -n dev -o ./kubeconfig — write to file, do not paste tokens into the agent context

Scale Workers for Load Testing

Judgment: scaling is fast but reverse-scaling is destructive — workers are deleted, in-flight pods lost. Treat scale-down like a delete.

  1. tkc get dev-cluster -n dev → record current worker count and any pending pods
  2. Scale-up: tkc scale dev-cluster -n dev --workers 6 → safe, additive operation
  3. Verify new workers reach Ready in kubectl get nodes before sending traffic
  4. Scale-down: drain pods first via kubectl drain on the to-be-deleted nodes, THEN tkc scale --workers 3. Skipping drain causes pod restarts on remaining nodes — measurable user impact.
  5. Confirm namespace quota leftover supports the new size — quota is enforced at scheduling, not at scale request

Namespace Resource Management

Judgment: quota changes are atomic but consequences are not. Reducing quota below current usage doesn't evict pods — they keep running, but no new pods schedule, looking like a "namespace is broken" symptom.

  1. namespace list → see all namespaces and their phase
  2. storage -n dev → check current CPU/memory/storage usage; never reduce quota below current usage + 20% headroom
  3. namespace update dev --cpu --memory --dry-run → preview, then real
  4. Validate by attempting a small pod scale-up; if it pends with Insufficient cpu, quota is still the bottleneck

Architecture

User (Natural Language)
  ↓
AI Agent (Claude Code / Goose / Cursor)
  ↓ reads SKILL.md
  ↓
vmware-vks CLI  ─── or ───  vmware-vks MCP Server (stdio)
  │
  ├─ Layer 1: pyVmomi → vCenter REST API
  │   Supervisor status, storage policies, Namespace CRUD, VM classes, Harbor
  │
  └─ Layer 2: kubernetes client → Supervisor K8s API endpoint
      TKC CR apply / get / delete  (cluster.x-k8s.io/v1beta1)
      Kubeconfig bearer token from POST /wcp/login (Supervisor JWT)
  ↓
vCenter Server 8.x+ (Workload Management enabled)
  ↓
Supervisor Cluster → vSphere Namespaces → TanzuKubernetesCluster

Usage Mode

ScenarioRecommendedWhy
Local/small models (Ollama, Qwen)CLI~2K tokens vs ~8K for MCP
Cloud models (Claude, GPT-4o)EitherMCP gives structured JSON I/O
Automated pipelinesMCPType-safe parameters, structured output

MCP Tools (23 — 16 read, 7 write)

All accept optional target parameter to specify a named vCenter.

list_namespaces, list_supervisor_storage_policies and list_vm_classes return the family list envelope — {items, returned, limit, total, truncated, hint} — rather than a bare array. Read the rows from items; truncated says whether the listing is complete, so it never has to be guessed from the row count. These three read their collection in one un-paged call, so total is the real count and truncated is always false.

CategoryToolType
Supervisorcheck_vks_compatibilityRead
get_supervisor_statusRead
list_supervisor_storage_policiesRead
Namespacelist_namespacesRead
get_namespaceRead
create_namespaceWrite
update_namespaceWrite
delete_namespaceWrite
list_vm_classesRead
TKClist_tkc_clustersRead
get_tkc_clusterRead
get_tkc_available_versionsRead
create_tkc_clusterWrite
scale_tkc_clusterWrite
upgrade_tkc_clusterWrite
delete_tkc_clusterWrite
VM Servicelist_vm_snapshotsRead
list_vm_groupsRead
list_vm_network_interfacesRead
Accessget_supervisor_kubeconfigRead
get_tkc_kubeconfigRead
get_harbor_infoRead
list_namespace_storage_usageRead

create_namespace / create_tkc_cluster — defaults to dry_run=True, returns a YAML plan for review. Pass dry_run=False to apply.

delete_namespace — requires confirmed=True and rejects if TKC clusters still exist (prevents orphaned clusters).

delete_tkc_cluster — requires confirmed=True and checks for running workloads. Rejects if found unless force=True.

Credential handling: get_supervisor_kubeconfig and get_tkc_kubeconfig return short-lived session tokens (not long-lived credentials). Tokens are derived from the authenticated vCenter session and expire when the session ends. Kubeconfig output is intended for local kubectl use — agents should write it to a file (-o ) rather than displaying tokens in conversation context.

Full capability details and safety features: see references/capabilities.md

CLI Quick Reference

# Supervisor
vmware-vks check [--target ]
vmware-vks preflight-auth [--target ]   # live-validate POST /wcp/login (issue #13)
vmware-vks supervisor status  [--target ]
vmware-vks supervisor storage-policies [--target ]

# Namespace
vmware-vks namespace list [--target ]
vmware-vks namespace get  [--target ]
vmware-vks namespace create  --cluster  [--cpu ] [--memory ] [--storage-policy ] [--apply]
vmware-vks namespace update  [--cpu ] [--memory ] [--target ]
vmware-vks namespace delete  [--target ]

# TKC Clusters
vmware-vks tkc list [-n ] [--target ]
vmware-vks tkc create  -n  [--version ] [--workers ] [--vm-class ] [--apply]
vmware-vks tkc scale  -n  --workers  [--pool ] [--target ]
vmware-vks tkc upgrade  -n  --version  [--target ]
vmware-vks tkc delete  -n  [--skip-workload-check] [--target ]

# Kubeconfig
vmware-vks kubeconfig supervisor -n  [--target ]
vmware-vks kubeconfig get  -n  [-o ] [--target ]

# Harbor & Storage
vmware-vks harbor [--target ]
vmware-vks storage -n  [--target ]

Full CLI reference with all flags and interactive creation: see references/cli-reference.md

Troubleshooting

"VKS not compatible" error

Workload Management must be enabled in vCenter. Check: vCenter UI → Workload Management. Requires vSphere 8.x+ with Enterprise Plus or VCF license.

Namespace creation fails with "storage policy not found"

List policies first: vmware-vks supervisor storage-policies, then pass the Policy ID column value (not the display name) as --storage-policy.

TKC cluster stuck in "Creating" phase

Check Supervisor events in vCenter. Common causes: insufficient resources on ESXi hosts, network issues with NSX-T, or storage policy not available on target datastore.

Validating Supervisor auth (POST /wcp/login)

Supervisor/TKC Kubernetes auth uses a JWT obtained from POST https:///wcp/login (HTTP Basic → JSON session_id bearer token), not the pyVmomi SOAP session key. To validate this end-to-end against your real Supervisor, run:

vmware-vks preflight-auth [--target ]

It performs the real login (no mocks) and reports, per target: vCenter reachable → /wcp/login HTTP status → parseable session_id → does the JWT authenticate a trivial Supervisor K8s API call. A healthy result is all four steps ✓ PASS ending in target '': /wcp/login auth flow validated end-to-end. (exit code 0). On failure each step prints a teaching message — e.g. a 404 on /wcp/login means the endpoint path differs on your Supervisor version (capture the real path), a 401 on the K8s probe means session_id is not the bearer token on your version. It never tracebacks — every failure is status output.

Kubeconfig retrieval fails

Supervisor API endpoint must be reachable from the machine running vmware-vks. Check firewall rules for port 6443.

Scale operation has no effect

Verify the cluster is in "Running" phase before scaling. Clusters in "Creating" or "Updating" phase reject scale operations.

Delete namespace rejected unexpectedly

The namespace delete guard prevents deletion when TKC clusters exist inside. Delete all TKC clusters in the namespace first, then retry.

Prerequisites

  • vSphere 8.x+ with Workload Management enabled
  • Enterprise Plus or VCF license
  • NSX-T (recommended) or VDS + HAProxy networking
  • Supervisor Cluster configured and running

Setup

uv tool install vmware-vks
mkdir -p ~/.vmware-vks
vmware-vks init

All tools are automatically audited via vmware-policy. Audit logs: vmware-audit log --last 20

Full setup guide, security details, and AI platform compatibility: see references/setup-guide.md

Audit & Safety

All operations are automatically audited via vmware-policy (@vmware_tool decorator):

  • Every tool call logged to ~/.vmware/audit.db (SQLite, framework-agnostic) with a local JSON-Lines mirror at ~/.vmware-vks/audit.log
  • Policy rules enforced via ~/.vmware/rules.yaml (deny rules, maintenance windows, risk levels)
  • Risk classification: each tool tagged as low/medium/high/critical
  • View recent operations: vmware-audit log --last 20
  • View denied operations: vmware-audit log --status denied

In-memory kubeconfig (v1.5.18+): kubeconfig for the Supervisor and TKC clusters — which embeds the vCenter session bearer token — is built as a Python dict and loaded into the kubernetes client via load_kube_config_from_dict(). The token never touches disk during normal MCP/CLI flow, eliminating the previous temp-file TOCTOU window. The explicit kubeconfig get -o CLI export still writes to the user-chosen path for kubectl use.

vmware-policy is automatically installed as a dependency — no manual setup needed.

License

MIT — github.com/vmware-skills/VMware-VKS

相关技能

管理 vSphere 存储——数据存储、iSCSI 和 vSAN——通过 12 个 MCP 工具或 CLI。

53 次安装1 星标

Use this skill whenever the user needs to manage VMware NSX networking — segments, gateways, NAT, routing, and IP pools. Directly handles: create/manage network segments, configure Tier-0/Tier-1 gateways, set up NAT rules, manage static routes, configure IP pools, check transport node and edge cluster health. Always use this skill for "create segment", "set up gateway", "create NAT rule", "check network health", "troubleshoot connectivity", or any NSX/networking/segment task. Do NOT use for DFW firewall rules or security groups (use vmware-nsx-security), VM lifecycle (use vmware-aiops), or AVI/ALB load balancing (use vmware-avi). For multi-step workflows use vmware-pilot.

49 次安装

Use this skill whenever the user needs VMware Aria Operations (rebranded VMware VCF Operations in VCF 9 and later) data — performance metrics, alerts, capacity planning, anomaly detection, and automated reports. Directly handles: query resource metrics, list/acknowledge/cancel alerts, manage alert definitions, check capacity and time-remaining forecasts, detect anomalies, generate and manage reports. Always use this skill for "check vSphere capacity", "what Aria Operations alerts are active", "show VMware anomalies", "generate an Aria report", "rightsizing recommendations", "VCF Operations alerts", or any Aria Operations / VCF Operations / vRealize Operations task. Combined with LLM, Aria data powers natural language reports: "give me a capacity report" → Aria collects data → LLM formats the report. Do NOT use for real-time vCenter alarms/events (use vmware-monitor), VM operations (use vmware-aiops), or NSX networking (use vmware-nsx). For load balancing/AVI/AKO use vmware-avi.

52 次安装

Use this skill whenever the user mentions load balancing, ingress, virtual services, pool members, AVI, NSX ALB, AKO, or application delivery in a VMware/NSX ALB or Tanzu/vSphere Kubernetes context. Directly handles: virtual service listing and enable/disable, pool member drain/enable, SSL certificate expiry checks, analytics and error logs, service engine health, AKO pod troubleshooting, AKO Helm config management, Ingress annotation validation, K8s-to-Controller sync diagnostics, and multi-cluster AKO overview. Always use it for "virtual service", "pool member", "AKO status", "AKO logs", "ingress diagnose", "ssl expiry", "load balancer", "NSX ALB", "AVI controller", "AKO sync", or "负载均衡" tasks. Do NOT use to set up or configure nginx/HAProxy/Traefik from scratch — those are not AVI tasks. For VM lifecycle use vmware-aiops, for NSX networking use vmware-nsx, for Kubernetes cluster lifecycle (Supervisor/TKC) use vmware-vks.

50 次安装

Use this skill whenever the user needs to manage VMware NSX security — distributed firewall (DFW) policies, security groups, microsegmentation, and IDS/IPS. Directly handles: create/manage DFW policies and rules, security groups, VM tags, network traceflow diagnostics, IDPS profiles and status. Always use this skill for "create firewall rule", "set up microsegmentation", "add VM to security group", "run traceflow", "check IDS status", or any NSX security/DFW task. Do NOT use for NSX networking operations like segments, gateways, NAT, or routing (use vmware-nsx), or VM lifecycle (use vmware-aiops). For load balancing/AVI/AKO use vmware-avi.

53 次安装

改任何文件前先看 diff、备份、等用户确认,安全升级已安装的代理技能。

92 次安装6 星标