Use this skill whenever the user needs to operate a Nutanix estate through Prism Central (v4 REST API) — an estate/cluster health overview, cluster & host inventory and utilization, VM lifecycle across AHV and ESXi (list/get/power/create/update/clone/delete/migrate), storage containers, subnets/network, images & categories, data protection / DR (snapshots, recovery points, protection domains, VM protect, failover), alerts & events with alert RCA (analyze_alert), LCM firmware/software upgrades, capacity runway forecasting, and read-only diagnostics/RCA over the whole estate (cluster_health_rca, alert_triage_rca). Always use this skill for "Nutanix", "Prism Central", "AHV", "cluster health", "list VMs", "power on/off a VM", "clone/migrate a VM", "delete a VM", "snapshot", "recovery point", "protection domain", "failover", "why did this alert fire" / "root cause this alert", "LCM upgrade / firmware", "days until storage is full" / "capacity runway", "what's wrong with my cluster" / "diagn
Memory
truenas-aiops
Try itUse this skill whenever the user needs to operate TrueNAS SCALE storage — a one-shot health overview, system info, read-only diagnostics / RCA (pool health, alerts & dataset capacity), inspect ZFS pools (list/get/status, capacity, scrub status, start a scrub), datasets (list/get/create), snapshots (list/create/delete), physical disks and S.M.A.R.T. self-test results, system alerts, services (list/restart), and replication / cloud-sync tasks. Always use this skill for "list truenas pools", "truenas dataset", "create zfs snapshot", "start a scrub", "diagnose truenas pool health", "why is my pool degraded", "truenas disk health", "truenas smart test", "truenas alerts", "restart truenas service", or "truenas replication" when the context is explicitly TrueNAS / TrueNAS SCALE / a ZFS NAS appliance. Do NOT use when the target is not a TrueNAS SCALE appliance — other NAS/storage products, backup software, hypervisor VM lifecycle, container clusters, and network devices are out of scope (negat
What it does
Use this skill whenever the user needs to operate TrueNAS SCALE storage — a one-shot health overview, system info, read-only diagnostics / RCA (pool health, alerts & dataset capacity), inspect ZFS pools (list/get/status, capacity, scrub status, start a scrub), datasets (list/get/create), snapshots (list/create/delete), physical disks and S.M.A.R.T. self-test results, system alerts, services (list/restart), and replication / cloud-sync tasks. Always use this skill for "list truenas pools", "truenas dataset", "create zfs snapshot", "start a scrub", "diagnose truenas pool health", "why is my pool degraded", "truenas disk health", "truenas smart test", "truenas alerts", "restart truenas service", or "truenas replication" when the context is explicitly TrueNAS / TrueNAS SCALE / a ZFS NAS appliance. Do NOT use when the target is not a TrueNAS SCALE appliance — other NAS/storage products, backup software, hypervisor VM lifecycle, container clusters, and network devices are out of scope (negative routing hints only). Common TrueNAS SCALE operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). Live-verified against real TrueNAS SCALE 25.04 and 26 appliances over both transports; see docs/VERIFICATION.md for what is and is not covered.
The skill document
TrueNAS AIops
Disclaimer: This is a community-maintained open-source project and is not affiliated with, endorsed by, or sponsored by iXsystems or the TrueNAS project. "TrueNAS" is a trademark of its owner. Source code is publicly auditable at github.com/AIops-tools/TrueNAS-AIops under the MIT license.
Governed TrueNAS SCALE storage operations — 25 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.truenas-aiops/, policy engine, token/runaway budget guard, undo-token recording, and descriptive risk tiers. The TrueNAS API key is stored encrypted (~/.truenas-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.
Standalone: the governance harness is bundled in the package (
truenas_aiops.governance) — truenas-aiops has no external skill-family dependency. Verification: coverage focuses on common TrueNAS operations and is not exhaustive, but it is no longer mock-only — reads, governed writes with audit + undo, the WebSocket transport, degraded-pool RCA, replication and cloud-sync have all been exercised against live TrueNAS SCALE 25.04 and 26 appliances.docs/VERIFICATION.mdrecords exactly what was checked and what is still open.
What This Skill Does
| Category | Tools | Count | Read or Write |
|---|---|---|---|
| Overview / System | health overview, system info | 2 | 2 read |
| Diagnostics / RCA | pool health RCA, alert & capacity RCA | 2 | 2 read |
| Pools | list, get, status, scrub status, capacity | 5 | 5 read |
| scrub start | 1 | 1 write (medium) | |
| Datasets | list, get | 2 | 2 read |
| create | 1 | 1 write (medium) | |
| Snapshots | list | 1 | 1 read |
| create (medium), delete (high) | 2 | 2 write | |
| Disks | list, S.M.A.R.T. results | 2 | 2 read |
| Alerts | list | 1 | 1 read |
| Services | list | 1 | 1 read |
| restart | 1 | 1 write (medium) | |
| Replication | replication tasks, cloud-sync tasks | 2 | 2 read |
Quick Install
uv tool install truenas-aiops
truenas-aiops init # interactive wizard: connection + encrypted API key
truenas-aiops doctor
When to Use This Skill
- Triage a TrueNAS appliance (
overview): pool capacity/health, alerts, running services - Root-cause a degraded/full pool (
diagnose pool-health) or a wall of alerts (diagnose alerts) — worst-first findings that cite the measured number - List/inspect ZFS pools, datasets, and snapshots
- Create a snapshot before a risky change; start a pool scrub
- Check disk health and S.M.A.R.T. self-test results
- List and restart system services (smb/nfs/ssh)
- Inspect replication and cloud-sync tasks
Do NOT use when the target is not a TrueNAS SCALE appliance — other NAS/storage or backup products, hypervisor VM lifecycle, Kubernetes/containers, and network devices are out of scope for this skill.
Related Skills — Skill Routing
| If the user wants… | Use |
|---|---|
| TrueNAS pools / datasets / snapshots / ZFS health | truenas-aiops (this skill) |
| Backup software job/restore operations | a backup-software ops skill |
| Hypervisor VM lifecycle (power, snapshot, migrate) | a hypervisor ops skill |
| Container/cluster lifecycle | a cluster ops skill |
Common Workflows
Root-cause a degraded or full pool (start here)
truenas-aiops diagnose pool-health→ worst-first findings: bad ZFS state (DEGRADED/FAULTED/OFFLINE), non-zero read/write/checksum/scan error counters, and pools over 80%/90% capacity — each citing the measured numbertruenas-aiops pool status→ inspect the topology / scan detail the finding citedtruenas-aiops pool scrub-start→ kick an integrity scrub (governed, medium risk); poll withpool scrub-statustruenas-aiops diagnose alerts→ cross-check active alerts by level and any datasets nearing their quota/available ceiling
Snapshot a dataset before a change, then roll back if needed
truenas-aiops dataset list→ confirm the dataset id (e.g.tank/data)truenas-aiops snapshot create tank/data pre-change→ records an inversesnapshot_deleteundo descriptor- Make your change; if it went wrong, the snapshot is your recovery point
truenas-aiops snapshot delete tank/data@pre-change --dry-run→ preview; then without--dry-run(double confirm) — IRREVERSIBLE, captures BEFORE state, no undo
Scrub a pool and follow it
truenas-aiops pool list→ find the pool name and healthtruenas-aiops pool scrub-start tank→ starts the integrity scrubtruenas-aiops pool scrub-status→ checkstate/percentage; do not re-issue (the runaway budget guard backs a tight poll loop)
Usage Mode
| Scenario | Recommended | Why |
|---|---|---|
| Local/small models | CLI | fewer tokens than MCP |
| Cloud models (Claude, GPT) | Either | MCP gives structured JSON I/O |
| Automated pipelines | MCP | type-safe parameters, audited |
MCP Tools (25 — 19 read, 6 write)
| Category | Tools | R/W |
|---|---|---|
| Overview / System | overview, system_info | Read |
| Diagnostics / RCA | pool_health_rca, alert_and_capacity_rca | Read |
| Pools | pool_list, pool_get, pool_status, scrub_status, pool_capacity | Read |
pool_scrub_start | Write | |
| Datasets | dataset_list, dataset_get | Read |
dataset_create | Write | |
| Snapshots | snapshot_list | Read |
snapshot_create, snapshot_delete | Write | |
| Disks | disk_list, smart_test_results | Read |
| Alerts | alert_list | Read |
| Services | service_list | Read |
service_restart | Write | |
| Replication | replication_list, cloudsync_list | Read |
| Undo (governance) | undo_list | Read |
undo_apply | Write |
Harness features that light up: snapshot_create passes an undo= lambda so the harness records an inverse snapshot_delete descriptor (with _undo_id) to the undo store. snapshot_delete is tagged risk_level=high, captures the snapshot's BEFORE state, and declares no undo (it is irreversible). pool_scrub_start, dataset_create, and service_restart are medium risk and capture prior state where relevant. All 25 tools are audit-logged under ~/.truenas-aiops/ and pass through the budget/runaway guard, with a descriptive risk-tier label on each audit row. Start any triage with overview.
CLI Quick Reference
truenas-aiops init # onboarding wizard (encrypted API key)
truenas-aiops overview [--target ] # health summary
truenas-aiops system [--target ] # version / hostname / memory / uptime
truenas-aiops diagnose pool-health # RCA: pool state / error counters / capacity (worst first)
truenas-aiops diagnose alerts # RCA: active alerts by level + datasets near full
truenas-aiops pool list
truenas-aiops pool get
truenas-aiops pool status
truenas-aiops pool scrub-status
truenas-aiops pool capacity # size / allocated / free / used%
truenas-aiops pool scrub-start
truenas-aiops dataset list
truenas-aiops dataset get # e.g. tank/data
truenas-aiops dataset create [--dry-run]
truenas-aiops snapshot list [--dataset tank/data] [--limit 200]
truenas-aiops snapshot create
truenas-aiops snapshot delete [--dry-run] # double confirm, IRREVERSIBLE
truenas-aiops disk list
truenas-aiops disk smart # S.M.A.R.T. self-test results
truenas-aiops alert list
truenas-aiops service list
truenas-aiops service restart [--dry-run] # double confirm (smb/nfs/ssh)
truenas-aiops replication list
truenas-aiops replication cloudsync
truenas-aiops secret set # store API key encrypted
truenas-aiops secret list # names only
truenas-aiops secret migrate # import legacy plaintext .env
truenas-aiops secret rotate-password
truenas-aiops doctor
truenas-aiops mcp # start MCP server (stdio)
See references/cli-reference.md for the full command list, and
references/agent-guardrails.md when driving these tools with a smaller /
local model (enforced guardrails, ready-to-paste system prompt).
Troubleshooting
"Config file not found"
Run truenas-aiops init to set up your first target (writes ~/.truenas-aiops/config.yaml and stores the API key encrypted).
"No API key for target ''"
Add it to the encrypted store: truenas-aiops secret set (prompts hidden), or run truenas-aiops init. Create the key in the TrueNAS UI under Credentials → API Keys. For non-interactive use (MCP/CI), also export TRUENAS_AIOPS_MASTER_PASSWORD so the store can be unlocked without a prompt.
"Master password not set" / "Wrong master password"
The encrypted store ~/.truenas-aiops/secrets.enc is unlocked by TRUENAS_AIOPS_MASTER_PASSWORD (or an interactive prompt). If you forgot it, delete secrets.enc and re-run truenas-aiops init. Rotate it with truenas-aiops secret rotate-password.
"Authentication/authorization failed (401/403)"
The API key is wrong or revoked, or the account lacks permission. Regenerate the key in the TrueNAS UI (Credentials → API Keys) and update it: truenas-aiops secret set .
"Could not reach TrueNAS … check the host/port"
Confirm the TrueNAS web/REST endpoint is reachable on the configured port (default 443) and api_path is /api/v2.0. For self-signed certificates set verify_ssl: false on the target (lab only).
"Resource not found (404)"
The pool/dataset/snapshot id is stale. List the parent collection first (pool list, dataset list, snapshot list) to get a current id.
Audit & Safety
The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (scope the TrueNAS API key to a limited-privilege account and writes then fail at the appliance). There is no read-only switch, policy file, or approval gate.
- API key stored encrypted in
~/.truenas-aiops/secrets.enc(Fernet/AES-128 + scrypt key derivation; chmod 600) — never plaintext on disk; the master password is never stored, only a per-store salt + ciphertext. - Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to
~/.truenas-aiops/audit.db(relocatable viaTRUENAS_AIOPS_HOME): params (secrets redacted), result, status, duration, and the risk tier. The CLI writes the same row the MCP path does. TRUENAS_AUDIT_APPROVED_BY/TRUENAS_AUDIT_RATIONALEare optional annotations recorded on the audit row (who/why); they are never required and never block.- Runaway guard — a safety backstop, not authorization: cumulative tool calls and wall-time are capped, and a tight scrub/poll loop trips a circuit breaker.
- Writes support
--dry-run/dry_run=Trueand double confirmation at the CLI; CLI writes execute through the same governed tools, so they are audited + undo-recorded. - Reversible writes capture the real fetched before-state and record an inverse descriptor (e.g.
snapshot_create→snapshot_delete) that replays against the tool's own signature.
The harness is bundled in the package — no external dependency, no manual setup. See references/setup-guide.md for security details.
Contributing & feature requests
Coverage is intentionally focused, and what has actually been verified against live appliances is recorded in docs/VERIFICATION.md. Missing a capability you need, or hit an endpoint that needs fixing for your TrueNAS version? Open an issue or pull request at github.com/AIops-tools/TrueNAS-AIops — feature requests, contributions, and comments are all welcome.
License
Related skills
Use this skill whenever the user needs to manage VMs and containers on Proxmox VE — list/inspect/configure VMs, power and lifecycle (start/stop/shutdown/reboot/reconfigure/clone/delete/migrate), snapshots (create/delete/list/rollback), disk grow/move, vzdump backups (create/list/restore), LXC containers (list/start/stop), cluster/node status, cluster resource inventory, async task polling + logs, free-VMID lookup, HA status, resource pools, firewall inspection, guest-agent ping, and storage listing. Also use it to diagnose cluster health — rank nodes by CPU/memory/disk pressure and scan guests for saturation (read-only RCA). Always use this skill for "list proxmox vms", "start proxmox vm", "stop proxmox vm", "proxmox snapshot", "proxmox backup", "restore proxmox vm", "resize proxmox disk", "proxmox vm status", "migrate proxmox vm", "proxmox container", "proxmox ha", "proxmox pool", "proxmox firewall", "list proxmox storage", "proxmox node pressure", or "why is proxmox slow" when the co
Use this skill whenever the user needs to operate or diagnose MinIO object storage — explain why the cluster is filling up or refusing writes (capacity_rca), find publicly exposed buckets and hygiene gaps (bucket_exposure_audit), find storage that lifecycle/ILM should be reclaiming but isn't, including noncurrent versions and incomplete multipart uploads (lifecycle_gap_analysis), check heal backlog and erasure-set write-quorum risk (healing_health), read service health / cluster status / per-bucket config (policy, versioning, lifecycle, encryption, quota, tags) — plus governed writes (set or delete bucket policy, enable/suspend versioning, set or delete lifecycle rules, set bucket quota, purge incomplete uploads, delete an empty bucket). Always use this skill for "minio health", "why is my object storage full", "which bucket is biggest", "is any bucket public / anonymous access", "versioning / noncurrent versions piling up", "incomplete multipart uploads", "lifecycle / ILM rules", "buc
Use this skill whenever the user needs to operate or diagnose a Ceph cluster via its ceph-mgr Dashboard REST API — decode a HEALTH_WARN/ERR state into cause + action (cluster_health), read the cluster status, inspect OSDs (tree/df/perf), placement groups (summary/stuck/scrub), pools (list/usable capacity), RBD images and snapshots, CephFS/MDS and RGW status, monitors/managers, slow ops and capacity forecast — plus governed writes (set cluster flags, reweight/mark-in/mark-out/purge OSDs, trigger scrubs, set pool quota/pg_num/autoscale/size, create/delete pools, create/delete RBD images and snapshots, throttle recovery/backfill). Always use this skill for "ceph health", "what does this HEALTH_WARN mean", "PG_DEGRADED / OSD_NEARFULL / SLOW_OPS / MON_DOWN", "ceph -s", "which OSD is most full", "drain an OSD", "purge an OSD", "stuck PGs", "overdue scrub", "pool usable capacity", "set pool size / quota", "rebalance is too slow / throttle backfill", "RBD image or snapshot", "MDS behind on tri
Use this skill whenever the user needs to operate a network device — read device facts, interfaces (+ counters/IP), BGP/LLDP neighbors (summary and detail), ARP/MAC tables, VLANs, routes, hardware environment (fans/temp/power/CPU/mem), optics, NTP, users, SNMP info, VRFs, and an aggregated device-health summary; run read-only RCA diagnostics on interface health and BGP neighbors; back up a switch/router config, diff a candidate config (dry-run), and merge/replace/rollback config — across Cisco IOS/IOS-XE, Nexus NX-OS, IOS-XR, Arista EOS, and Juniper Junos via NAPALM. An optional NetBox block adds source-of-truth lookups. Always use this skill for "back up switch config", "show bgp neighbors", "diff network config", "push config to router", "show interfaces on the switch", or tasks mentioning "cisco", "arista", "juniper", "nexus", "ios-xr", or "napalm". Do NOT use when the target is not a NAPALM-supported network device (Kubernetes clusters, hypervisor VMs, and cloud consoles are out of
Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-health, workload-readiness); scale deployments/statefulsets; rollout status/history/undo/pause/resume and set image; delete pods/deployments/jobs; create/delete namespaces; and cordon/uncordon/drain nodes. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS). Always use this skill for "list k8s pods", "scale deployment", "kubernetes pod logs", "describe pod", "why is my pod crashing", "diagnose pods", "which deployments are unhealthy", "rollout undo", "set image", "top pods", "drain node", "cordon node", "restart deployment", "k3s", or "kubectl"-style tasks when the context is exp