Infrastructure

Managed Proxmox Day-2 Operations and Support

Change workflow, support intake, maintenance, troubleshooting, incidents, and ownership boundaries


Day-2 operations and support

Day-2 operations keep Managed Proxmox reliable after onboarding. The goal is a predictable workflow for changes, incidents, maintenance, access, troubleshooting, and ongoing ownership decisions.

Operating cadence#

CadenceActivities
ContinuousMonitor alerts, backup status, host/storage health, support tickets, and security signals
Weekly or biweeklyReview open changes, capacity trends, failed jobs, documentation gaps, and customer requests
Monthly or quarterlyAccess review, restore exercise where scoped, patch planning, capacity forecast, and risk register update
Before major changesConfirm backups, maintenance window, affected workloads, rollback plan, and customer approver
After incidentsPublish notes, root-cause or contributing-factor summary, and follow-up actions

Change requests#

Use planned changes for:

  • new VM/LXC provisioning;
  • CPU, memory, disk, or storage pool changes;
  • VLAN, firewall, public IP, DNS, or reverse-proxy updates;
  • backup scope or retention changes;
  • restore requests;
  • Proxmox, firmware, kernel, or host package updates;
  • user or service-account access changes;
  • hardware replacement, host addition, or storage expansion.

A complete change request includes business reason, environment, affected workloads, requested window, approver, validation step, rollback plan, and communication expectations.

Support intake#

For support requests, include:

FieldExample
Environment and VMproduction, staging, VM name, or cluster name
Business impactUsers affected, release blocked, revenue impact, internal-only issue
UrgencyCritical outage, degraded service, planned question, advisory request
Recent changesDeployments, firewall updates, host maintenance, storage expansion, user changes
EvidenceError messages, timestamps, screenshots, logs, monitoring links, ticket links
Desired outcomeRestore VM, add capacity, confirm health, investigate alert, schedule migration

Troubleshooting quick reference#

SymptomFirst checksLikely handoff
VM downProxmox task log, host health, storage availability, guest console, recent changesAssistance for platform; customer for guest OS/app root cause unless scoped
VM slowHost contention, storage latency, backup job, guest CPU/RAM, application loadShared: Assistance checks platform, customer checks application behavior
Disk fullVM disk usage, storage pool usage, snapshots, backup retention, logsCustomer approves data cleanup; Assistance expands platform storage when scoped
Network unreachableVLAN/bridge, guest IP/gateway, firewall, DNS, public IP, provider statusDepends on firewall/provider ownership
Backup failedTarget capacity, credentials, VM locks, snapshots, storage healthAssistance fixes platform job; customer validates data policy impact
Restore neededBackup availability, target isolation, overwrite approval, validation ownerAssistance restores; customer accepts application/data result
Suspected compromisePreserve logs, isolate if approved, review access, recent changes, affected VMsIncident workflow; customer owns application/legal communications unless contracted

Maintenance workflow#

  1. Identify maintenance need and affected hosts/VMs.
  2. Confirm backup status and rollback considerations.
  3. Request customer approval if impact is possible.
  4. Announce maintenance window and expected symptoms.
  5. Apply change, monitoring throughout.
  6. Validate platform and named workloads.
  7. Record outcome, known issues, and follow-up actions.

Ownership boundaries#

TopicAssistance ownsCustomer owns
Proxmox platformHost and cluster operations, storage/network configuration in scope, monitoring, backups, and runbooksApprovals, risk acceptance, application priority, business timing
Guest OSOnly if explicitly includedOS packages, users, services, agents, and hardening inside customer-managed VMs
ApplicationsOnly if separately contractedCode, configuration, releases, data correctness, and customer-facing behavior
Data policyImplement scoped backup/retention controlsClassification, legal retention, deletion, and restore acceptance
Provider/facilityCoordinate if scopedContracts, billing, remote hands, circuits, warranties, and provider escalations unless contracted
Security/complianceTechnical controls and evidence in scopeLegal conclusions, internal policy, access approvals, audit submissions, and breach communications

Incident workflow#

For critical issues, use the agreed incident channel and define roles: incident lead, platform operator, customer application owner, communications owner, and provider/vendor contact. Preserve evidence, stabilize first, avoid destructive changes without approval, and document the timeline as the incident unfolds.

After recovery, record customer impact, technical cause or contributing factors, what worked, what failed, and actions for monitoring, backup, architecture, documentation, or process improvement.

Offboarding or scope change#

When Managed Proxmox scope ends or changes, agree the handoff: final inventory, backups, credentials, documentation, outstanding risks, support end date, provider ownership, and data deletion/retention steps. Assistance should not retain access beyond the agreed period.