Attaqué par trois lentilles sur le code écrit, avant de le lancer pour de vrai. Deux fautes valaient à elles seules l'exercice. Il n'aurait JAMAIS fonctionné. L'installeur était lancé par « sh », or il porte « set -euo pipefail » et un shebang bash : sur Debian /bin/sh est dash, qui répond « set: Illegal option -o pipefail » et sort à la PREMIÈRE ligne. Chaque étage aurait échoué sur l'installation, à tous les coups. Et « --detruire » pouvait emporter une machine étrangère. Il prenait toute entrée ssh dont le nom CONTENAIT « deep-pve », puis sur son rebond détruisait toute VM dont le nom contenait « deep-pve » — une « deep-pve-lab » de production tombait dedans, et « --purge » emporte les disques. Son tri « du plus profond au plus haut » comptait les « + » de l'alias, or alias_etage remplace le « + » du parent par un « - » : chaque alias en portait exactement UN, le tri ne triait rien, et la destruction partait du plus HAUT — le disque du parent emportait ses enfants sans qu'on les ait nommés. Il ignorait « --dry-run », ne lisait aucun code de retour, concluait « ✓ défait », et le menu le lançait d'une touche. Il ne détruit plus que ce que le RAPPORT nomme : un couple (parent, VMID) par étage, du plus profond d'après le niveau lu, égalité stricte du nom, arrêt CONSTATÉ avant destruction, codes de retour lus, et une confirmation par « OUI » après la liste. Six autres constats, tous réels. Le redémarrage se prouve par btime et non par le seul noyau — rejoué sur un étage déjà installé, on validait un redémarrage qui n'avait pas eu lieu, exactement le piège corrigé la semaine dernière dans le suivi. La sonde de disponibilité ne demande plus sudo, sinon un sudo lent se lisait « jamais joignable en ssh ». Les délais suivent la profondeur : le script existe pour mesurer un ralentissement de 36x, et un plafond fixe déclarait échouée une installation qui avançait. L'adresse fixe est contrôlée AVANT de télécharger une image et de démarrer une VM. Le DNS de l'hôte suit la spec, sinon apt meurt sans rien expliquer. Et l'essai à blanc ne prétend plus avoir atteint quoi que ce soit — son rapport était indiscernable d'une réussite, JSON compris. L'algorithme aussi : profondeur 0 rendait un plan d'UN étage, donc « --depth 0 » créait une VM ; et sur un hôte de quatre cœurs le premier étage recevait UN vCPU quand son invité en recevait deux — un parent plus étroit que son enfant. Les tests mordent, prouvé par mutation : remplacer le calcul du premier étage par la valeur imbriquée les laissait verts. --- EN --- Attacked by three lenses on the written code, before running it for real. Two faults alone justified the exercise. It would NEVER have worked. The installer was run by "sh", yet it carries "set -euo pipefail" and a bash shebang: on Debian /bin/sh is dash, which answers "set: Illegal option -o pipefail" and exits on the FIRST line. Every level would have failed at install, every time. And "--detruire" could take a stranger's machine. It took every ssh entry whose name CONTAINED "deep-pve", then on its jump host destroyed every VM whose name contained "deep-pve" — a production "deep-pve-lab" fell in, and "--purge" takes the disks. Its "deepest first" sort counted the "+" in the alias, yet alias_etage replaces the parent's "+" with a "-": every alias had exactly ONE, the sort sorted nothing, and destruction started from the TOP — the parent's disk took its children with it, unnamed. It ignored "--dry-run", read no return code, concluded "✓ done", and the menu fired it on one key. It now destroys only what the REPORT names: a (parent, VMID) pair per level, deepest first by the recorded level, strict name equality, shutdown VERIFIED before destruction, return codes read, and a "OUI" confirmation after the list. Six more findings, all real. The reboot is proven by btime, not by the kernel alone — replayed on an already-installed level, we validated a reboot that never happened, exactly the trap fixed last week in the monitor. The liveness probe no longer asks for sudo, or a slow sudo read as "never reachable by ssh". Timeouts follow the depth: the script exists to measure a 36x slowdown, and a fixed ceiling declared failed an install that was progressing. The static address is checked BEFORE downloading an image and starting a VM. The host's DNS follows the spec, or apt dies explaining nothing. And the dry run no longer claims to have reached anything — its report was indistinguishable from a success, JSON included. The algorithm too: depth 0 returned a ONE-level plan, so "--depth 0" created a VM; and on a four-core host the first level got ONE vCPU while its guest got two — a parent narrower than its child. The tests bite, proven by mutation: replacing the first level's computation with the nested value left them green. Assisted-by: Claude Opus 5 (cherry picked from commit 64b8e5063bd7f420cdeb27b88f94043190b5ecd4) |
||
|---|---|---|
| .. | ||
| install_proxmox.sh | ||
| nesting.py | ||
| proxmox_deploy.py | ||
| README.base.md | ||
| README.fr.md | ||
| README.md | ||
Deploying VMs on a Proxmox VE host
Two different things live in this directory:
install_proxmox.shturns a Debian into a Proxmox hypervisor. Seescript/qemu/README.md, which documents theproxmoxdistro of the deployment catalog.proxmox_deploy.pydeploys VMs on such a host, fromTODO › Execute › Deploy › Proxmox VE, right underQEMU/KVM.
The whole difference: the hypervisor is elsewhere
With QEMU/KVM, the hypervisor is the machine running the script. With Proxmox it is somewhere else, so the first question is which host — and the answer is remembered for the session. Three ways, all offered by the menu:
- From the local QEMU VMs — a
proxmoxVM deployed here. Its address comes from the DHCP lease, nothing to retype. - By address —
user@host, plus an optional SSH jump. - From
~/.ssh/config— the alias already carries user, port and ProxyJump; nothing else is asked.
The chosen host is then checked, not assumed: pveversion proves it is a
Proxmox, id -u and sudo -n true decide whether commands need sudo, and an
unknown SSH host key is offered for recording (with ssh-keyscan, never by
disabling the check — a hypervisor is not a throwaway VM).
# Ce que l'outil envoie, et qu'on peut rejouer à la main :
ssh erplibre@pve1 sudo sh -c 'qm list'
ssh erplibre@pve1 sudo sh -c 'pvesm status --content images'
Why SSH and qm, not the REST API
The API needs a token or a ticket to create and renew. qm is the path every
Proxmox administrator knows, the repository already manages SSH access
(~/.ssh/config, ProxyJump, keys), and the commands stay readable in the log —
so they can be replayed by hand. That is how every failure of this module was
diagnosed.
sudo sh -c '<whole command>' and not sudo <command>: these commands are
sequences and redirections. Prefixing with sudo would elevate only the first
word, and the redirection would still be the unprivileged shell's.
Four traps met on a real host
A Proxmox installed on Debian has no vmbr0 — the ISO installer creates
one, that procedure does not. And qm create requires a bridge.
The menu offers an internal bridge (vmbr0, 10.10.10.1/24, NAT through
the uplink). Never adding the physical NIC to a bridge is deliberate: that
moves the host address and cuts the SSH session in progress — remotely, there
is no way back. A LAN-facing bridge is printed as a stanza to apply from a
console.
On an internal bridge no DHCP answers, so the address is static, derived from the VMID — and therefore known before the VM boots. Looking for it afterwards was absurd.
The Debian cloud image does not ship qemu-guest-agent, so Proxmox cannot tell
the address of a DHCP guest: it does not hand out the leases. The fallback is
the host's own neighbour table (ip neigh), which needs nothing from the guest.
The menu, entry by entry
The seventeen QEMU/KVM entries have their counterpart. Four of them are the
same code, because it is the same work: reopening the install monitoring,
the remote desktop tunnel, the Android emulator and the image catalog. They
reach Proxmox guests through the ~/.ssh/config entries that entry 13 writes,
with the Proxmox host as ProxyJump.
[1] Déployer une VM [8] Redimensionner un disque [15] Émulateur Android *
[2] Prévisualiser [9] Effacer des VM [16] Catalogue d'images *
[3] Télécharger une image [10] Nettoyer (orphelins) [17] Exemple (dry-run)
[4] Rouvrir le suivi * [11] Tester une VM (Odoo) [18] Changer d'hôte
[5] Lister (qm list) [12] Statistiques
[6] Adresse IP d'une VM [13] Configuration SSH
[7] Console d'une VM [14] Tunnel bureau distant *
* code partagé avec le menu QEMU/KVM
Verified
Deploying a VM inside a Proxmox that itself runs in a libvirt VM: image
downloaded on the host, internal bridge created, static address, cloud-init
user and key, disk resized, qm start. Then ssh vm-essai from the outside
reaches it through the jump — three nested levels. Resize 12G → 16G, delete
with --purge, orphan scan: all checked against Proxmox VE 9.2.11.