Trois pannes trouvées en LANÇANT la descente, aucune vue en la relisant — ni
par moi, ni par l'attaque adversariale.
Le premier « apt update » d'une image cloud échoue sur un verrou qui n'est pas
celui qu'on croit. Mesuré une seconde après le premier ssh :
E: Could not get lock /var/lib/apt/lists/lock.
It is held by process 1026 (apt-get)
Ce n'est pas cloud-init — « status --wait » avait rendu la main. C'est
apt-daily, le minuteur de Debian, qui se déclenche au démarrage. Et le verrou
des LISTES n'est pas couvert par « DPkg::Lock::Timeout », que l'installeur
réglait pourtant déjà à 600 s : cette attente ne vaut que pour dpkg. On arrête
donc les minuteurs, puis on RÉESSAIE — arrêter une unité n'interrompt pas
l'apt-get déjà en vol. Le défaut touchait tout déploiement Proxmox, pas
seulement ce test.
Le premier étage n'avait pas d'alias ssh. « deploy_qemu.py » en ligne de
commande n'écrit pas d'entrée ~/.ssh/config — le menu le fait, la CLI non. La
descente aurait attendu son plein délai avant de conclure « jamais joignable »
sur une VM qui répondait à son adresse. Elle l'écrit maintenant elle-même,
depuis l'adresse résolue, et refuse d'avancer si la VM n'en a pas.
Et la réserve de l'hôte est proportionnelle. Quatre gigaoctets sur une machine
de soixante, c'était 6 % laissés au système : le jour où les invités touchent
vraiment leur mémoire, c'est l'hôte qui part en swap — et la mesure serait
celle du swap, pas de l'imbrication. Un huitième, avec le plancher d'avant
pour les petites machines.
Ce que la descente a établi en trois étages : 392 s, 644 s, 1120 s, soit 1,7
fois par étage. Puis la poignée de main ssh passe de 77 à 1664 secondes au
quatrième — vingt fois d'un seul cran. Le coude est là.
Et le mur que j'avais pris pour une limite d'imbrication n'en était pas une.
La VM qui gelait au quatrième étage avait douze vCPU ; celle-ci en a deux et
elle passe, en écrivant. C'était une limite de parallélisme SOUS imbrication —
exactement ce que l'algorithme borne, vérifié pour la première fois plutôt que
supposé.
--- EN ---
Three faults found by RUNNING the descent, none seen by reading it — neither
by me nor by the adversarial attack.
A cloud image's first "apt update" fails on a lock that is not the one you
expect. Measured one second after the first ssh:
E: Could not get lock /var/lib/apt/lists/lock.
It is held by process 1026 (apt-get)
It is not cloud-init — "status --wait" had returned. It is apt-daily, Debian's
timer, firing at boot. And the LISTS lock is not covered by
"DPkg::Lock::Timeout", which the installer already set to 600 s: that wait
only applies to dpkg. So we stop the timers, then RETRY — stopping a unit does
not interrupt the apt-get already in flight. The defect affected every Proxmox
deployment, not just this test.
The first level had no ssh alias. "deploy_qemu.py" on the command line does not
write a ~/.ssh/config entry — the menu does, the CLI does not. The descent
would have waited its full timeout before concluding "never reachable" about a
VM answering at its address. It now writes the entry itself, from the resolved
address, and refuses to proceed if the VM has none.
And the host's reserve is proportional. Four gigabytes on a sixty-gigabyte
machine left 6 % to the system: the day the guests really touch their memory,
the host swaps — and the measurement would be of swap, not of nesting. One
eighth now, keeping the old floor for small machines.
What the descent established over three levels: 392 s, 644 s, 1120 s — 1.7x per
level. Then the ssh handshake goes from 77 to 1664 seconds at the fourth:
twenty times in one step. That is the elbow.
And the wall I had taken for a nesting limit was not one. The VM that froze at
the fourth level had twelve vCPU; this one has two and it gets through,
writing. It was a limit of parallelism UNDER nesting — exactly what the
algorithm caps, verified for the first time rather than assumed.
Assisted-by: Claude Opus 5
(cherry picked from commit 7f86562cbd8f10017dcb88fe4272efc162cbccbc)
|
||
|---|---|---|
| .. | ||
| install_proxmox.sh | ||
| nesting.py | ||
| proxmox_deploy.py | ||
| README.base.md | ||
| README.fr.md | ||
| README.md | ||
Deploying VMs on a Proxmox VE host
Two different things live in this directory:
install_proxmox.shturns a Debian into a Proxmox hypervisor. Seescript/qemu/README.md, which documents theproxmoxdistro of the deployment catalog.proxmox_deploy.pydeploys VMs on such a host, fromTODO › Execute › Deploy › Proxmox VE, right underQEMU/KVM.
The whole difference: the hypervisor is elsewhere
With QEMU/KVM, the hypervisor is the machine running the script. With Proxmox it is somewhere else, so the first question is which host — and the answer is remembered for the session. Three ways, all offered by the menu:
- From the local QEMU VMs — a
proxmoxVM deployed here. Its address comes from the DHCP lease, nothing to retype. - By address —
user@host, plus an optional SSH jump. - From
~/.ssh/config— the alias already carries user, port and ProxyJump; nothing else is asked.
The chosen host is then checked, not assumed: pveversion proves it is a
Proxmox, id -u and sudo -n true decide whether commands need sudo, and an
unknown SSH host key is offered for recording (with ssh-keyscan, never by
disabling the check — a hypervisor is not a throwaway VM).
# Ce que l'outil envoie, et qu'on peut rejouer à la main :
ssh erplibre@pve1 sudo sh -c 'qm list'
ssh erplibre@pve1 sudo sh -c 'pvesm status --content images'
Why SSH and qm, not the REST API
The API needs a token or a ticket to create and renew. qm is the path every
Proxmox administrator knows, the repository already manages SSH access
(~/.ssh/config, ProxyJump, keys), and the commands stay readable in the log —
so they can be replayed by hand. That is how every failure of this module was
diagnosed.
sudo sh -c '<whole command>' and not sudo <command>: these commands are
sequences and redirections. Prefixing with sudo would elevate only the first
word, and the redirection would still be the unprivileged shell's.
Four traps met on a real host
A Proxmox installed on Debian has no vmbr0 — the ISO installer creates
one, that procedure does not. And qm create requires a bridge.
The menu offers an internal bridge (vmbr0, 10.10.10.1/24, NAT through
the uplink). Never adding the physical NIC to a bridge is deliberate: that
moves the host address and cuts the SSH session in progress — remotely, there
is no way back. A LAN-facing bridge is printed as a stanza to apply from a
console.
On an internal bridge no DHCP answers, so the address is static, derived from the VMID — and therefore known before the VM boots. Looking for it afterwards was absurd.
The Debian cloud image does not ship qemu-guest-agent, so Proxmox cannot tell
the address of a DHCP guest: it does not hand out the leases. The fallback is
the host's own neighbour table (ip neigh), which needs nothing from the guest.
The menu, entry by entry
The seventeen QEMU/KVM entries have their counterpart. Four of them are the
same code, because it is the same work: reopening the install monitoring,
the remote desktop tunnel, the Android emulator and the image catalog. They
reach Proxmox guests through the ~/.ssh/config entries that entry 13 writes,
with the Proxmox host as ProxyJump.
[1] Déployer une VM [8] Redimensionner un disque [15] Émulateur Android *
[2] Prévisualiser [9] Effacer des VM [16] Catalogue d'images *
[3] Télécharger une image [10] Nettoyer (orphelins) [17] Exemple (dry-run)
[4] Rouvrir le suivi * [11] Tester une VM (Odoo) [18] Changer d'hôte
[5] Lister (qm list) [12] Statistiques
[6] Adresse IP d'une VM [13] Configuration SSH
[7] Console d'une VM [14] Tunnel bureau distant *
* code partagé avec le menu QEMU/KVM
Verified
Deploying a VM inside a Proxmox that itself runs in a libvirt VM: image
downloaded on the host, internal bridge created, static address, cloud-init
user and key, disk resized, qm start. Then ssh vm-essai from the outside
reaches it through the jump — three nested levels. Resize 12G → 16G, delete
with --purge, orphan scan: all checked against Proxmox VE 9.2.11.