Le conseil « rejouer install_proxmox.sh sur l'hôte » ne pouvait PAS marcher : la VM clone le dépôt distant, donc sa copie du script est celle du distant — tant que le correctif n'y est pas, celle qui ne corrige rien. Trois hôtes de suite sont tombés dessus, avec le même message inutile. L'écran répare donc : gel de cloud-init, réécriture de /etc/hosts, relance des unités, constat du montage — le pendant exact de l'offre de créer un pont. Écrit, puis ATTAQUÉ par trois lentilles sur le code réel. Ce qu'elles ont mesuré valait la peine. /etc/hosts se réécrivait en DEUX écritures — « sed -i » puis « printf >> » — alors que la docstring promettait l'inverse. Sed refusé et ajout réussi, la ligne 127.0.1.1 survivait EN PREMIER et la nôtre s'ajoutait une fois par tentative ; sed réussi et ajout refusé, l'hôte perdait l'entrée de son nom, et chaque sudo y attendait ensuite le résolveur. C'est maintenant un fichier complet bâti dans un temporaire, VÉRIFIÉ, puis recopié — « cat > » et non « mv », qui remplacerait l'inode et perdrait mode et propriétaire. Le contrôle final s'en remettait à « getent hosts », qui réussit via mDNS même quand rien n'a été écrit — et acceptait les fe80:: que notre propre code rejette. Il relit désormais ce qui a été écrit. awk remplace sed pour filtrer : « print » émet un saut de ligne, donc un /etc/hosts non terminé par un — cloud-init n'en met pas — est normalisé. Sans ça notre ligne se collait à la précédente et le nom du nœud partait sur l'adresse d'une autre machine. Trois autres, du même acabit. Les dépendants de pmxcfs sont relancés eux aussi : actifs pendant la panne, ils échouaient sur ipcc_send_rec, et les laisser donnait une GUI en « communication failure » juste après notre ✓. Un silence du lien n'est plus lu comme une absence de montage. Et l'adresse n'est mise en cause que si pve-cluster a réellement démarré. Les tests exécutent les commandes au lieu de les relire, bouchons capables d'ÉCHOUER : écriture refusée, fichier sans saut de ligne final, tabulations, start qui rate, montage qui disparaît pendant la reconfirmation. Prouvé par mutation — trois HOSTS-KO changés en HOSTS-OK font rougir le test. --- EN --- The advice "replay install_proxmox.sh on the host" could NOT work: the VM clones the remote, so its copy of the script is the remote's — while the fix is not there, the one that fixes nothing. Three hosts in a row hit it with the same useless message. So the screen repairs: freeze cloud-init, rewrite /etc/hosts, restart the units, verify the mount — the exact counterpart of the offer to create a bridge. Written, then ATTACKED by three lenses on the real code. What they measured was worth it. /etc/hosts was rewritten in TWO writes — "sed -i" then "printf >>" — while the docstring promised the opposite. Sed refused and append succeeded: the 127.0.1.1 line survived FIRST and ours was added once per attempt; sed succeeded and append refused: the host lost its own name entry, and every sudo then waited on the resolver. It is now a complete file built in a temporary, VERIFIED, then copied over — "cat >" not "mv", which would replace the inode and lose mode and owner. The final check relied on "getent hosts", which succeeds via mDNS even when nothing was written — and accepted the fe80:: our own code rejects. It now re-reads what was written. awk replaces sed for filtering: "print" emits a newline, so an /etc/hosts with no final one — cloud-init omits it — gets normalised. Without that our line glued onto the previous one and the node's name pointed at another machine's address. Three more of the same kind. pmxcfs's dependents are restarted too: active throughout the outage, they failed on ipcc_send_rec, and leaving them gave a GUI in "communication failure" right after our ✓. A silent link is no longer read as a missing mount. And the address is only blamed if pve-cluster actually started. The tests execute the commands instead of reading them, with stubs able to FAIL: refused write, file with no final newline, tabs, a start that fails, a mount that vanishes during reconfirmation. Proven by mutation — three HOSTS-KO turned into HOSTS-OK make the test go red. Assisted-by: Claude Opus 5 (cherry picked from commit d4f9358c6cb562029cc2ca9eb478d80c6a0a19a4) |
||
|---|---|---|
| .. | ||
| install_proxmox.sh | ||
| proxmox_deploy.py | ||
| README.base.md | ||
| README.fr.md | ||
| README.md | ||
Deploying VMs on a Proxmox VE host
Two different things live in this directory:
install_proxmox.shturns a Debian into a Proxmox hypervisor. Seescript/qemu/README.md, which documents theproxmoxdistro of the deployment catalog.proxmox_deploy.pydeploys VMs on such a host, fromTODO › Execute › Deploy › Proxmox VE, right underQEMU/KVM.
The whole difference: the hypervisor is elsewhere
With QEMU/KVM, the hypervisor is the machine running the script. With Proxmox it is somewhere else, so the first question is which host — and the answer is remembered for the session. Three ways, all offered by the menu:
- From the local QEMU VMs — a
proxmoxVM deployed here. Its address comes from the DHCP lease, nothing to retype. - By address —
user@host, plus an optional SSH jump. - From
~/.ssh/config— the alias already carries user, port and ProxyJump; nothing else is asked.
The chosen host is then checked, not assumed: pveversion proves it is a
Proxmox, id -u and sudo -n true decide whether commands need sudo, and an
unknown SSH host key is offered for recording (with ssh-keyscan, never by
disabling the check — a hypervisor is not a throwaway VM).
# Ce que l'outil envoie, et qu'on peut rejouer à la main :
ssh erplibre@pve1 sudo sh -c 'qm list'
ssh erplibre@pve1 sudo sh -c 'pvesm status --content images'
Why SSH and qm, not the REST API
The API needs a token or a ticket to create and renew. qm is the path every
Proxmox administrator knows, the repository already manages SSH access
(~/.ssh/config, ProxyJump, keys), and the commands stay readable in the log —
so they can be replayed by hand. That is how every failure of this module was
diagnosed.
sudo sh -c '<whole command>' and not sudo <command>: these commands are
sequences and redirections. Prefixing with sudo would elevate only the first
word, and the redirection would still be the unprivileged shell's.
Four traps met on a real host
A Proxmox installed on Debian has no vmbr0 — the ISO installer creates
one, that procedure does not. And qm create requires a bridge.
The menu offers an internal bridge (vmbr0, 10.10.10.1/24, NAT through
the uplink). Never adding the physical NIC to a bridge is deliberate: that
moves the host address and cuts the SSH session in progress — remotely, there
is no way back. A LAN-facing bridge is printed as a stanza to apply from a
console.
On an internal bridge no DHCP answers, so the address is static, derived from the VMID — and therefore known before the VM boots. Looking for it afterwards was absurd.
The Debian cloud image does not ship qemu-guest-agent, so Proxmox cannot tell
the address of a DHCP guest: it does not hand out the leases. The fallback is
the host's own neighbour table (ip neigh), which needs nothing from the guest.
The menu, entry by entry
The seventeen QEMU/KVM entries have their counterpart. Four of them are the
same code, because it is the same work: reopening the install monitoring,
the remote desktop tunnel, the Android emulator and the image catalog. They
reach Proxmox guests through the ~/.ssh/config entries that entry 13 writes,
with the Proxmox host as ProxyJump.
[1] Déployer une VM [8] Redimensionner un disque [15] Émulateur Android *
[2] Prévisualiser [9] Effacer des VM [16] Catalogue d'images *
[3] Télécharger une image [10] Nettoyer (orphelins) [17] Exemple (dry-run)
[4] Rouvrir le suivi * [11] Tester une VM (Odoo) [18] Changer d'hôte
[5] Lister (qm list) [12] Statistiques
[6] Adresse IP d'une VM [13] Configuration SSH
[7] Console d'une VM [14] Tunnel bureau distant *
* code partagé avec le menu QEMU/KVM
Verified
Deploying a VM inside a Proxmox that itself runs in a libvirt VM: image
downloaded on the host, internal bridge created, static address, cloud-init
user and key, disk resized, qm start. Then ssh vm-essai from the outside
reaches it through the jump — three nested levels. Resize 12G → 16G, delete
with --purge, orphan scan: all checked against Proxmox VE 9.2.11.