erplibre/script/proxmox
Mathieu Benoit 4bc2fa6097 [ADD] proxmox : l'écran remet pmxcfs debout lui-même
Le conseil « rejouer install_proxmox.sh sur l'hôte » ne pouvait PAS marcher :
la VM clone le dépôt distant, donc sa copie du script est celle du distant —
tant que le correctif n'y est pas, celle qui ne corrige rien. Trois hôtes de
suite sont tombés dessus, avec le même message inutile. L'écran répare donc :
gel de cloud-init, réécriture de /etc/hosts, relance des unités, constat du
montage — le pendant exact de l'offre de créer un pont.

Écrit, puis ATTAQUÉ par trois lentilles sur le code réel. Ce qu'elles ont
mesuré valait la peine.

/etc/hosts se réécrivait en DEUX écritures — « sed -i » puis « printf >> » —
alors que la docstring promettait l'inverse. Sed refusé et ajout réussi, la
ligne 127.0.1.1 survivait EN PREMIER et la nôtre s'ajoutait une fois par
tentative ; sed réussi et ajout refusé, l'hôte perdait l'entrée de son nom, et
chaque sudo y attendait ensuite le résolveur. C'est maintenant un fichier
complet bâti dans un temporaire, VÉRIFIÉ, puis recopié — « cat > » et non
« mv », qui remplacerait l'inode et perdrait mode et propriétaire.

Le contrôle final s'en remettait à « getent hosts », qui réussit via mDNS même
quand rien n'a été écrit — et acceptait les fe80:: que notre propre code
rejette. Il relit désormais ce qui a été écrit.

awk remplace sed pour filtrer : « print » émet un saut de ligne, donc un
/etc/hosts non terminé par un — cloud-init n'en met pas — est normalisé. Sans
ça notre ligne se collait à la précédente et le nom du nœud partait sur
l'adresse d'une autre machine.

Trois autres, du même acabit. Les dépendants de pmxcfs sont relancés eux
aussi : actifs pendant la panne, ils échouaient sur ipcc_send_rec, et les
laisser donnait une GUI en « communication failure » juste après notre ✓. Un
silence du lien n'est plus lu comme une absence de montage. Et l'adresse n'est
mise en cause que si pve-cluster a réellement démarré.

Les tests exécutent les commandes au lieu de les relire, bouchons capables
d'ÉCHOUER : écriture refusée, fichier sans saut de ligne final, tabulations,
start qui rate, montage qui disparaît pendant la reconfirmation. Prouvé par
mutation — trois HOSTS-KO changés en HOSTS-OK font rougir le test.

--- EN ---

The advice "replay install_proxmox.sh on the host" could NOT work: the VM
clones the remote, so its copy of the script is the remote's — while the fix
is not there, the one that fixes nothing. Three hosts in a row hit it with the
same useless message. So the screen repairs: freeze cloud-init, rewrite
/etc/hosts, restart the units, verify the mount — the exact counterpart of the
offer to create a bridge.

Written, then ATTACKED by three lenses on the real code. What they measured
was worth it.

/etc/hosts was rewritten in TWO writes — "sed -i" then "printf >>" — while the
docstring promised the opposite. Sed refused and append succeeded: the
127.0.1.1 line survived FIRST and ours was added once per attempt; sed
succeeded and append refused: the host lost its own name entry, and every sudo
then waited on the resolver. It is now a complete file built in a temporary,
VERIFIED, then copied over — "cat >" not "mv", which would replace the inode
and lose mode and owner.

The final check relied on "getent hosts", which succeeds via mDNS even when
nothing was written — and accepted the fe80:: our own code rejects. It now
re-reads what was written.

awk replaces sed for filtering: "print" emits a newline, so an /etc/hosts with
no final one — cloud-init omits it — gets normalised. Without that our line
glued onto the previous one and the node's name pointed at another machine's
address.

Three more of the same kind. pmxcfs's dependents are restarted too: active
throughout the outage, they failed on ipcc_send_rec, and leaving them gave a
GUI in "communication failure" right after our ✓. A silent link is no longer
read as a missing mount. And the address is only blamed if pve-cluster
actually started.

The tests execute the commands instead of reading them, with stubs able to
FAIL: refused write, file with no final newline, tabs, a start that fails, a
mount that vanishes during reconfirmation. Proven by mutation — three
HOSTS-KO turned into HOSTS-OK make the test go red.

Assisted-by: Claude Opus 5
(cherry picked from commit d4f9358c6cb562029cc2ca9eb478d80c6a0a19a4)
2026-08-29 01:53:03 -04:00
..
install_proxmox.sh [FIX] proxmox : ne pas démarrer le pare-feu depuis l'extérieur 2026-08-29 01:53:03 -04:00
proxmox_deploy.py [ADD] proxmox : l'écran remet pmxcfs debout lui-même 2026-08-29 01:53:03 -04:00
README.base.md [ADD] proxmox : déployer des VM sur un hôte Proxmox distant 2026-08-23 05:54:38 -04:00
README.fr.md [ADD] proxmox : déployer des VM sur un hôte Proxmox distant 2026-08-23 05:54:38 -04:00
README.md [ADD] proxmox : déployer des VM sur un hôte Proxmox distant 2026-08-23 05:54:38 -04:00

Deploying VMs on a Proxmox VE host

Two different things live in this directory:

  • install_proxmox.sh turns a Debian into a Proxmox hypervisor. See script/qemu/README.md, which documents the proxmox distro of the deployment catalog.
  • proxmox_deploy.py deploys VMs on such a host, from TODO › Execute › Deploy › Proxmox VE, right under QEMU/KVM.

The whole difference: the hypervisor is elsewhere

With QEMU/KVM, the hypervisor is the machine running the script. With Proxmox it is somewhere else, so the first question is which host — and the answer is remembered for the session. Three ways, all offered by the menu:

  1. From the local QEMU VMs — a proxmox VM deployed here. Its address comes from the DHCP lease, nothing to retype.
  2. By address — user@host, plus an optional SSH jump.
  3. From ~/.ssh/config — the alias already carries user, port and ProxyJump; nothing else is asked.

The chosen host is then checked, not assumed: pveversion proves it is a Proxmox, id -u and sudo -n true decide whether commands need sudo, and an unknown SSH host key is offered for recording (with ssh-keyscan, never by disabling the check — a hypervisor is not a throwaway VM).

# Ce que l'outil envoie, et qu'on peut rejouer à la main :
ssh erplibre@pve1 sudo sh -c 'qm list'
ssh erplibre@pve1 sudo sh -c 'pvesm status --content images'

Why SSH and qm, not the REST API

The API needs a token or a ticket to create and renew. qm is the path every Proxmox administrator knows, the repository already manages SSH access (~/.ssh/config, ProxyJump, keys), and the commands stay readable in the log — so they can be replayed by hand. That is how every failure of this module was diagnosed.

sudo sh -c '<whole command>' and not sudo <command>: these commands are sequences and redirections. Prefixing with sudo would elevate only the first word, and the redirection would still be the unprivileged shell's.

Four traps met on a real host

A Proxmox installed on Debian has no vmbr0 — the ISO installer creates one, that procedure does not. And qm create requires a bridge.

The menu offers an internal bridge (vmbr0, 10.10.10.1/24, NAT through the uplink). Never adding the physical NIC to a bridge is deliberate: that moves the host address and cuts the SSH session in progress — remotely, there is no way back. A LAN-facing bridge is printed as a stanza to apply from a console.

On an internal bridge no DHCP answers, so the address is static, derived from the VMID — and therefore known before the VM boots. Looking for it afterwards was absurd.

The Debian cloud image does not ship qemu-guest-agent, so Proxmox cannot tell the address of a DHCP guest: it does not hand out the leases. The fallback is the host's own neighbour table (ip neigh), which needs nothing from the guest.

The menu, entry by entry

The seventeen QEMU/KVM entries have their counterpart. Four of them are the same code, because it is the same work: reopening the install monitoring, the remote desktop tunnel, the Android emulator and the image catalog. They reach Proxmox guests through the ~/.ssh/config entries that entry 13 writes, with the Proxmox host as ProxyJump.

[1] Déployer une VM        [8]  Redimensionner un disque   [15] Émulateur Android *
[2] Prévisualiser          [9]  Effacer des VM             [16] Catalogue d'images *
[3] Télécharger une image  [10] Nettoyer (orphelins)       [17] Exemple (dry-run)
[4] Rouvrir le suivi *     [11] Tester une VM (Odoo)       [18] Changer d'hôte
[5] Lister (qm list)       [12] Statistiques
[6] Adresse IP d'une VM    [13] Configuration SSH
[7] Console d'une VM       [14] Tunnel bureau distant *
                                        * code partagé avec le menu QEMU/KVM

Verified

Deploying a VM inside a Proxmox that itself runs in a libvirt VM: image downloaded on the host, internal bridge created, static address, cloud-init user and key, disk resized, qm start. Then ssh vm-essai from the outside reaches it through the jump — three nested levels. Resize 12G → 16G, delete with --purge, orphan scan: all checked against Proxmox VE 9.2.11.