erplibre/script/proxmox
Mathieu Benoit e3138bea7e [FIX] proxmox : ne pas démarrer le pare-feu depuis l'extérieur
Une révision adversariale de la réparation à distance a rendu un constat que
ses TROIS lentilles — réseau, systemd, shell — ont trouvé indépendamment :
démarrer pve-firewall peut couper le ssh qui répare. Sa configuration vit dans
/var/lib/pve-cluster/config.db, donc elle est invisible tant que /etc/pve
n'est pas monté — c'est-à-dire exactement dans l'état qu'on répare. On
appliquerait des règles qu'on ne peut pas lire, sur la seule voie d'accès à la
machine.

Il n'est pas nécessaire au but : le stockage et le suivi demandent pve-cluster
et pvestatd, l'interface web pveproxy. Il repartira au prochain démarrage,
quand /etc/pve sera monté à temps. Le retirer de la liste coûte donc rien et
supprime le seul geste qui pouvait isoler un hôte.

Deux autres constats de la même révision, également réels.

Le gel de cloud-init gardait sur l'EXISTENCE du fichier. Or « printf … > » le
TRONQUE avant d'écrire : une coupure au mauvais moment laisse zéro octet, et
la garde annonce « déjà gelé » pour toujours. cloud-init continue de remettre
127.0.1.1 à chaque démarrage et le défaut redevient invisible — celui-là même
que ce code existe pour supprimer. La garde porte maintenant sur le CONTENU.

Et les adresses de lien-local passaient pour routables. Mesuré : « hostname
--ip-address » peut ne rendre QUE des fe80::, et une APIPA en 169.254 passait
le seul test « ne commence pas par 127. ». pmxcfs n'a alors rien
d'utilisable, mais le diagnostic concluait l'inverse et renvoyait vers
journalctl au lieu de /etc/hosts.

Enfin « la sonde n'a pas répondu » n'est plus lu comme « rien n'est monté » :
un dépassement de délai rend les mêmes vides, et on affirmait une cause qu'on
n'avait pas constatée.

--- EN ---

An adversarial review of the remote repair produced one finding all THREE of
its lenses — network, systemd, shell — reached independently: starting
pve-firewall can cut the ssh doing the repair. Its configuration lives in
/var/lib/pve-cluster/config.db, so it is invisible while /etc/pve is unmounted
— exactly the state being repaired. We would apply rules we cannot read, over
the machine's only way in.

It is not needed for the goal: storage and monitoring need pve-cluster and
pvestatd, the web interface pveproxy. It will come back at the next boot, when
/etc/pve mounts in time. Removing it from the list costs nothing and removes
the one gesture that could isolate a host.

Two more findings from the same review, equally real.

The cloud-init freeze guarded on the file's EXISTENCE. But "printf … >"
TRUNCATES before writing: an ill-timed cut leaves zero bytes, and the guard
then reports "already frozen" forever. cloud-init keeps putting 127.0.1.1 back
at every boot and the defect becomes invisible again — the very one this code
exists to remove. The guard now looks at the CONTENT.

And link-local addresses counted as routable. Measured: "hostname
--ip-address" can return ONLY fe80:: entries, and an APIPA 169.254 passed the
lone "does not start with 127." test. pmxcfs then has nothing usable, yet the
diagnosis concluded the opposite and pointed at journalctl instead of
/etc/hosts.

Finally "the probe did not answer" is no longer read as "nothing is mounted": a
timeout returns the same emptiness, and we were asserting a cause we had not
measured.

Assisted-by: Claude Opus 5
(cherry picked from commit fa9fb729d82e8d1a8e4fb549cc8061c7281b5dcb)
2026-08-29 01:53:03 -04:00
..
install_proxmox.sh [FIX] proxmox : ne pas démarrer le pare-feu depuis l'extérieur 2026-08-29 01:53:03 -04:00
proxmox_deploy.py [FIX] proxmox : ne pas démarrer le pare-feu depuis l'extérieur 2026-08-29 01:53:03 -04:00
README.base.md [ADD] proxmox : déployer des VM sur un hôte Proxmox distant 2026-08-23 05:54:38 -04:00
README.fr.md [ADD] proxmox : déployer des VM sur un hôte Proxmox distant 2026-08-23 05:54:38 -04:00
README.md [ADD] proxmox : déployer des VM sur un hôte Proxmox distant 2026-08-23 05:54:38 -04:00

Deploying VMs on a Proxmox VE host

Two different things live in this directory:

  • install_proxmox.sh turns a Debian into a Proxmox hypervisor. See script/qemu/README.md, which documents the proxmox distro of the deployment catalog.
  • proxmox_deploy.py deploys VMs on such a host, from TODO › Execute › Deploy › Proxmox VE, right under QEMU/KVM.

The whole difference: the hypervisor is elsewhere

With QEMU/KVM, the hypervisor is the machine running the script. With Proxmox it is somewhere else, so the first question is which host — and the answer is remembered for the session. Three ways, all offered by the menu:

  1. From the local QEMU VMs — a proxmox VM deployed here. Its address comes from the DHCP lease, nothing to retype.
  2. By address — user@host, plus an optional SSH jump.
  3. From ~/.ssh/config — the alias already carries user, port and ProxyJump; nothing else is asked.

The chosen host is then checked, not assumed: pveversion proves it is a Proxmox, id -u and sudo -n true decide whether commands need sudo, and an unknown SSH host key is offered for recording (with ssh-keyscan, never by disabling the check — a hypervisor is not a throwaway VM).

# Ce que l'outil envoie, et qu'on peut rejouer à la main :
ssh erplibre@pve1 sudo sh -c 'qm list'
ssh erplibre@pve1 sudo sh -c 'pvesm status --content images'

Why SSH and qm, not the REST API

The API needs a token or a ticket to create and renew. qm is the path every Proxmox administrator knows, the repository already manages SSH access (~/.ssh/config, ProxyJump, keys), and the commands stay readable in the log — so they can be replayed by hand. That is how every failure of this module was diagnosed.

sudo sh -c '<whole command>' and not sudo <command>: these commands are sequences and redirections. Prefixing with sudo would elevate only the first word, and the redirection would still be the unprivileged shell's.

Four traps met on a real host

A Proxmox installed on Debian has no vmbr0 — the ISO installer creates one, that procedure does not. And qm create requires a bridge.

The menu offers an internal bridge (vmbr0, 10.10.10.1/24, NAT through the uplink). Never adding the physical NIC to a bridge is deliberate: that moves the host address and cuts the SSH session in progress — remotely, there is no way back. A LAN-facing bridge is printed as a stanza to apply from a console.

On an internal bridge no DHCP answers, so the address is static, derived from the VMID — and therefore known before the VM boots. Looking for it afterwards was absurd.

The Debian cloud image does not ship qemu-guest-agent, so Proxmox cannot tell the address of a DHCP guest: it does not hand out the leases. The fallback is the host's own neighbour table (ip neigh), which needs nothing from the guest.

The menu, entry by entry

The seventeen QEMU/KVM entries have their counterpart. Four of them are the same code, because it is the same work: reopening the install monitoring, the remote desktop tunnel, the Android emulator and the image catalog. They reach Proxmox guests through the ~/.ssh/config entries that entry 13 writes, with the Proxmox host as ProxyJump.

[1] Déployer une VM        [8]  Redimensionner un disque   [15] Émulateur Android *
[2] Prévisualiser          [9]  Effacer des VM             [16] Catalogue d'images *
[3] Télécharger une image  [10] Nettoyer (orphelins)       [17] Exemple (dry-run)
[4] Rouvrir le suivi *     [11] Tester une VM (Odoo)       [18] Changer d'hôte
[5] Lister (qm list)       [12] Statistiques
[6] Adresse IP d'une VM    [13] Configuration SSH
[7] Console d'une VM       [14] Tunnel bureau distant *
                                        * code partagé avec le menu QEMU/KVM

Verified

Deploying a VM inside a Proxmox that itself runs in a libvirt VM: image downloaded on the host, internal bridge created, static address, cloud-init user and key, disk resized, qm start. Then ssh vm-essai from the outside reaches it through the jump — three nested levels. Resize 12G → 16G, delete with --purge, orphan scan: all checked against Proxmox VE 9.2.11.