Attaqué par trois lentilles sur le code écrit, avant de le lancer pour de vrai. Deux fautes valaient à elles seules l'exercice. Il n'aurait JAMAIS fonctionné. L'installeur était lancé par « sh », or il porte « set -euo pipefail » et un shebang bash : sur Debian /bin/sh est dash, qui répond « set: Illegal option -o pipefail » et sort à la PREMIÈRE ligne. Chaque étage aurait échoué sur l'installation, à tous les coups. Et « --detruire » pouvait emporter une machine étrangère. Il prenait toute entrée ssh dont le nom CONTENAIT « deep-pve », puis sur son rebond détruisait toute VM dont le nom contenait « deep-pve » — une « deep-pve-lab » de production tombait dedans, et « --purge » emporte les disques. Son tri « du plus profond au plus haut » comptait les « + » de l'alias, or alias_etage remplace le « + » du parent par un « - » : chaque alias en portait exactement UN, le tri ne triait rien, et la destruction partait du plus HAUT — le disque du parent emportait ses enfants sans qu'on les ait nommés. Il ignorait « --dry-run », ne lisait aucun code de retour, concluait « ✓ défait », et le menu le lançait d'une touche. Il ne détruit plus que ce que le RAPPORT nomme : un couple (parent, VMID) par étage, du plus profond d'après le niveau lu, égalité stricte du nom, arrêt CONSTATÉ avant destruction, codes de retour lus, et une confirmation par « OUI » après la liste. Six autres constats, tous réels. Le redémarrage se prouve par btime et non par le seul noyau — rejoué sur un étage déjà installé, on validait un redémarrage qui n'avait pas eu lieu, exactement le piège corrigé la semaine dernière dans le suivi. La sonde de disponibilité ne demande plus sudo, sinon un sudo lent se lisait « jamais joignable en ssh ». Les délais suivent la profondeur : le script existe pour mesurer un ralentissement de 36x, et un plafond fixe déclarait échouée une installation qui avançait. L'adresse fixe est contrôlée AVANT de télécharger une image et de démarrer une VM. Le DNS de l'hôte suit la spec, sinon apt meurt sans rien expliquer. Et l'essai à blanc ne prétend plus avoir atteint quoi que ce soit — son rapport était indiscernable d'une réussite, JSON compris. L'algorithme aussi : profondeur 0 rendait un plan d'UN étage, donc « --depth 0 » créait une VM ; et sur un hôte de quatre cœurs le premier étage recevait UN vCPU quand son invité en recevait deux — un parent plus étroit que son enfant. Les tests mordent, prouvé par mutation : remplacer le calcul du premier étage par la valeur imbriquée les laissait verts. --- EN --- Attacked by three lenses on the written code, before running it for real. Two faults alone justified the exercise. It would NEVER have worked. The installer was run by "sh", yet it carries "set -euo pipefail" and a bash shebang: on Debian /bin/sh is dash, which answers "set: Illegal option -o pipefail" and exits on the FIRST line. Every level would have failed at install, every time. And "--detruire" could take a stranger's machine. It took every ssh entry whose name CONTAINED "deep-pve", then on its jump host destroyed every VM whose name contained "deep-pve" — a production "deep-pve-lab" fell in, and "--purge" takes the disks. Its "deepest first" sort counted the "+" in the alias, yet alias_etage replaces the parent's "+" with a "-": every alias had exactly ONE, the sort sorted nothing, and destruction started from the TOP — the parent's disk took its children with it, unnamed. It ignored "--dry-run", read no return code, concluded "✓ done", and the menu fired it on one key. It now destroys only what the REPORT names: a (parent, VMID) pair per level, deepest first by the recorded level, strict name equality, shutdown VERIFIED before destruction, return codes read, and a "OUI" confirmation after the list. Six more findings, all real. The reboot is proven by btime, not by the kernel alone — replayed on an already-installed level, we validated a reboot that never happened, exactly the trap fixed last week in the monitor. The liveness probe no longer asks for sudo, or a slow sudo read as "never reachable by ssh". Timeouts follow the depth: the script exists to measure a 36x slowdown, and a fixed ceiling declared failed an install that was progressing. The static address is checked BEFORE downloading an image and starting a VM. The host's DNS follows the spec, or apt dies explaining nothing. And the dry run no longer claims to have reached anything — its report was indistinguishable from a success, JSON included. The algorithm too: depth 0 returned a ONE-level plan, so "--depth 0" created a VM; and on a four-core host the first level got ONE vCPU while its guest got two — a parent narrower than its child. The tests bite, proven by mutation: replacing the first level's computation with the nested value left them green. Assisted-by: Claude Opus 5 (cherry picked from commit 64b8e5063bd7f420cdeb27b88f94043190b5ecd4) |
||
|---|---|---|
| .. | ||
| deep_proxmox.py | ||
| README.base.md | ||
| README.fr.md | ||
| README.md | ||
LongTest — tests that create real machines
These are not unit tests. They create virtual machines, install systems on
them, and take hours. They live here and not in test/, which the unit
runner sweeps: ./script/test/run_unit_test.sh must stay runnable in seconds
on any machine, including one without virtualisation.
Run them from the menu — TODO › Execute › Test › Long tests — or directly.
deep_proxmox.py — how deep does Proxmox-in-Proxmox go?
The practicable nesting depth cannot be deduced, only measured. A manual measurement found, at the fourth level, a guest 36 times slower than real time — 583 seconds of wall clock for 16 seconds of guest time, each ACPI line taking a second — then a guest kernel frozen at the same byte whatever the resources. A number obtained once, on one machine, is not a number: this script redoes it on demand and says exactly where it breaks.
./LongTest/deep_proxmox.py --depth 10 --dry-run # the plan, nothing created
./LongTest/deep_proxmox.py --depth 10 # hours
./LongTest/deep_proxmox.py --detruire # undo it
The descent is uniform. Every level, the first included, goes through the
same six steps: create, wait for ssh, install Proxmox, reboot and check the
kernel, bring pmxcfs back up, check the storage. Only creation differs —
libvirt locally, qm afterwards.
It sends our install_proxmox.sh over scp instead of letting the VM clone
the repository: it is our code we want to exercise, and the remote is often
behind the checkout — a fix absent from the remote made the same defect "come
back" on three VMs in a row.
The resource algorithm
Two things run out going down, and a third degrades. What runs out is
arithmetic, and script/proxmox/nesting.py computes it:
- memory — each level keeps what its own daemons need (
pve-cluster,pvestatd,pvedaemon,pveproxy) before handing the rest down; - disk — the child's disk lives inside the parent's, which must also hold its own system.
What degrades is measured, not assumed: past the second level, vendors document nothing. Hence one capped number — 2 vCPU for every nested level. Twelve vCPU at the fourth level froze the guest kernel in early boot; the same two progressed. Bringing twelve processors online costs as many round trips through the whole stack.
Memory is not capped: the same VM froze at the same byte with 9 GB and with 2 GB, so trimming it would gain nothing and starve the level below.
The plan is printed before anything is created, and the script never promises a depth it knows will not fit — better to announce six levels and reach six than to promise ten and die at the seventh without knowing why.