erplibre/LongTest
Mathieu Benoit adf0f275d4 [FIX] proxmox : apt-daily tient le verrou au démarrage
Trois pannes trouvées en LANÇANT la descente, aucune vue en la relisant — ni
par moi, ni par l'attaque adversariale.

Le premier « apt update » d'une image cloud échoue sur un verrou qui n'est pas
celui qu'on croit. Mesuré une seconde après le premier ssh :

  E: Could not get lock /var/lib/apt/lists/lock.
     It is held by process 1026 (apt-get)

Ce n'est pas cloud-init — « status --wait » avait rendu la main. C'est
apt-daily, le minuteur de Debian, qui se déclenche au démarrage. Et le verrou
des LISTES n'est pas couvert par « DPkg::Lock::Timeout », que l'installeur
réglait pourtant déjà à 600 s : cette attente ne vaut que pour dpkg. On arrête
donc les minuteurs, puis on RÉESSAIE — arrêter une unité n'interrompt pas
l'apt-get déjà en vol. Le défaut touchait tout déploiement Proxmox, pas
seulement ce test.

Le premier étage n'avait pas d'alias ssh. « deploy_qemu.py » en ligne de
commande n'écrit pas d'entrée ~/.ssh/config — le menu le fait, la CLI non. La
descente aurait attendu son plein délai avant de conclure « jamais joignable »
sur une VM qui répondait à son adresse. Elle l'écrit maintenant elle-même,
depuis l'adresse résolue, et refuse d'avancer si la VM n'en a pas.

Et la réserve de l'hôte est proportionnelle. Quatre gigaoctets sur une machine
de soixante, c'était 6 % laissés au système : le jour où les invités touchent
vraiment leur mémoire, c'est l'hôte qui part en swap — et la mesure serait
celle du swap, pas de l'imbrication. Un huitième, avec le plancher d'avant
pour les petites machines.

Ce que la descente a établi en trois étages : 392 s, 644 s, 1120 s, soit 1,7
fois par étage. Puis la poignée de main ssh passe de 77 à 1664 secondes au
quatrième — vingt fois d'un seul cran. Le coude est là.

Et le mur que j'avais pris pour une limite d'imbrication n'en était pas une.
La VM qui gelait au quatrième étage avait douze vCPU ; celle-ci en a deux et
elle passe, en écrivant. C'était une limite de parallélisme SOUS imbrication —
exactement ce que l'algorithme borne, vérifié pour la première fois plutôt que
supposé.

--- EN ---

Three faults found by RUNNING the descent, none seen by reading it — neither
by me nor by the adversarial attack.

A cloud image's first "apt update" fails on a lock that is not the one you
expect. Measured one second after the first ssh:

  E: Could not get lock /var/lib/apt/lists/lock.
     It is held by process 1026 (apt-get)

It is not cloud-init — "status --wait" had returned. It is apt-daily, Debian's
timer, firing at boot. And the LISTS lock is not covered by
"DPkg::Lock::Timeout", which the installer already set to 600 s: that wait
only applies to dpkg. So we stop the timers, then RETRY — stopping a unit does
not interrupt the apt-get already in flight. The defect affected every Proxmox
deployment, not just this test.

The first level had no ssh alias. "deploy_qemu.py" on the command line does not
write a ~/.ssh/config entry — the menu does, the CLI does not. The descent
would have waited its full timeout before concluding "never reachable" about a
VM answering at its address. It now writes the entry itself, from the resolved
address, and refuses to proceed if the VM has none.

And the host's reserve is proportional. Four gigabytes on a sixty-gigabyte
machine left 6 % to the system: the day the guests really touch their memory,
the host swaps — and the measurement would be of swap, not of nesting. One
eighth now, keeping the old floor for small machines.

What the descent established over three levels: 392 s, 644 s, 1120 s — 1.7x per
level. Then the ssh handshake goes from 77 to 1664 seconds at the fourth:
twenty times in one step. That is the elbow.

And the wall I had taken for a nesting limit was not one. The VM that froze at
the fourth level had twelve vCPU; this one has two and it gets through,
writing. It was a limit of parallelism UNDER nesting — exactly what the
algorithm caps, verified for the first time rather than assumed.

Assisted-by: Claude Opus 5
(cherry picked from commit 7f86562cbd8f10017dcb88fe4272efc162cbccbc)
2026-08-29 01:53:03 -04:00
..
deep_proxmox.py [FIX] proxmox : apt-daily tient le verrou au démarrage 2026-08-29 01:53:03 -04:00
README.base.md [ADD] LongTest : jusqu'à quel étage un Proxmox imbriqué tient-il 2026-08-29 01:53:03 -04:00
README.fr.md [ADD] LongTest : jusqu'à quel étage un Proxmox imbriqué tient-il 2026-08-29 01:53:03 -04:00
README.md [ADD] LongTest : jusqu'à quel étage un Proxmox imbriqué tient-il 2026-08-29 01:53:03 -04:00

LongTest — tests that create real machines

These are not unit tests. They create virtual machines, install systems on them, and take hours. They live here and not in test/, which the unit runner sweeps: ./script/test/run_unit_test.sh must stay runnable in seconds on any machine, including one without virtualisation.

Run them from the menu — TODO › Execute › Test › Long tests — or directly.

deep_proxmox.py — how deep does Proxmox-in-Proxmox go?

The practicable nesting depth cannot be deduced, only measured. A manual measurement found, at the fourth level, a guest 36 times slower than real time — 583 seconds of wall clock for 16 seconds of guest time, each ACPI line taking a second — then a guest kernel frozen at the same byte whatever the resources. A number obtained once, on one machine, is not a number: this script redoes it on demand and says exactly where it breaks.

./LongTest/deep_proxmox.py --depth 10 --dry-run   # the plan, nothing created
./LongTest/deep_proxmox.py --depth 10             # hours
./LongTest/deep_proxmox.py --detruire             # undo it

The descent is uniform. Every level, the first included, goes through the same six steps: create, wait for ssh, install Proxmox, reboot and check the kernel, bring pmxcfs back up, check the storage. Only creation differs — libvirt locally, qm afterwards.

It sends our install_proxmox.sh over scp instead of letting the VM clone the repository: it is our code we want to exercise, and the remote is often behind the checkout — a fix absent from the remote made the same defect "come back" on three VMs in a row.

The resource algorithm

Two things run out going down, and a third degrades. What runs out is arithmetic, and script/proxmox/nesting.py computes it:

  • memory — each level keeps what its own daemons need (pve-cluster, pvestatd, pvedaemon, pveproxy) before handing the rest down;
  • disk — the child's disk lives inside the parent's, which must also hold its own system.

What degrades is measured, not assumed: past the second level, vendors document nothing. Hence one capped number — 2 vCPU for every nested level. Twelve vCPU at the fourth level froze the guest kernel in early boot; the same two progressed. Bringing twelve processors online costs as many round trips through the whole stack.

Memory is not capped: the same VM froze at the same byte with 9 GB and with 2 GB, so trimming it would gain nothing and starve the level below.

The plan is printed before anything is created, and the script never promises a depth it knows will not fit — better to announce six levels and reach six than to promise ten and die at the seventh without knowing why.