Proxmox VE n'existe qu'après un redémarrage : tant que la VM tourne le noyau
de son image cloud, elle n'a aucun module netfilter — ni pont NAT, ni invité.
install_proxmox.sh pose le noyau puis s'arrête, à raison, car lancé par ssh un
reboot couperait sa session et ferait passer l'installation pour un échec. On
le découvrait donc des jours plus tard, en créant un pont.
Le redémarrage revient à l'enveloppe de lancement, qui tourne sur NOTRE
machine et survit à celui de la VM : installation, reboot, attente, puis
vérification du noyau. Le ✅ ne s'écrit qu'après, et il veut donc dire
« hyperviseur utilisable ».
Trois choix méritent d'être dits. On ne redémarre qu'après un SUCCÈS —
redémarrer après un échec effacerait la seule machine sur laquelle on pouvait
chercher. On n'attend pas que ssh « revienne » mais que « uname -r » porte le
motif attendu : sshd répond encore une seconde ou deux après l'ordre, et on
lirait l'ancien noyau en croyant avoir la réponse. Et l'absence du noyau
attendu est un vrai ÉCHEC, pas un avertissement.
Le shell est exécuté par les tests, ssh bouchonné, dans les quatre cas — dont
celui où les deux premières lectures rendent l'ancien noyau. Un garde qu'on ne
sait pas éprouver s'ouvre le jour où il casse.
La note du sommaire ne paraît plus que sans suivi, où rien ne redémarre :
réclamer un redémarrage déjà fait est une consigne fausse.
--- EN ---
Proxmox VE only exists after a reboot: while the VM runs its cloud image's
kernel it has no netfilter module — no NAT bridge, no guest.
install_proxmox.sh installs the kernel then stops, rightly, since run over ssh
a reboot would cut its own session and make the install look failed. So you
found out days later, when creating a bridge.
The reboot moves to the launch wrapper, which runs on OUR machine and survives
the VM's: install, reboot, wait, then verify the kernel. The ✅ is written only
after, and therefore means "usable hypervisor".
Three choices worth stating. We reboot only after SUCCESS — rebooting after a
failure would wipe the one machine you could investigate. We do not wait for
ssh to "come back" but for "uname -r" to carry the expected pattern: sshd
answers for another second or two after the order, and we would read the old
kernel believing we had the answer. And a missing expected kernel is a real
FAILURE, not a warning.
The shell is executed by the tests, ssh stubbed, in all four cases — including
the one where the first two reads return the old kernel. A guard you cannot
exercise opens the day it breaks.
The summary note now appears only without monitoring, where nothing reboots:
asking for a reboot already done is a false instruction.
Assisted-by: Claude Opus 5
Sur trois VM d'un même Proxmox, une seule avait ses colonnes vides — et les
deux autres montraient les chiffres d'une AUTRE machine. Deux fautes, dont
une était le miroir d'un correctif précédent.
Le code de sortie de la suite distante est celui de son DERNIER maillon, la
sonde Odoo. Tant qu'Odoo n'écoute pas — c'est-à-dire pendant TOUTE
l'installation, précisément quand on regarde — la boucle finit en échec et le
relevé, parfait, était jeté. On avait corrigé l'erreur inverse, un code 0 pris
pour une réponse ; exiger 0 était la même faute retournée. Seule une liste de
ressources analysable prouve une réponse.
Pendant ce temps, « virsh domstats » indexe par NOM, et un nom se partage :
les deux VM qui avaient un homonyme LOCAL affichaient ses chiffres. Mesuré —
1,5 Gio de RAM sur 12 et 58 Gio de disque sur 65, quand la vraie tournait avec
3 Gio et 25. Les relevés locaux d'une VM qui vit ailleurs sont donc retirés
AVANT d'ajouter ceux de l'hôte : un hôte muet laisse la colonne VIDE, ce qui
est vrai. Une colonne vide se remarque ; une colonne juste et fausse, non.
L'alias enfin. Prendre le nom court quand il se trouvait libre donnait un parc
incohérent : sur ce même déploiement, deux VM ont reçu « hôte+vm » — leurs
noms étaient pris par des domaines locaux — et la troisième son nom court. Une
convention qui dépend de ce qui traîne dans le fichier n'est pas une
convention. Le nom chaîné est systématique.
--- EN ---
Of three VMs on one Proxmox, only one had empty columns — and the other two
showed ANOTHER machine's figures. Two defects, one the mirror of an earlier
fix.
A remote pipeline's exit code is its LAST link's, the Odoo probe. While Odoo
is not listening — that is, during the WHOLE install, exactly when you are
watching — the loop ends in failure and the reading, perfectly good, was
thrown away. We had fixed the opposite error, a 0 taken for an answer;
demanding 0 was the same mistake reversed. Only a parsable resource list
proves an answer.
Meanwhile "virsh domstats" indexes by NAME, and a name is shared: the two VMs
with a LOCAL namesake displayed its figures. Measured — 1.5 GiB of RAM out of
12 and 58 GiB of disk out of 65, while the real one ran on 3 GiB and 25. Local
readings for a VM that lives elsewhere are therefore dropped BEFORE the host's
are added: a silent host leaves the column EMPTY, which is true. An empty
column gets noticed; a plausible wrong one does not.
The alias, finally. Taking the short name while it happened to be free gave an
inconsistent fleet: in that same deployment two VMs got "host+vm" — their
names were held by local domains — and the third its short name. A convention
that depends on what happens to sit in the file is not a convention. The
chained name is now systematic.
Assisted-by: Claude Opus 5
Le tableau de bord se rouvre sur un manifeste passé — c'est fait pour, les
installations partent détachées. Mais un nom de domaine se réemploie et un
VMID libéré est RÉATTRIBUÉ : effacer « le 101 » d'un run de mars, c'est
effacer ce qui porte le 101 aujourd'hui, et « erplibre-ubuntu-2604 » de mars
n'est pas celui d'aujourd'hui. Même famille que tout le reste — on jugeait
sur le nom, avec ici la pire conséquence.
La commande porte donc son garde, et non l'écran : elle protège ainsi tous
ses appelants, et la vérification se fait SUR la machine, à l'instant
d'effacer. Sur Proxmox, le VMID doit encore porter ce nom. En local, l'UUID
du domaine — relevé au lancement, seul instant où l'on sait que ce nom
désigne bien cette machine-là. Un manifeste écrit avant ce correctif n'en a
pas : il retombe sur la protection d'avant plutôt que de bloquer.
Le garde du VMID est une fonction à part, exécutable telle quelle. Il
traverse deux « shlex.quote » avant d'atteindre un dash, et un garde qu'on ne
sait pas éprouver s'OUVRE le jour où il casse. Vérifié sur erplibre-proxmox-9
sans rien détruire : le VMID 100 refusé sous un nom périmé, accepté sous le
sien.
--- EN ---
The dashboard reopens on a past manifest — by design, since installs run
detached. But a domain name gets reused and a freed VMID is REASSIGNED:
deleting "the 101" from a March run deletes whatever holds 101 today, and
March's "erplibre-ubuntu-2604" is not today's. Same family as the rest — we
judged by name, here with the worst consequence.
The command carries its guard, not the screen: that protects every caller,
and the check happens ON the machine, at the moment of deletion. On Proxmox
the VMID must still bear that name. Locally, the domain's UUID — recorded at
launch, the only moment we know that name means that machine. A manifest
written before this fix has none: it falls back to the previous protection
rather than blocking.
The VMID guard is its own function, runnable as is. It crosses two
"shlex.quote" layers before reaching a dash, and a guard you cannot exercise
OPENS the day it breaks. Verified on erplibre-proxmox-9 without destroying
anything: VMID 100 refused under a stale name, accepted under its own.
Assisted-by: Claude Opus 5
La confirmation de suppression promettait à TOUTE VM « son disque qcow2
EFFACÉ », puis nommait /var/lib/libvirt/images/<nom>.qcow2. Sur une VM
Proxmox ce fichier n'existe pas — au mieux, au pire c'est celui d'une autre
VM du même nom. C'est la peur exacte qui avait fait remonter le nettoyage.
Elle nomme désormais l'hôte, le VMID et « qm destroy ».
« Console de l'hyperviseur » lisait le port par « virsh vncdisplay ». Un
Proxmox n'a pas de libvirt : l'échec se lisait « écran fermé » et on
conseillait « sudo virsh edit » sur une machine sans ce binaire. Ce n'est pas
un écran fermé, c'est la mauvaise question — Proxmox sert le sien par un
ticket. Les deux vrais chemins sont nommés : la console série, l'interface
web par tunnel.
La colonne Odoo était un 🟢 acquis pour toujours : « Odoo ne redescend pas en
cours d'install » est faux — le service redémarre au moins une fois, et il
lui arrive de mourir. Elle est relue, gratuitement sur Proxmox, toutes les
trente secondes ailleurs. Un hôte muet reste distinct d'un Odoo tombé.
Enfin « Versions principales » (F7) manquait à l'écran Proxmox, qui affiche
pourtant le même catalogue. Les trois gestes du catalogue vivent maintenant
dans le socle du plan, où ils ne peuvent plus diverger.
--- EN ---
The delete confirmation promised EVERY VM "its qcow2 disk ERASED", then named
/var/lib/libvirt/images/<name>.qcow2. On a Proxmox VM that file does not
exist — at best; at worst it is another VM's, of the same name. That is the
very fear that surfaced the cleanup report. It now names the host, the VMID
and "qm destroy".
"Hypervisor console" read the port through "virsh vncdisplay". Proxmox has no
libvirt: the failure read as "screen closed" and we advised "sudo virsh edit"
on a machine without that binary. It is not a closed screen, it is the wrong
question — Proxmox serves its own by ticket. Both real paths are named: the
serial console, the web interface through a tunnel.
The Odoo column was a 🟢 acquired forever: "Odoo does not go back down during
the install" is false — the service restarts at least once, and it does die.
It is re-read, free on Proxmox, every thirty seconds elsewhere. A silent host
stays distinct from a dead Odoo.
Finally "Main versions" (F7) was missing from the Proxmox screen, which shows
the same catalog. The catalog's three gestures now live in the plan
foundation, where they can no longer diverge.
Assisted-by: Claude Opus 5
Trois autres chemins menaient au 🗑 sur un seul incident, et « effacée » gèle
la ligne pour de bon. Un « virsh list » en échec condamnait TOUT le parc
local. Un statut Proxmox hors des trois attendus — prelaunch, suspended,
internal-error — passait pour une disparition. Et le code de sortie de la
suite distante est celui de son DERNIER maillon : un pvesh en panne se lisait
« l'hôte a répondu sans elle ». Ce qui prouve une réponse, c'est désormais une
liste de ressources analysable.
Le plan annonçait « 25G » quand « qm resize » recevait 20 : la marge
d'ERPLibre se perdait en route, la VM naissait trop petite. « Changer l'état »
choisissait par NOM, or seul le VMID est unique sur un hôte — cocher une VM
en éteignait deux homonymes. L'entrée 13 volait son alias à une VM locale du
même nom. Enfin le déploiement par QUESTIONS avait vieilli seul : il partage
maintenant l'épilogue de l'écran, donc le guide, l'alias protégé, les colonnes
vivantes et le sommaire.
--- EN ---
Three more paths led to 🗑 on a single incident, and "deleted" freezes the row
for good. One failing "virsh list" condemned the WHOLE local fleet. A Proxmox
status outside the three expected ones — prelaunch, suspended, internal-error
— passed for a disappearance. And a remote pipeline's exit code is its LAST
link's: a broken pvesh read as "the host answered without it". Proof of an
answer is now a parsable resource list.
The plan announced "25G" while "qm resize" got 20: ERPLibre's margin was lost
on the way and the VM was born too small. "Change state" selected by NAME,
yet only the VMID is unique on a host — ticking one VM shut down two
namesakes. Menu entry 13 stole its alias from a local VM of the same name.
Finally the QUESTION-driven deployment had aged alone: it now shares the
screen's epilogue — guide, protected alias, live columns and summary.
Assisted-by: Claude Opus 5
Rapporté : une VM Arch à peine déployée sur Proxmox s'affichait 🗑 dès le
premier tour. « Effacée » est un état TERMINAL — la ligne gèle et ne revient
jamais — et il se déduisait d'UN relevé manquant. Or l'hôte peut être occupé,
la VM en train de naître, le relevé en cache d'avant sa création. On distingue
désormais « l'hôte n'a pas répondu » (on ne sait rien) de « l'hôte a répondu
sans elle » (on compte, trois fois), et la case part de « - » plutôt que d'un
sablier qui affirmerait qu'on attend quelque chose.
L'écran Proxmox n'offrait pas le choix de l'interpréteur Python : il envoyait
donc toujours « automatique », et comme mise n'est jamais installé d'office,
c'était pyenv — qui COMPILE Python depuis le tar.xz. Le choix existe
maintenant des deux côtés, avec le même garde-fou : rien n'est imposé quand
aucune architecture retenue n'est servie par mise.
--- EN ---
Reported: an Arch VM barely deployed on Proxmox showed 🗑 on the very first
pass. "Deleted" is a TERMINAL state — the row freezes and never comes back —
and it was inferred from ONE missing reading. Yet the host may be busy, the VM
may be starting, the reading may be cached from before it existed. We now tell
"the host did not answer" (we know nothing) from "the host answered without
it" (count, three times), and the cell starts at "-" rather than an hourglass
claiming we await something.
The Proxmox screen offered no Python interpreter choice: it therefore always
sent "automatic", and since mise is never installed by default, that meant
pyenv — which COMPILES Python from the tar.xz. The choice now exists on both
sides, with the same guard: nothing is imposed when no selected architecture
is served by mise.
Assisted-by: Claude Opus 5
La colonne Odoo testait le port 8069 depuis le poste : une VM sur pont
interne n'y répond jamais, elle restait « — » quel que soit l'état d'Odoo.
Le test part maintenant DE L'HÔTE, glissé dans l'appel des statistiques déjà
payé — aucun aller-retour de plus. Éprouvé sur une VM d'essai servant sur
8069 : 🟢, et « — » sur la voisine qui n'a rien.
Deux dernières actions visaient encore la mauvaise machine. La touche « w »
ouvrait une page morte : l'adresse d'un pont interne n'est pas routable
d'ici, elle passe donc par un tunnel local le temps de la visite — tenu par
son PID, car « pkill -f <motif> » tuait le shell qui l'avait lancé, le motif
figurant dans sa propre ligne de commande. Et la SUPPRESSION appelait
« virsh undefine <nom> », qui aurait effacé le domaine local homonyme : elle
passe par « qm destroy <vmid> » sur l'hôte.
--- EN ---
The Odoo column probed port 8069 from the workstation: a VM on an internal
bridge never answers there, so it stayed "—" whatever Odoo was doing. The
probe now runs FROM THE HOST, folded into the stats call already paid for —
no extra round trip. Proven on a test VM serving on 8069: 🟢, and "—" on the
neighbour that serves nothing.
Two last actions still aimed at the wrong machine. Key "w" opened a dead
page: an internal-bridge address is not routable from here, so it now goes
through a local tunnel for the length of the visit — held by its PID, since
"pkill -f <pattern>" killed the very shell that had launched it, the pattern
being in its own command line. And DELETION called "virsh undefine <name>",
which would have erased the homonymous local domain: it now goes through
"qm destroy <vmid>" on the host.
Assisted-by: Claude Opus 5
« s » ouvrait encore la VM locale homonyme : la vue de progression n'avait
que le NOM de la VM, et l'entrée ~/.ssh/config n'existe pas encore à ce
moment. Le déploiement lui passe maintenant « ssh -J <hôte> user@<ip> », qui
ne dépend de rien. Deux voisines du même défaut, jamais rapportées mais aussi
graves : la console ouvrait « virsh console <nom> » — celle de la VM LOCALE —
et la pause suspendait la locale. Les deux passent par le VMID sur l'hôte.
La colonne Disque annonçait « 6.0G/6.0G » sur une VM qui n'avait écrit que
1,2 Go : « du -sb » rend la taille APPARENTE, et un disque raw creux la donne
entière. « du -sB1 » compte les blocs. Enfin l'écran de déploiement dit ce qui
l'attend : « Quitter (q) pour lancer l'installation d'ERPLibre » — on
attendait devant une fenêtre terminée sans le savoir.
--- EN ---
"s" still opened the homonymous local VM: the progress view only had the VM's
NAME, and the ~/.ssh/config entry does not exist yet at that point. The
deployment now hands it "ssh -J <host> user@<ip>", which depends on nothing.
Two neighbours of the same defect, never reported but just as serious: the
console opened "virsh console <name>" — the LOCAL VM's — and pause suspended
the local one. Both now go through the VMID on the host.
The Disk column claimed "6.0G/6.0G" on a VM that had written 1.2 GB: "du -sb"
returns the APPARENT size, and a sparse raw disk gives it in full. "du -sB1"
counts blocks. Finally the deployment screen says what awaits it: "Quit (q)
to start the ERPLibre install" — one waited before a finished window without
knowing.
Assisted-by: Claude Opus 5
Rapporté, et c'est le plus grave de la série. Une VM déployée sur Proxmox
sous le nom « erplibre-ubuntu-2604 » — nom déjà porté par un domaine LOCAL —
a vu son installation d'ERPLibre + Odoo partir sur la VM locale. Deux causes
enchaînées : l'entrée ~/.ssh/config volait l'alias de la locale, et le
lanceur détaché ré-résout l'adresse par virsh à chaque tour, qui a répondu
avec le domaine homonyme. Le journal l'écrivait — « → 192.168.123.118 » —
sans que rien n'alerte.
Une VM distante n'est plus ré-résolue : son alias est la seule vérité,
puisqu'il porte le rebond. Et son alias suit la convention des VM imbriquées,
« hôte+vm », le nom court n'étant ajouté que s'il est libre — l'écran le dit.
S'y ajoute le sommaire final qui manquait, à l'image de QEMU/KVM : ce qui
existe, son adresse, sa commande ssh, son journal.
--- EN ---
Reported, and the worst of the series. A VM deployed on Proxmox under the
name "erplibre-ubuntu-2604" — a name already held by a LOCAL domain — had its
ERPLibre + Odoo install land on the local VM. Two chained causes: the
~/.ssh/config entry stole the local one's alias, and the detached launcher
re-resolves the address through virsh on every pass, which answered with the
homonymous domain. The log said so — "→ 192.168.123.118" — with nothing to
raise an alarm.
A remote VM is no longer re-resolved: its alias is the only truth, since it
carries the jump. And its alias follows the nested-VM convention, "host+vm",
the short name being added only when free — the screen says so. Plus the
final summary that was missing, mirroring QEMU/KVM: what exists, its address,
its ssh command, its log.
Assisted-by: Claude Opus 5
Les trois manques restants, demandés. Les colonnes vivantes du suivi — écrit
par seconde, RAM, disque — venaient de virsh, qui ne connaît pas les VM d'un
hôte Proxmox : elles restaient vides, et le relevé d'état les déclarait même
« effacées », ce qui éteignait le reste. L'hôte sait tout cela en un appel
(« pvesh get /cluster/resources »), et le relevé prend la forme de celui de
virsh pour que rien en aval ne distingue la source. Un appel par hôte, mis en
cache cinq secondes : une poignée de main ssh coûte 1 s, virsh 0,03 s.
« Lister les VM » propose maintenant d'en changer l'état, comme QEMU/KVM —
éteindre proprement avant de couper le courant, l'ordre le dit. Et le menu
accepte les commandes ajoutées par todo.json. Enfin « models » manquait aux
traductions : le rapport de migration sortait « 812 models » en français.
--- EN ---
The three remaining gaps, as asked. The dashboard's live columns — written per
second, RAM, disk — came from virsh, which knows nothing of a Proxmox host's
VMs: they stayed empty, and the state probe even declared them "gone", which
switched off the rest. The host knows all of it in one call ("pvesh get
/cluster/resources"), and the reading takes virsh's own shape so nothing
downstream tells the sources apart. One call per host, cached five seconds: an
ssh handshake costs 1 s, virsh 0.03 s.
"List VMs" now offers to change their state, like QEMU/KVM — clean shutdown
before pulling the plug, the order says so. And the menu accepts the commands
todo.json adds. Finally "models" was missing from the translations: the
migration report printed "812 models" in French.
Assisted-by: Claude Opus 5
Décocher l'installation d'ERPLibre faisait disparaître le tableau de bord. La
case « suivi » vivait DANS le groupe de l'installation, build_spec ne la
recopiait même pas dans la spec, et l'épilogue était gardé par « if install or
desktop » : sans rien à installer, il ne se passait rien.
Le suivi devient un choix du DÉPLOIEMENT. Et sans rien à installer, la
commande distante ne vaut plus « true » — journal vide, ✅ instantané : elle
regarde la VM ARRIVER, attend cloud-init, puis relève système, noyau, adresse,
disque et mémoire. Le journal cesse aussi d'annoncer une installation ERPLibre
qui n'a pas lieu.
--- EN ---
Unchecking the ERPLibre install made the dashboard vanish. The "monitoring"
checkbox lived INSIDE the install group, build_spec did not even copy it into
the spec, and the deploy epilogue was gated by "if install or desktop": with
nothing to install, nothing happened.
Monitoring is now a DEPLOYMENT-level choice. And with nothing to install, the
remote command is no longer "true" — empty log, instant ✅: it watches the VM
ARRIVE, waits for cloud-init, then reports system, kernel, address, disk and
memory. The log also stops announcing an ERPLibre install that never happens.
Assisted-by: Claude Opus 5
Le tableau disait la durée et la taille du disque. Il ne disait pas si une VM
TRAVAILLAIT : une installation figée et une qui compile s'y ressemblaient.
Trois chiffres par VM, d'un seul appel « virsh domstats » pour tout le parc
(0,03 s) : ce qu'elle écrit, sa RAM occupée/totale, son disque occupé/total.
Le débit est une moyenne sur DIX secondes — le disque d'une installation
travaille par rafales, et l'instantané n'y montrait que des 0 et des pics.
Le ballon mémoire est réarmé au tour lent : sans période de collecte, libvirt
rend le dernier rapport du pilote, vieux d'une demi-heure. Et cinq caractères
récupérés sur trois colonnes trop larges font tenir la ligne en 150 colonnes.
--- EN ---
The table showed elapsed time and disk size. It did not show whether a VM was
WORKING: a stalled install and a compiling one looked alike.
Three numbers per VM, from a single "virsh domstats" call for the whole fleet
(0.03 s): bytes written, RAM used/total, disk used/total. The write figure is
a TEN-second average — an install's disk works in bursts, and the snapshot
showed only zeros and spikes.
The memory balloon is re-armed on the slow tick: with no collection period,
libvirt hands back the driver's last report, half an hour old. And five
characters reclaimed from three oversized columns keep the row inside 150.
Assisted-by: Claude Opus 5
Le sablier ne distingue pas une installation qui travaille d'une qui est morte :
le marqueur de sortie manque dans les deux cas. Une session ssh emportée, et le
tableau de bord a affiché « ⏳ » pendant 54 minutes sans que rien ne cloche à
l'œil.
La colonne d'état porte maintenant le silence du journal — « ⏳ silence 48min ».
C'est un chiffre, pas un verdict. Le seuil de dix minutes vient d'une mesure :
le téléchargement d'Android Studio tient ~5 min sans une ligne, et l'étape
« APK debug » davantage, son détail partant dans le journal de la VM. Plus bas,
chaque installation deviendrait une alerte, et l'alerte cesserait d'être lue.
--- EN ---
The hourglass does not tell a working install from a dead one: the exit marker
is missing in both cases. An ssh session was reaped, and the dashboard showed
"⏳" for 54 minutes with nothing looking wrong.
The state column now carries the log's silence — "⏳ silent 48min". It is a
figure, not a verdict. The ten-minute threshold comes from a measurement: the
Android Studio download holds ~5 min without a line, and the "debug APK" step
longer, its detail going to the VM's own log. Any lower and every install would
become an alert, and the alert would stop being read.
Assisted-by: Claude Opus 5
La barre montrait le CPU et le disque, jamais la mémoire. C'est pourtant la
seule des trois dont l'épuisement ne se voit nulle part ailleurs : une
compilation mobile s'est fait tuer par le noyau sur une VM de 12 Go sans swap
pendant que la barre affichait une charge tranquille et du disque de reste.
Lue dans /proc/meminfo, sans dépendance — ce suivi tourne sur l'hyperviseur,
donc sous Linux, d'où viennent déjà getloadavg et libvirt. « MemAvailable »
plutôt que « MemFree », presque nul dès que le cache travaille. Le swap
n'occupe la barre que s'il existe, et alors même à zéro : une machine qui
commence à échanger explique une lenteur. Au passage, « charge » et « libre »
s'affichaient en français dans une session anglaise.
--- EN ---
The bar showed CPU and disk, never memory. Yet memory is the one of the three
whose exhaustion shows up nowhere else: a mobile build was killed by the kernel
on a 12 GB VM with no swap while the bar displayed a quiet load and disk to
spare.
Read from /proc/meminfo, no dependency — this monitor runs on the hypervisor,
so on Linux, where getloadavg and libvirt already come from. "MemAvailable"
rather than "MemFree", near zero as soon as the cache is working. Swap takes
room in the bar only when it exists, and then even at zero: a machine starting
to swap explains a slowdown. Along the way, "load" and "free" were showing in
French in an English session.
Assisted-by: Claude Opus 5
Le volet « d » et le compteur du tableau de bord cherchaient la sous-chaîne
« error ». Le journal de l'installation qui vient d'échouer — APK tué par le
noyau sur erplibre-ubuntu-2604-gnome — n'en contient AUCUNE : 0 ligne sur
8765, mesuré. Le volet annonçait « aucune erreur détectée » sur une machine
morte, et le tableau de bord 0 erreur.
Comptent désormais les marqueurs qui ne disent jamais « error » : « ⚠ ÉCHEC »,
« FAILURE », une trace Python, un « fatal: » de git, une mort par mémoire. Le
volet ouvre sur un résumé — l'étape en échec et son diagnostic, les signaux
durs, puis les répétitions comptées par forme. Sur ce journal : 1 étape nommée,
2 signaux, là où il n'affichait rien.
--- EN ---
The "d" pane and the dashboard counter looked for the "error" substring. The
log of the install that just failed — APK killed by the kernel on
erplibre-ubuntu-2604-gnome — contains NONE: 0 lines out of 8765, measured. The
pane said "no error detected" about a dead machine, and the dashboard 0 errors.
Markers that never say "error" now count: "⚠ ÉCHEC", "FAILURE", a Python
traceback, a git "fatal:", a death by memory. The pane opens on a summary — the
failed step with its diagnostic, the hard signals, then repeats counted by
shape. On that log: 1 named step and 2 signals, where it showed nothing.
Assisted-by: Claude Opus 5
The architecture was missing, though it alone explains why an install
takes ten times longer: s390x and arm64 are EMULATED on an amd64 host.
Without it you look for the fault elsewhere. An "Arch" column, right after
the name.
The table was pinned at 74 columns: beyond that nothing was reachable and
no scrollbar showed, for want of reserved space. Automatic width capped at
60% of the screen, scrolling both ways, and VISIBLE bars rather than
guessable ones.
Columns can finally be adjusted: "+" widens the cursor's, "-" narrows it,
"0" returns them all to their original width -- the same keys as the mail
TUI. auto_width is turned off along the way, otherwise the setting did not
survive the table's first update, and the resize aimed at the "#" column
whatever the cursor was on.
--- FR ---
L'architecture manquait, alors qu'elle explique à elle seule pourquoi une
installation dure dix fois plus : s390x et arm64 sont ÉMULÉES sur un hôte
amd64. Sans elle, on cherche la faute ailleurs. Une colonne « Arch »,
juste après le nom.
Le tableau était figé à 74 colonnes : au-delà rien n'était atteignable et
aucune barre de défilement n'apparaissait, faute de place réservée.
Largeur automatique plafonnée à 60 % de l'écran, défilement dans les deux
sens, et barres VISIBLES plutôt que devinables.
Les colonnes s'ajustent enfin : « + » élargit celle du curseur, « - » la
rétrécit, « 0 » les ramène toutes à leur largeur d'origine — les mêmes
touches que la TUI courriel. auto_width est désactivé au passage, sans
quoi le réglage ne survivait pas à la première mise à jour du tableau, et
le redimensionnement visait la colonne « # » quel que soit le curseur.
Assisted-by: Claude Opus 5
The monitor could only watch. The "a" key now opens the selected VM's
actions: update, restart Odoo, delete.
The update is chosen by parts, and the order is not free: system packages,
then git repositories, then Python dependencies -- the latter compile
against the former. Nothing is chained with "&&": one failing part must
not take the others down. Deleting demands a second hand, a confirmation
screen naming the disk to be erased.
Two defects the dashboard revealed. Detached installs were stealing the
shell's keystrokes, the child inheriting a terminal it had no business
reading. And the elapsed time restarted from zero every time the dashboard
was reopened, since it was counted from the view rather than from the
run's own first write.
--- FR ---
Le suivi ne savait que regarder. La touche « a » ouvre désormais les
actions de la VM sélectionnée : mettre à jour, redémarrer Odoo, supprimer.
La mise à jour se choisit par parties, et l'ordre n'est pas libre :
paquets système, puis dépôts git, puis dépendances Python — ces dernières
compilent contre les premiers. Rien n'est enchaîné par « && » : une partie
en échec ne doit pas emporter les autres. Supprimer exige une seconde
main, un écran de confirmation nommant le disque à effacer.
Deux défauts que le tableau de bord a révélés. Les installations détachées
volaient les frappes du shell, l'enfant héritant d'un terminal qu'il
n'avait pas à lire. Et la durée repartait de zéro à chaque réouverture,
comptée depuis la vue plutôt que depuis la première écriture du run.
Assisted-by: Claude Opus 5
VM type was global: the whole fleet as servers, or all of them GNOME. It
now lives on the VM, all the way down -- the creation flag, and one remote
command per machine at install time, where a single one served them all.
The right pane is no longer a table but a row of widgets per VM: vCPU,
RAM, disk and type, each with its usual values and a free entry. The
left-hand fields become the shared default, which the screen now says. The
scope selector and the F2 modal go away: given two ways to do the same
thing, keep the visible one.
Three traps came out of it, and the tests lock them. Widget ids carry a
RANK, and the rank shifts when an entry is ticked, so an event from an
already destroyed widget applied to the VM that took its place -- rows now
carry a generation, marked BEFORE mounting, since mount_all empties the
pending children and marking after it was a race. The x1..x4 profile no
longer reached any VM. And a total of zero never said it had counted
nothing.
--- FR ---
Le type de VM était global : tout le parc en serveur, ou tout en GNOME. Il
vit désormais sur la VM, jusqu'au bout — le drapeau de création, et une
commande distante par machine à l'installation, là où une seule les
servait toutes.
Le panneau de droite n'est plus un tableau mais une rangée de widgets par
VM : vCPU, RAM, disque et type, chacun avec ses valeurs usuelles et une
saisie libre. Les champs de gauche deviennent le défaut commun, ce que
l'écran dit maintenant. Le sélecteur de portée et la modale F2 partent :
entre deux façons de faire la même chose, on garde la visible.
Trois pièges en sont sortis, et les tests les verrouillent. Les
identifiants de widgets portent un RANG, et le rang se décale quand on
coche une entrée : un événement émis par un widget déjà détruit
s'appliquait à la VM qui avait pris sa place — les rangées portent
maintenant une génération, marquée AVANT le montage, car mount_all vide
les enfants en attente et marquer après était une course. Le profil x1..x4
n'atteignait plus aucune VM. Et un total à zéro ne disait pas qu'il
n'avait rien compté.
Assisted-by: Claude Opus 5
Monitoring stayed on the same address for 1170 s, never catching up with
the VM. Two holes, in the very re-resolution meant to prevent that.
The lease fallback required an answer on port 22. But dnsmasq keeps one
lease per MAC: when cloud-init sets the real hostname and the DHCP client
asks again, the lease MOVES the address. The old one no longer belongs to
the VM and sshd will never answer there. The lease therefore wins as soon
as it stops listing the current address.
The other hole explains the silence: virsh was muted on both branches, so
an unreachable libvirt kept the initial IP without a single line saying
so. The log went quiet for a quarter of an hour for the same reason --
cloud-init holds the package lock while writing nothing.
--- FR ---
Le suivi restait sur la même adresse pendant 1170 s, sans jamais rattraper
la VM. Deux trous, dans la re-résolution censée l'éviter.
Le repli par bail exigeait une réponse sur le port 22. Or dnsmasq garde un
bail par MAC : quand cloud-init pose le vrai nom d'hôte et que le client
DHCP redemande, le bail DÉPLACE l'adresse. L'ancienne n'appartient plus à
la VM et sshd n'y répondra jamais. Le bail l'emporte donc dès qu'il cesse
de lister l'adresse courante.
L'autre trou explique le silence : virsh était muet sur les deux branches,
si bien qu'un libvirt injoignable conservait l'IP initiale sans une ligne
pour le dire. Le log se taisait un quart d'heure pour la même raison —
cloud-init tient le verrou des paquets sans rien écrire.
Assisted-by: Claude Opus 5
Installs run detached (setsid -f): closing the terminal does not stop
them, but it lost the only view onto them. With no way to resume, the
only way out was deleting the VMs and starting over.
Picking "Deploy" now looks at the latest run: if VMs there still lack an
exit marker, it offers to reopen its monitoring instead of starting
another. Time since the last write is shown, a dead run being otherwise
indistinguishable from a live one.
Only the latest run is examined: an old one left without a marker would
flag a phantom install forever. Dry-run creates nothing, so it does not
ask. "Reopen monitoring" moves to the Deployment section, where one
looks for it.
--- FR ---
Les installs partent détachées (setsid -f) : fermer le terminal ne les
arrête pas, mais faisait perdre la seule vue dessus. Sans moyen de
reprendre, la seule issue était d'effacer les VM et de recommencer.
Choisir « Déployer » regarde donc le dernier run : s'il lui reste des VM
sans marqueur de sortie, il propose de rouvrir son suivi plutôt que d'en
lancer un autre. Le silence depuis la dernière écriture est affiché, un
run mort n'étant pas distinguable autrement d'un run vivant.
Seul le dernier run est examiné : un run ancien laissé sans marqueur
signalerait éternellement une install fantôme. L'aperçu ne crée rien, il
ne pose pas la question. « Rouvrir le suivi » rejoint la section
Déploiement, où on le cherche.
Assisted-by: Claude Opus 5
The monitor froze the address given at launch, so it lost the VM as soon as
cloud-init renamed the host and DHCP handed out another lease. It now
re-resolves at each attempt, in the views too, and reads virsh without sudo
— « sudo -n » fails in a detached session with no tty.
First boot also stopped paying for what it does not need: the guest agent
leaves cloud-init, snapd and locale-gen go, apt takes the fastest mirror.
The timezone follows the host. And when KVM is missing, the deployment says
so before the wait instead of being mysteriously fifteen times slower.
--- FR ---
Le suivi figeait l'adresse connue au lancement : il perdait donc la VM dès
que cloud-init posait le vrai nom d'hôte et que DHCP donnait un autre bail.
Il la ré-résout désormais à chaque tentative, dans les vues aussi, et lit
virsh sans sudo — « sudo -n » échoue dans une session détachée, sans tty.
Le premier démarrage cesse aussi de payer l'inutile : l'agent invité sort
de cloud-init, snapd et locale-gen disparaissent, apt prend le miroir le
plus rapide. Le fuseau suit l'hôte. Et faute de KVM, le déploiement le dit
avant l'attente, au lieu d'être quinze fois plus lent sans raison visible.
Assisted-by: Claude Opus 5
One Textual dashboard per install run: live status, logs, host telemetry,
error counts and history. The installs run detached, so closing it does not
stop them.
--- FR ---
Un tableau de bord Textual par exécution : état en direct, logs, télémétrie
de l'hôte, comptes d'erreurs et historique. Les installations tournent
détachées : le fermer ne les arrête pas.
Assisted-by: Claude Opus 4.8