Premier fichier de script/analyse/, le paquet d'outils qui répondent à « qu'y
a-t-il dans cette base, et qu'est-ce qui cassera à la montée de version ». Ce
commit ne livre que ce que le premier outil consommera, pas une bibliothèque
spéculative.
psql en sous-processus, pas psycopg2. La raison est déjà écrite dans
reset_stale_cow_views.py : ces outils tournent sur des bases dont le registre
Odoo ne charge PAS — une base 12.0 sur un checkout 18.0 — et c'est précisément
le moment où on veut les analyser. Accessoirement psycopg2 n'est pas dans
.venv.erplibre, qui est l'interpréteur de ces outils.
La lecture seule est une garantie du serveur, pas une promesse du code.
PGOPTIONS porte default_transaction_read_only=on, donc PostgreSQL refuse toute
écriture sur la connexion. Un SET dans le même -c n'aurait rien donné : psql -c
ouvre une transaction implicite unique, et default_transaction_read_only ne vaut
que pour les transactions suivantes. PGOPTIONS borne aussi statement_timeout —
un scan parti en vrille ne doit pas figer un menu.
Le résultat voyage en JSON, pas en champs séparés. Une arch de vue contient des
sauts de ligne et des « | » : tout séparateur maison finit par couper au mauvais
endroit. json_agg côté PostgreSQL, json.loads côté Python, et le problème
n'existe plus. C'est le même choix que snapshot_cow_views.py.
Les sondes lisent pg_attribute, pas information_schema, qui est filtré par les
droits : avec un rôle non propriétaire elle rendrait un ensemble vide, et
l'analyse concluerait « aucune colonne website, donc aucune vue COW » sans le
moindre avertissement. Faux, et silencieux.
tr_col décide sur le TYPE réel de la colonne, pas sur un numéro de version. Un
champ traduit est du jsonb à partir de 16.0 et du texte avant ; une base à
moitié migrée porte les deux, et un numéro de version mentirait. Une colonne
inconnue rend NULL::text plutôt que du SQL invalide, ce qui permet de sonder et
d'interroger d'un trait.
MODEL_TABLE_OVERRIDE est dérivée des sources, pas écrite de mémoire : parcours
AST des 27 843 .py de odoo18.0/ et addons/, en ne gardant que les classes dont
le _table diffère du défaut. Onze entrées. Beaucoup de modules déclarent un
_table égal au défaut — les compter ferait croire à 22 surcharges là où il y en
a 11. Un modèle absent de la table et dont la table est introuvable est classé
« table inconnue », un fait, jamais « table orpheline », une anomalie.
Le nom de base est validé au lieu d'être échappé : il finit dans une commande
lancée avec shell=True, et l'alphabet qu'une base Odoo utilise vraiment est
court. Ce contrôle attrape aussi « Traceback (most recent call last): », que la
liste des bases pouvait offrir jusqu'au commit précédent.
Pas de __init__.py : script/todo/, script/database/ et script/odoo/migration/
n'en ont pas et sont importés tous les jours. Paquets d'espace de noms.
Vérifié. 31 tests sans base — fonctions pures, construction de SQL,
« False » de config.conf traité comme non défini, nom de base hostile refusé.
Puis à la main sur une base synthétique dont le SQL est dans le docstring du
test et a été rejoué tel quel : jsonb et text distingués, le « | » et le saut de
ligne préservés à travers json_query, ir.actions.act_window résolu en
ir_act_window, une base sans ir_module_module refusée, et un CREATE TABLE
refusé par le serveur. La suite passe de 262 à 293 tests, sans nouvel échec.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Textual était déclaré sans version. Les quatre écrans TUI du dépôt sont écrits
pour la 8, et la bibliothèque casse son API entre majeures : un « pip install -U »
les casse tous sans prévenir.
Borner le fichier de requirements ne suffisait pas. install_command() faisait
« pip install textual », sans borne : « make install » aurait pris la 8, et
l'installation proposée à l'écran la majeure suivante. Deux chemins pour la même
dépendance, qui ne disent pas la même chose. La borne vit donc dans une seule
constante, TEXTUAL_SPEC, que install_command() utilise, et que le fichier de
requirements recopie avec un commentaire qui pointe dessus.
La borne ne s'applique qu'à l'installation : ensure() vérifie « est-ce
importable », pas « à quelle version ». Un Textual 9 déjà présent passe, et c'est
volontaire — refuser de démarrer sur une version qui marche peut-être serait pire
que le problème.
lxml devient une dépendance déclarée. Il n'arrivait que par pykeepass,
openupgradelib et odoo-module-migrator ; le jour où l'un d'eux s'en passe, il
disparaît d'un venv sans que rien ne le réclame. Sans borne : cyclonedx-python-lib
demande déjà « lxml >=4,<7 », en ajouter une seconde n'apporterait qu'un conflit
possible.
Vérifié : install_command() porte la borne, l'insertion de « --user » hors venv
reste au bon rang, et le Textual installé (8.2.8) la satisfait.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
« db --list » rendait un code de retour que personne ne lisait. Comme la sortie
et l'erreur passent par le même flux (stderr=STDOUT, execute.py), les lignes de
la trace d'appel devenaient les entrées du menu :
[1] Traceback (most recent call last):
[2] File "…/odoo-bin", line 8, in <module>
[0] Retour
Choisir [1] renvoyait « Traceback (most recent call last): » comme nom de base à
l'appelant, qui le passait à sa commande — sauvegarde, restauration, shell. Le
code de retour était pourtant là, 120, dans la variable « status », écrasée deux
lignes plus bas par la réponse de click.prompt.
Le contrôle se fait donc AVANT de construire le menu : code non nul, on affiche
les dernières lignes de l'erreur réelle et on rend False. Une liste vide dit
aussi ce qu'elle est, au lieu d'un écran ne portant que « Retour ».
Au passage, la branche de saisie invalide appelait t("cmd_not_found") — une clé
absente de TRANSLATIONS, donc affichée telle quelle, en anglais technique. Les
autres menus disent t("Command not found !") ; celui-ci le dit maintenant aussi.
« quiet=True » s'ajoute parce que la sortie brute n'a plus de raison d'être vue :
le menu affiche la liste, et en cas d'échec c'est l'erreur qui est montrée.
Vérifié sur les quatre chemins, PGPORT=1 pour simuler la panne : injoignable
rend False sans afficher un seul menu ; [0] rend False ; [1] rend « test » ; une
saisie invalide affiche un message traduit dans les deux langues.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Un littéral de dict garde la DERNIÈRE valeur : une clé déclarée deux fois écrase
silencieusement la première. Cinq clés étaient dans ce cas — « SSH host is
required! », « Total: », « max », « Cancelled. », « failed » — et rien ne le
signalait, ni au lancement, ni en revue, ni dans la suite de tests.
Quatre étaient d'inoffensifs copier-coller : les deux déclarations portaient la
même paire fr/en. La cinquième ne l'était pas. « failed » valait « en échec » au
premier endroit et « échouées » au second ; c'est « échouées » qui gagnait, pour
les TROIS appels, y compris le bilan de déploiement écrit pour « en échec ». Une
traduction avait donc été rédigée, relue, et n'a jamais été affichée une fois.
C'est l'occurrence masquée qui part, jamais celle qui gagne : les cinq clés
rendent exactement ce qu'elles rendaient avant, la sortie ne bouge pas d'un
caractère. Choisir un meilleur français pour « failed » est un autre changement,
à faire en le voyant à l'écran.
Vérifié : 865 clés, 865 uniques, aucun doublon ; les cinq clés rendent la même
valeur qu'avant. La suite de tests garde ses 262 tests et ses échecs
préexistants, à l'identique — dont trois qui relèvent du même angle mort :
test_todo_i18n teste la clé « menu_quit », absente de TRANSLATIONS.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
« Créer la clé SSH si absente et la déployer ? » was asked before anything had
been attempted — so before knowing whether a key was needed at all. On a fleet
already reachable it was pure noise, and answering no left every later failure
reported as a flat « injoignable ».
The question now appears where the problem does:
🔒 hote: SSH refused the identity.
Permission denied (publickey).
Aucune clé SSH dans ~/.ssh.
En créer une et la déployer ? (O/n)
and the host is probed again straight after, so the walk carries on into its
guests instead of stopping.
That required telling a refused identity from an unreachable host, which the
probe could not do: both returned None. It now reports which — auth or net —
by matching ssh's own wording (permission denied, too many authentication
failures, no such identity, host key verification failed). An unreachable host
never triggers the key question, because a key would not help it, and the real
ssh message is printed either way rather than a generic label.
A key created mid-walk is picked up by the entries written afterwards, so
their IdentityFile names it.
Verified against ssh's actual messages: the eight classified correctly,
including « Identity file not accessible » which is a warning about a missing
file, not a refusal. Then the flow: a fleet that answers straight away is
never asked about keys at all, a refusal asks and — once accepted — creates,
deploys, re-probes and configures the guest, declining reports « accès
refusé » without deploying, and an unreachable host says so without mentioning
identity.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
« Installez textual pour le TUI (pip) » left the user to work out which pip,
which interpreter, which package — and it was printed from eight different
places. Every TUI screen now asks:
⚠ Textual est nécessaire pour cet écran.
L'installer maintenant ? (O/n, défaut : oui)
The command targets sys.executable, the interpreter that will have to import
it — installing a distribution package would land somewhere the venv never
looks. Outside a venv it adds --user, which is also what gets past the refusal
of distributions whose environment is externally managed (PEP 668).
Two details that decide whether this works at all:
· importlib.invalidate_caches() after installing. A failed import is
remembered, so without it textual stays « missing » for the rest of the
session despite having just been installed.
· a pip that exits non-zero never reports success. The check is « is it
importable NOW », not « did pip return 0 », and the failure suggests the
distribution package by name.
It lives in its own module rather than as a TODO method: todo_upgrade needs it
too and is imported BY todo, so putting it there would close a cycle.
One call site is deliberately NOT converted. The statistics screen only reads
files; it never touches Textual, and its old message claimed otherwise. An
import failure there is a real module problem and now says so.
Verified: already-present asks nothing and runs nothing; refusal installs
nothing; pip failing returns False and points at python3-textual; pip
succeeding returns True; prompt=False reports without asking. Then each of the
four TUI entries — telemetry, deploy form, deploy progress, migration resume —
offers and falls back cleanly on refusal, while the statistics screen stays
silent.
Also caught by those tests: « import importlib » alone does not expose
importlib.util, so availability could not be checked at all.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adopting a host from ~/.ssh/config only to write « User erplibre » under its
guests defeats the point: those hosts are not necessarily ERPLibre VMs, and a
guest follows its parent's convention because the parent created it.
The declared User is now read and propagated — to the guest entries, to the
summary line, and to the qemu+ssh URI handed to virt-manager. Only when
nothing is declared does QEMU_VM_USER apply, the cloud-init account of VMs
deployed from here, now a named constant instead of a literal repeated at
each call site.
Resolution follows OpenSSH: the FIRST User among the matching blocks wins,
wildcards included. That is the opposite of what one expects, so it was
checked against ssh -G rather than the manual: with « Host * / User global »
placed first, ssh really does report global even for a host that declares its
own — which is why the manual tells you to put « Host * » last. The
implementation matches on both layouts.
« user@host » is also accepted when typing a raw address, and otherwise the
account is asked with QEMU_VM_USER as default.
ssh-copy-id needed no change: it is given the alias, whose block now carries
the right User.
Verified on a config with a per-host User, a host without one and a trailing
« Host * »: the three resolutions correct, the guest written with mathben and
not erplibre, and virt-manager offered ('hyperviseur', 'mathben').
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The two entries — « local VMs » and « nested, recursive » — wrote the same
thing and differed only by how far they walked. That is a question, not a
menu, so there is now a single flow and the depth is asked:
Profondeur (1 = ces machines seulement, défaut : 2)
Depth counts LEVELS of machines including the roots, so 1 writes the roots and
probes nothing. The old code always probed, which is why « local VMs only »
needed to be a separate entry at all.
The first question is where to look, because the machine to configure is not
always a VM of this host:
[1] Local QEMU VMs (virsh) the previous behaviour
[2] Hosts from ~/.ssh/config pick among what is already known; each is
probed over SSH and only those actually
running QEMU get their guests configured
[3] Type a host or an IP a raw IP gets an alias, without which
neither the children's ProxyJump nor
virt-manager would have a name to use
A root taken from ~/.ssh/config carries no IP: its address is already in the
file, so nothing is rewritten — the walk simply starts from it. That is the
single change that let the two flows merge.
Verified on a config holding a wildcard, a bastion and a hypervisor: local VMs
reaching their nested guest, ssh-config hosts skipping the bastion (no QEMU)
while configuring the hypervisor's guest and offering only the hypervisor to
virt-manager, a raw IP producing qemu-10-9-9-9 and its child, depth 1 probing
nothing, and [0] asking nothing further.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two dead ends the menu let you walk into.
Without libvirt, every entry answered « sudo: virsh: command not found » and
carried on as if nothing had happened — « Aucune VM trouvée » reads like an
empty host, not a missing package. Entering the QEMU menu now checks for virsh
and offers to install it. The packages are not guessed here: deploy_qemu.py
--setup-host already knows them per distribution, so the offer just runs it.
Refusing keeps the menu open — listing images or previewing a deployment needs
no libvirt. If virsh is still absent afterwards, that is said too: on a
rolling-release kernel the setup can need a reboot before the modules load.
Without an SSH key, « Chemin de la clé publique SSH (aucune): » accepted an
empty answer and deployed VMs nobody could log into — cloud-init injects no
key, so there is no install, no check, nothing but a console. The prompt now
says what the consequence is and offers to generate one, reusing the same
_ssh_ensure_key as the SSH configuration tool so the whole fleet shares one
key.
Verified: silence when virsh is present, the sudo --setup-host command issued
when accepted and nothing run when refused; and on an empty ~/.ssh, the key
generated, both halves on disk, and its public path carried into the spec.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A stopped service on the far side showed up only as a wall of « channel N:
open failed: connect failed: Connection refused », one line per browser
request, saying nothing about WHICH end refused. The tunnel itself was fine;
there was simply nothing to reach. Diagnosing it took three commands.
The port is now probed first, and the answer is plain:
⚠ Rien n'écoute sur le port 8069 de test-vm_02+erplibre-ubuntu-2404
Démarrer le service là-bas, ou continuer quand même.
Continuer quand même ? (o/N)
The probe opens a real TCP connection to « localhost:<port> » from the remote
host rather than reading its listening table. That is exactly what the tunnel
will do — same host resolution, same IPv4/IPv6 choice — so it cannot say open
where the tunnel would fail. It also needs no ss or netstat, which minimal
images lack.
Three outcomes, three behaviours: listening goes straight through, closed
warns and asks (default no), and an inconclusive probe — unreachable host, no
bash — says so and continues rather than blocking on its own uncertainty.
Verified against the real VM: port 22 open, port 9999 closed, an unknown host
inconclusive, and port 8069 correctly reported closed after the Odoo service
had been stopped — the very case that prompted this. Then at flow level:
refusing aborts without opening anything, forcing opens the tunnel anyway, and
an inconclusive probe still opens it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A failed update_addons_all left one option, « [1] to redo the command » —
useless against the failure that actually causes it, since redoing an upgrade
whose COW copy has drifted fails again identically. The prompt now offers to
run reset_stale_cow_views on the database concerned:
[1] to redo the command
[2] Check the COW views that drifted (technolibre_…_upgrade_15)
Choosing [2] runs the checker and comes back to the same prompt, so the usual
sequence — look, repair, redo — happens without leaving the migration.
The database is read from the failed command itself, since the executor is not
told which one it is. Three forms are recognised: « -d <db> », « --database
<db> », and the positional argument of the addons scripts. No database found,
no [2] offered — « make format » gets the old prompt unchanged.
The option only ever REPORTS. Resetting a copy can erase a real customisation,
so it stays a separate deliberate command, printed with the exact line to run
and a reminder to read the diff first.
Verified: the six command forms parsed correctly (and two that hold no
database at all), [2] invoking the tool then re-offering the choice, no [2] on
a command without a database, and the whole chain against the real
_upgrade_15, where it prints the two drifted copies followed by the reset
line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reaching a VM's Odoo from the workstation browser meant remembering the
-L syntax and which side « localhost » refers to. Entry [3] of the Deploy
menu asks for a host, a remote port (8069 by default) and a local one, then
holds the tunnel open.
Nothing has to be said about jumps: the ProxyJump already in ~/.ssh/config
applies on its own, which is what makes a NESTED VM reachable — its address
means nothing from here, only from its parent.
Two guards, both from getting it wrong by hand:
· a local port already in use is reported before ssh fails on it;
· a local port that differs from the remote one gets a warning, because
Odoo redirects using web.base.url and would send the browser to an
address that does not exist locally. With matching ports and the usual
web.base.url = http://localhost:8069, there is nothing to adjust.
The host list is read from ~/.ssh/config by a small shared helper: it expands
a Host line carrying several names and drops the wildcard patterns, which are
rules rather than machines.
Verified against a config holding « Host * », a plain host and a two-name
line: the three real names listed in order, selection by number and by name,
8069/8069 by default, 9072:localhost:8072 warning about web.base.url,
matching ports staying silent, an empty host cancelling without running
anything, and the busy-port probe answering correctly on a socket bound then
released.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Third COW incident of this migration, and the first two tools did not cover
it. A copy freezes the module view it came from; version after version the
MODULE view is modernised and the copy is not, until a child xpath lands on an
anchor the copy never had:
Element '<xpath expr="//head/script[@id='web.layout.odooscript']">'
cannot be located in parent view
On 14.0 -> 15.0 that was web.layout, copied in the 12.0 era: no id on the
script tag, QWeb still saying t-raw. It had survived only because the 14.0
module xpath carried a fallback — //head/script[@id='…'] | //head/script
[last()] — that 15.0 removed. The breakage was years old; 15.0 merely stopped
hiding it.
The test is DIFFERENTIAL, and it has to be. Resolving a child's xpath against
its parent's own arch proves nothing: Odoo resolves against the COMBINED arch
of the whole chain, so a child of website.layout legitimately targets //header
coming from an ancestor. Measured on the real database, that naive rule
reported 5 copies where only 1 had a problem. Resolving twice — against the
module twin and against the copy — and keeping only what the twin satisfies
and the copy cannot, isolates drift and nothing else.
Two accuracy fixes the real data forced:
· inactive children are skipped; Odoo never applies them
· a missing lxml is now a LOUD failure. The first version fell back to
« everything resolves », so the checker answered « all clean » while
checking nothing — the worst possible outcome for a checker. Verified: the
bare system python3 has no lxml, .venv.erplibre does.
--reset copies the module arch over the copy, after saving the previous arch
to a timestamped file and printing the diff, because a copy can hold a real
customisation buried in the drift — on web.layout it was one <meta viewport>
among six differences.
Timing matters: drift only exists once the module views carry the new version,
so run this AFTER OpenUpgrade and BEFORE update_addons_all. Run on the 14.0
database it finds nothing about web.layout, and rightly so — there the module
view says t-raw too.
Verified end to end on a scratch database rebuilt from the real arches:
detection of the exact reported failure, dry-run changing nothing, --apply
resetting it, the backup holding the previous arch, and a clean re-check.
Then across the four migration databases, where it flags two latent drifts
(website_sale.product, website_blog.blog_post_complete) that no version bump
has surfaced yet.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The interface choice offered two ways to ACT on a migration and none to just
look at one. Entry [3] reads the progression file and the traces left under
private/, writes nothing, and can be opened as often as one likes — [0] there
returns to the interface choice, and a new [0] Cancel leaves the tool entirely
instead of forcing a migration to be started.
« stats » is deliberately not storable as a default: it does nothing, so
landing on it every time would only be in the way.
The dashboard answers what was asked — which modules were removed, and how far
between Odoo versions — plus what the recorded data already knew and nobody
was showing:
· elapsed time since the migration started
· module count per version, with the delta per bump — a jump losing 15
modules does not mean the same thing as one losing none
· modules reported missing or duplicated
· which fix hooks exist and which have run
· COW snapshots taken, with their view count over time
· how many commands ran, and the annotated decisions among them
Then, on demand: the removed modules with their justification, the same list
comma-separated to paste, a diff between any two COW snapshots (through the
existing snapshot_cow_views.py --diff), the decisions, and the last commands.
Computation lives in migration_stats.py as pure functions fed by
resume_context() — the same context the resume screen and its TUI already
render, so the step and version-bump state can never be described two
different ways. read_uninstall is INJECTED rather than reimplemented: it
resolves private-then-global itself, so the screen shows exactly the list that
would be applied.
Verified on a realistic progression: 1 j 03 h elapsed, 312 -> 297 -> 282
modules (-15 each), 5 removals across two bumps with their reasons, 4 COW
snapshots ordered by time, the fix applied for 14.0 and pending for 15.0, and
the progression dict byte-identical afterwards. The menu was checked to map
'' /1/2/3/0/9 to tui/tui/cli/stats/None/tui, and [0] to leave without asking
anything else.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Every session paid ~4 000 tokens for CLAUDE.md and .claude/rules/, and part of
that content no longer described this repository:
CLAUDE.md and 05-environments.md gave .venv.odoo18/bin/python — a path that
does not exist. The real one carries BOTH versions,
.venv.odoo18.0_python3.12.10. 02-project-structure.md announced addons/ and
odoo12.0/, absent from the checkout.
That is the reason for the cuts, more than the token count: the files drifted
from the repository, while « ls » cannot. What a session can rebuild by
reading the code is now left to the code.
02-project-structure.md deleted — a tree that ls gives, and gives right
05-environments.md deleted — a venv list that was simply false
04-code-conventions.md reduced to a pointer at .flake8, .editorconfig and
pyproject.toml, which already hold every value it
repeated, plus the Git conventions, which they
do not
01-versions.md table dropped for the pointer at
conf/supported_version_erplibre.json, whose keys
already pair Odoo with Python
Two blocks move to lazy loading — their body costs nothing until invoked:
03-commands.md -> skill erplibre-commands (kept whole: there is no
« make help » target, and the per-module test recipes
carry flags nobody would guess)
07-documentation.md -> skill erplibre-doc-i18n for the how-to, while the
prohibition « never edit a generated .md » STAYS in
the rules: a rule that must hold at all times cannot
live in a file loaded on demand
Resident guidance: ~3 997 -> ~1 939 est. tokens per session.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A nested VM was written as « Host erplibre-ubuntu-2404
test-vm_02+erplibre-ubuntu-2404 » — two patterns on one Host line. Valid ssh
syntax, but the short name only repeats the tail of the chain and buys
nothing, so it is gone. Nested hosts now carry the chained name alone.
Existing configs repair themselves: replacing a block drops any whose name
list intersects the new one, so the old two-name line is removed as a whole
and rewritten with the single chained name — no stale short entry left behind.
Verified on a config holding exactly the reported line, alongside an unrelated
host that must survive.
The naming rule is simpler too: a chain cannot collide with another machine,
so there is no case left where a short name has to be preferred or avoided.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The message claimed virt-manager rewrites its settings on exit and had to be
restarted. Reading the 5.1.0 source shows the opposite for the part that
matters: connection.py registers listen_perconn on /pretty-name and its
callback emits state-changed, so a running virt-manager picks a renamed
connection up immediately. Only the connection LIST is read at startup, so a
restart is needed to see a newly added one — which is what the message now
says.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A nested QEMU host appeared in virt-manager under its bare name, so nothing
said which machine it lived inside. It now carries the same « parent+child »
chain the ssh aliases use, at two levels of defence:
- the connection URI uses the CHAINED alias, so even a virt-manager that
ignores custom labels shows the nesting in the host part;
- pretty-name is set explicitly, which is the label virt-manager actually
displays.
pretty-name lives in a RELOCATABLE GSettings schema,
org.virt-manager.virt-manager.connection, one path per connection. The path is
built with no escaping at all: virt-manager just deletes every « / » from the
URI and uses the rest verbatim (virtManager/config.py, _make_perconn_key), so
qemu+ssh://erplibre@a+b/system becomes conns/qemu+ssh:erplibre@a+bsystem/.
That was verified for real rather than assumed: virt-manager is not installed
here, so its gschema was fetched from the distribution package, compiled into
a temporary schema dir, and the set/get/reset round trip run against it with a
genuinely chained URI — colons, @ and + all survive the dconf path. The key
was reset afterwards; nothing was left in dconf.
Making the URI carry the chain needs the chained name to resolve, so a nested
host is now written as « Host <short> <parent+child> » — ONE block, two names,
since ssh accepts several patterns on a Host line. The short name stays for
typing and is still only kept when free; the chain is always there. Checked
with ssh -G: both names yield the same HostName and ProxyJump, and a third
level chains as a+b+c. libvirt was checked too — it hands the URI host to ssh
untouched, + included.
Replacing a block by name meant a regex that could not see a Host line
carrying several names, which would have left the same name defined twice —
ssh honours the first, so an update would silently not apply. Removal now
parses the file into blocks and drops any whose name list intersects the new
one. Verified on a config with a global directive, a Match block and a
multi-name Host: only the targeted block goes, rewriting twice is idempotent,
and an unknown name leaves the file byte-identical.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Every reachable machine ended up in virt-manager, not just the ones hosting
VMs. The probe ran « for n in $(sudo virsh list --all --name 2>/dev/null) »:
with no virsh, the command substitution is empty, the loop body never runs and
the snippet exits 0 with no output — indistinguishable from « QEMU is here,
it just has no VM ». Both read as a libvirt host.
The probe now states it outright, « LIBVIRT<TAB>yes|no » as its first line, so
the two cases separate: a machine WITHOUT QEMU is skipped, a machine with QEMU
and no VM is still offered — that is where one would create some.
Verified against the real snippet: this host answers « yes » plus its two VMs,
and with virsh out of PATH it answers « no ». Then on a simulated fleet, only
vm-avec-qemu and vm-qemu-sans-vm are proposed; the plain dev VM and the
unreachable one are not.
Also: the ~/.ssh/config blocks were missing IdentityFile. Every entry now
names the private key it needs — the one cloud-init injected, or the one just
deployed — with IdentitiesOnly yes beside it. Without that flag IdentityFile
ADDS to the agent's identities instead of replacing them, and a slightly full
agent hits « Too many authentication failures » before reaching the right key.
The .pub suffix is stripped: IdentityFile wants the private half.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Writing ~/.ssh/config only helps if the key is accepted at the other end.
Both SSH options now offer to create an ed25519 key when none exists — the
same choice deploy_qemu.ensure_ssh_key makes, so created and adopted VMs share
one key — and to push it with ssh-copy-id.
Hosts that already accept the key are skipped, tested with
PasswordAuthentication=no: without it ssh would fall back to the password and
every host would look like it already had the key. ssh-copy-id runs on the
real terminal rather than through captured output, otherwise its password
prompt would be invisible.
In the recursive walk the key is deployed BEFORE probing each level, not at
the end. The probe uses BatchMode, so an un-keyed machine answers nothing and
the level below it stays invisible — deploying afterwards would find only the
first level.
virt-manager, when installed, gets the machines that actually run libvirt
added to its connection list, so their nested VMs can be driven from the local
GUI. The URI uses the SSH ALIAS rather than the raw IP: qemu+ssh goes through
the ssh binary, hence ~/.ssh/config, so the alias already carries both the
address and the ProxyJump — a bare IP could not reach a nested VM at all.
Connections live in GSettings, not a file. The list is READ first and written
back merged, so nothing already configured is lost, and a failed read (no
schema, no virt-manager) simply means the whole feature stays silent — no
prompt, no write. virt-manager rewrites its settings when it exits, so a
warning says to restart it.
Verified: key generated with 0600 on the private half and reused on the second
call; ssh-copy-id issued only for hosts that need it; the recursive walk
deploying level by level before each probe; GSettings merge keeping existing
URIs and skipping the write when nothing is missing; « @as [] », a populated
list and a missing schema all parsed correctly.
Caveat: virt-manager is not installed on this machine, so the absent path was
exercised for real and the write path only against a stubbed gsettings.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
~/.ssh/config entries were only ever written while deploying a VM. A fleet
that already exists — or one whose DHCP leases have moved — had no way to
refresh them. Entry [13] of the QEMU menu does it on demand:
[1] update ~/.ssh/config for the local VMs
[2] add the nested VMs through ProxyJump, recursively
The second one matters because of the « ERPLibre Deployment (+ QEMU + dev) »
profile: a VM built that way hosts VMs of its own, on its own private network.
Those are not reachable from this host at all — only from their parent. So
each level is written with a ProxyJump to the level above, and OpenSSH chains
the hops on its own. « ssh erplibre-fedora-42 » then works from here even
though the address only means something two machines away.
The recursion probes over « ssh <alias> », i.e. through the block just
written, so the parent's own ProxyJump applies automatically and one probe
works identically at any depth. One SSH connection per MACHINE, not per VM: a
single snippet returns every « name<TAB>ip » pair, falling back to the guest
agent when the dnsmasq lease is missing. Passwordless sudo is a given here —
the cloud-init config grants it (deploy_qemu.py:1175).
A nested VM keeps its short name, which is what one wants to type, and is only
prefixed with its parent on collision — so discovering a machine that already
exists elsewhere never overwrites the other one's entry. Depth defaults to 2
(host, VM, nested VM) and already-seen aliases are skipped, which is what
stops a cycle: a child that reports its own grandparent.
Verified on a simulated two-level fleet including a deliberate cycle: the
ProxyJump chain is correct at each level, the colliding name is prefixed, the
parent block is not overwritten, and only one probe per machine is issued. The
remote snippet itself was run for real on this host — valid POSIX sh, two VMs
with their addresses. A VM without an IP is skipped rather than written with
an empty HostName.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Uninstalling before the 13->14 bump asked for 15 modules; 12 of them no longer
exist in the 13.0 addons path. check_addons_exist.py rejects the command on
the first missing name, so the uninstall aborted whole — and the 3 modules
that WERE there stayed installed, blocking the bump for a reason that had
nothing to do with them.
The list is now split before the script is called, using check_addons_exist
which was already there and simply never consulted at this point. « Missing »
means the addons path has no code for it, not that the database lacks it —
the distinction is the whole point, since Odoo cannot uninstall a module whose
code is gone.
When some are missing, a choice is offered rather than a failure:
[1] uninstall the present ones, skip the missing (default)
[2] try the whole list anyway (it will fail)
[3] uninstall nothing, continue
Only what is really uninstalled is subtracted from the per-version module
list, so a module left in place stays counted as installed — which it is. The
missing names are written to the progression comments, so what was skipped is
still on record afterwards.
Verified on the exact list from the failure: 3 present and 12 missing,
option 1 issuing a command with only the 3 and leaving the 12 counted as
installed, option 2 sending all 15, option 3 running nothing.
Also swapped entries 4 and 5 of the Deploy menu — QEMU/KVM now sits at [4],
NTFY at [5].
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The migration now opens with the same choice the QEMU deployment does — TUI
or line-by-line prompts — settled in advance by TODO > Configuration if you
want it to be. Choosing the TUI loads the saved progression and asks the very
same first question: where does this migration stand, and where do we resume?
prompt_resume was one function that rendered, read and decided at once, so a
second view would have meant a second copy of the decision. It is now three:
resume_context() the progression -> plain data (file, database, target,
steps with their icon and detail, version bumps)
print_resume() renders it on the terminal
apply_resume_answer() answer -> (progression, changed)
The TUI returns the SAME answer strings as the prompt — c, n, r, q, 0..4,
4.<version> — so apply_resume_answer stays the only place that decides what a
choice means. Neither view can drift into describing the migration
differently, because both render the same context.
In the TUI the steps are a table and the version bumps a list: Enter on either
replays from there, which is what « [0-4] » and « [4.N] » meant in text. The
cursor opens on the first unfinished step and on the first unmigrated version
— where it stopped is where one usually wants to act.
Both views gain « q », quit without doing anything. A TUI needs Escape to do
something sane, and an escape hatch present in only one of the two views is
exactly the kind of divergence this split exists to prevent. execute_odoo_
upgrade returns immediately on it, writing nothing.
Verified on a 12->18 progression with steps 0-3 done and 2 of 6 bumps
migrated: identical context feeding both views, Enter on step 2 giving « 2 »,
Enter on the 15 bump giving « 4.15 », the c/n/r/q/Escape shortcuts, and every
answer producing the same progression through both paths — « 2 » leaving only
the state of steps 0 and 1, « 4.15 » resetting the clone list from the third
bump on so the half-migrated intermediate database gets rebuilt. « q » checked
to leave the progression file byte-identical.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The form opened with the four main versions already ticked, so the plan showed
four VMs nobody had asked for. F5 on an untouched form would have created
them. Deploying is expensive and hard to undo — the list now starts empty and
the choice is made, not inherited.
The « * » still marks each distro's main version, and F7 ticks exactly those
four in one keystroke, so nothing is lost but the presumption.
With an empty list a « 0 VM · 0 vCPU » total teaches nothing, so the footer
says what to do instead: tick what to deploy, F7 main versions, F6 all.
Verified headless: no box ticked on open, F5 producing no spec at all, F7
giving back the four, and the ticks surviving a switch to « all
architectures » (4 of 30) since identity is distro/version/arch, not rank.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Deploying a VM meant answering a dozen questions in a row, never seeing the
whole of one's choices, and starting over to revisit an answer. The first
question is now which interface to use — TUI form or the classic prompts —
defaulting to whatever TODO > Configuration says.
The form shows every setting at once with a plan that recomputes on each
change: names, resources, and the two collisions marked as they arise (a
defined VM is skipped, an orphan qcow2 will make deploy_qemu fail). F2 edits
one VM, F3 previews the commands, F5 deploys, F6/F7/F8 select all / main
versions / none.
Function keys rather than ctrl+letter: ctrl+p is Textual's command palette and
silently swallowed the shortcut, and a bare letter is eaten by whichever input
has focus.
Both interfaces go through _qemu_deploy_parts_for, so the same choices produce
the same command by construction — and a test now drives the CLI prompts and
the form to the same state and compares the argv, which is what keeps them
from drifting.
Nothing privileged or networked runs inside the form. Every virsh call in this
codebase goes through sudo, and a password prompt while Textual owns the
terminal would wreck the display; the domain list and the remote branches are
fetched before the app starts and arrive as plain data.
The progress view is optional (preference: CLI output stays the default, since
it is the easiest to copy from). It gives one collapsible block per VM,
expanded while running, folded on success — and left OPEN on failure, which is
the part worth reading. « c » copies the selected log, « C » all of them.
On copying from a TUI over SSH: copy_to_clipboard emits OSC 52, which the
LOCAL terminal emulator interprets, so it does reach the workstation's
clipboard. Two caveats are handled: the payload is capped at 100 kB keeping
the TAIL (some terminals truncate long OSC 52), and the notification says a
compatible terminal is required — macOS Terminal.app has none, tmux needs
set-clipboard on.
Verified: pure logic (profiles, per-VM overrides surviving a profile change,
totals, statuses) by direct calls; the form headless via run_test — default
selection, F6/F7/F8, profile switch, and the orphan guard demanding a second
F5; parity CLI/TUI on identical argv; the progress view on three fake jobs,
with the failure staying open and the clipboard filled. The semaphore bounding
concurrency was measured: four 0.35 s jobs take 1.63 s at 1 and 0.58 s at 4 —
without it, « 4 in parallel » launched every VM at once.
Also removed a duplicated @staticmethod on _qemu_install_dir, harmless since
Python 3.10 but misleading.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
_qemu_deploy was ~400 lines mixing a dozen prompts with the parallel
deployment, IP resolution, ~/.ssh/config and the ERPLibre install. A second
interface could not be added to it without duplicating all of that — and the
repository already shows where duplication leads: _qemu_choose_cli_browser
(todo.py) and Monitor._choose_browser (qemu_install_monitor.py) are the same
function twice, and they have already drifted apart.
The function now has three parts around a plain dict, the SPEC:
_qemu_collect_vms_cli arch, catalog, resources, names -> the VM list
_qemu_collect_options_cli SSH key, install, parallelism -> the spec
_qemu_run_spec consumes a spec, asks nothing
_qemu_deploy_parts_for is the single point every command goes through, so two
interfaces producing the same spec necessarily produce the same command —
which makes their divergence testable rather than a matter of discipline.
Pure, I/O-free helpers come out of the body: _qemu_catalog_entries (the flat
distro × version × arch list), _qemu_arches_for, _qemu_make_vm,
_qemu_split_existing and _qemu_orphan_disks. The last two matter beyond
tidiness — the form must recompute collisions on every keystroke, and every
virsh call in this file goes through sudo. Existence is now resolved with ONE
virsh list --all --name (_qemu_list_domains, already present) instead of one
sudo per VM, and the orphan-disk scan needs no privileges at all.
VMs travel as dicts rather than 6-tuples plus a parallel names list. The
tuple-based resource prompts are left untouched and converted at the boundary.
No behaviour change intended, and verified as such: replaying identical
scripted answers against a worktree pinned at the previous commit produces
byte-identical output — including the granular selection spanning three
architectures — once the repository path and the instantaneous free-RAM
reading are normalised.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Nothing in the CLI could be remembered from one session to the next except the
language, which lives in env_var.sh with its own ad-hoc parser. Entry [6]
Configuration, below Telemetry, now groups the settings that belong to the
USER rather than to the repository.
todo_prefs.py stores them in ~/.erplibre/todo_prefs.json — the same place and
the same best-effort shape as todo_telemetry.py, since both are per-user and
per-machine, and neither must ever prevent the CLI from starting. A DEFAULTS
table gives every known key its fallback, so a missing or corrupt file simply
reads as defaults.
The two keys added here prepare the QEMU deploy form: which interface to use
(ask / TUI / classic) and what to display while deploying (CLI output or TUI).
They are declared once in _PREF_CHOICES — the screen, the current-value label
and the editor all derive from that single table, so adding a preference is
one entry, not three edits.
Language keeps its own mechanism: it is read before the preferences file
exists and is consumed by shell scripts too.
Verified against a temporary HOME: defaults returned with no file on disk, a
change persisted and reflected in the menu, and reset falling back to
defaults.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Resources per VM offered x1..x4 only — anything else meant deploying with the
wrong sizing and resizing afterwards. « [5] Custom » now asks vCPU, RAM and
disk once for the whole fleet, and vCPU joins disk and RAM in the per-VM
customisation. Presets 1, 2, 4, 6, 8, 16, 24, 32.
The two places ask the SAME three questions, so they are one definition each
(_qemu_ask_disk / _qemu_ask_ram / _qemu_ask_cpu) instead of two copies that
would drift. vCPU is no longer a single global value: it travels per VM
through the whole chain, down to --vcpus.
Overcommit is allowed on an explicit choice — KVM permits more vCPU than
cores — and only warned about. The x1..x4 path still caps at the host core
count: that one is an automatic computation, not a decision.
Three prompts default to yes (install ERPLibre, ~/.ssh/config, deploy): they
are what one wants nearly every time, and typing « o » on each was noise.
Flipping a destructive-by-omission default demanded the guards below.
Before deploying, a final review lists everything that will change — VMs to
create with their effective disk (ERPLibre's +5 G included), VMs left
untouched, install profile and branch, SSH key, ~/.ssh/config, parallelism.
Answering no asks for confirmation rather than dropping every answer given
over a dozen prompts.
Name collisions are now reported BEFORE the wait, with their two very
different consequences: an already-defined VM is skipped and nothing is
overwritten, while a qcow2 left behind by a deleted VM makes deploy_qemu fail,
since it refuses to overwrite without --force. Continuing requires an explicit
yes; the default is no.
Also translated two strings of this flow that had no entry and stayed in
English (« Resolving VM IPs… », « no IP »).
Verified by driving the prompts with scripted answers: custom profile applied
and left intact when blank, per-VM override, 32 vCPU on 28 cores warned but
accepted, review with and without ERPLibre install, no → no → deploy, no →
yes → cancel, and collisions defaulting to no. Rendering checked in fr and en.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The install history was only surfaced as a « ~5m avg (3) » suffix next to a
distro. Entry [12] now shows what that history actually contains: totals and
success rate, period covered, median/min/max and cumulated duration, then a
breakdown by distribution, version and architecture, plus the current VMs and
the disk they occupy. « [r] » erases the history after confirmation.
record_duration() was only called on success, so no success rate could ever be
computed. Failures are now recorded with ok=False. They are counted separately
and EXCLUDED from the averages and the ETA: how long a failed install ran says
nothing about how long a successful one takes. Entries written before the flag
existed have no « ok » key and are read as successes, which they were.
Aggregation lives in qemu_install_monitor.py as pure functions (stats_summary,
stats_by, all_runs, reset_stats), display in todo.py — the split the file
already follows.
Disk presets extended to 400G, 600G, 800G, 1T, 1.5T and 2T. The parser only
understood G, so « 1T » would have been rejected: sizes are now normalised
through _qemu_parse_disk (1 T = 1024 G, decimal comma accepted), since the rest
of the chain reasons in gigabytes.
Verified with a synthetic history of 9 runs including 2 failures: the rate,
the per-group failure counts and the reset all behave; a group with no success
shows « — » rather than a misleading « ~0s ».
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The presets stopped at 32768 MB; current virtualization hosts go well beyond
that. The series now doubles up to 262144 MB (256G).
The prompt is in MB while people think in GB, so each preset carries its
equivalent: « [g] 65536 (64G) ». With nine of them the line no longer fits, so
suggestions are laid out five per row. Disk presets keep the plain format.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Both prompts expected a raw number, so common values had to be retyped every
time. They now list suggestions:
[a] 20G [b] 40G [c] 60G [d] 80G [e] 120G [f] 200G
New disk size in G, blank = keep (20G):
Letters start at « a » so they can never collide with a value typed directly:
anything starting with a digit is read as the value itself, and blank still
keeps the current one. A letter outside the range is rejected instead of being
silently taken for a size.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
_qemu_default_ssh_key() returns '' when the user has no public key, so the
deployment injects none through cloud-init: the VM boots with no SSH access and
its state can no longer be checked once created.
--setup-host now generates an ed25519 key without passphrase when neither
id_ed25519.pub nor id_rsa.pub is present, so the Deployment profile leaves a
host able to reach the VMs it creates.
The key is created AS the invoking user via « sudo -u », not as root: under
sudo « ~ » is /root and the key would land where cloud-init never looks.
Verified on a fresh Arch VM with no key: it is created as erplibre:erplibre
with 700/600/644, a rerun keeps the same fingerprint, and
_qemu_default_ssh_key() now returns it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When OpenUpgrade failed, todo_upgrade_execute offered the generic
« [1] to redo the command ». Replaying it re-ran OpenUpgrade on the database it
had just half migrated, which never recovers. That call now passes
wait_at_error=False, and the failure message explains that the clone step has
been reset so a relaunch DROPS and REBUILDS the intermediate database from the
previous version.
The resume menu gains « [4.N] »: replay the step 4 loop from one version bump.
Only the per-version lists are trimmed from that index, so earlier bumps stay
migrated and steps 0 to 3 are untouched. Resetting the clone entry is the whole
point: the intermediate database of a failed bump must be rebuilt, not upgraded
again. « [4] » still replays every bump.
The per-bump uninstall files were also read at step 1 only, for the source
version, so uninstall_module_list_odoo130_to_odoo140.txt existed in name but
was never read. The step 4 loop now reads the file of the bump it is about to
perform, merged with the answers already stored in the progression.
Verified against the real progression (13 and 14 migrated): the menu offers
13/14/15/16/17/18, and replaying from 14 turns every per-version list from
[True, True, False...] into [True, False, ...] while state_4_reach_open_upgrade
and steps 0-3 survive. A 13->14 list of 4 modules is read with its reasons.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The 13.0 -> 14.0 data migration died on:
duplicate key value violates unique constraint "res_groups_name_uniq"
Key (category_id, name)=(9, Allow to define fiscal years of more or less
than a year) already exists
In 13.0 that security group is declared by the core « account » module, so the
database holds it as account.group_fiscal_year. In 14.0 core account no longer
declares it and om_account_accountant (odoomates) does. Owning no XML id of its
own, the module tries to CREATE the group and collides with the existing row.
Renaming the XML id makes Odoo UPDATE the existing record instead. The record
id is untouched, so user assignments, access rights and record rules pointing
at the group survive -- on the reference database it holds none, but the fix
must not depend on that.
The fix hook also learns a « .sql » flavour, run through psql. The « .py »
flavour is piped into « odoo<target>-bin shell », which requires loading a
not-yet-migrated database with the TARGET version's registry -- precisely what
is failing at that point. A pure SQL fix needs no ORM and no registry.
Verified end to end on a throwaway copy of the stuck database: the statement
renames one row, a second run changes nothing, and the full 13->14 OpenUpgrade
then completes (« Modules loaded. », base at 14.0.1.3, zero
res_groups_name_uniq error).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On a rolling distro every fresh development VM was born unable to create VMs.
« make install_os » runs a full system upgrade, the kernel package is replaced
and /lib/modules/<running kernel> disappears. modprobe bridge then fails and
libvirt cannot create virbr0, so the « default » network stays inactive and
virt-install dies on « network 'default' is not active » -- with every package
correctly installed. Measured twice on a freshly created Arch VM: booted on
7.1.3-arch1-3 at 05:44, upgraded to 7.1.5.arch1-2 at 05:46.
Only a reboot fixes it, and one is enough: the default network is already
flagged autostart, so libvirt brings it up by itself once the modules match.
--setup-host therefore gains --reboot-if-needed, used by the deployment
profile. The reboot is scheduled through systemd-run --on-active=5 rather than
issued immediately, otherwise it would kill the installer's SSH session and the
orchestrator would report a failure for a VM that actually succeeded. Without
the flag the behaviour is unchanged: explain and exit 1, which is what a
workstation wants.
Also fix « Error setting up logfile: No write access to
/var/tmp/erplibre-virtinst/virt-manager »: that path was shared, so a first run
under sudo created it as root and later non-root runs could not write. It is
now per-UID.
Verified on the Arch VM that failed: without the flag it exits 1 with the
diagnosis; with it, the reboot is scheduled and the command still returns 0;
after the reboot the kernel and modules match, « default » is active on its own
and virbr0 exists. The cache directory is created as erplibre:erplibre 0700.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The « ERPLibre Deployment (+ QEMU + dev) » profile installed packages with a
one-liner ending in « || true » and hiding stderr, so a broken host looked
installed. Reproduced on erplibre-arch-latest: every binary was present, yet
virt-install failed and « virsh list --all » could not reach any hypervisor.
Three distinct causes, all unhandled:
- The user was never added to the libvirt group. Without it a non-root libvirt
client falls back to qemu:///session, where the « default » network does not
exist, so « --network network=default » fails while everything looks
installed. Being in the group grants the RIGHT to reach qemu:///system but
does NOT change the default URI, so virt-install and virsh now pass
« --connect qemu:///system » explicitly (new LIBVIRT_URI constant).
- dnsmasq was missing: ensure_tools only installed DAEMON_PACKAGES when the
daemon was absent, and libvirt was already there. setup_host now forces them.
- On a rolling distro, « make install_os » upgrades the kernel and the package
manager removes /lib/modules/<running kernel>. modprobe bridge then fails and
libvirt cannot create virbr0 (« Unable to create bridge virbr0: Package not
installed »). No package fixes that, only a reboot: it is now diagnosed and
reported instead of surfacing as an unreadable virsh error.
The profile now calls « deploy_qemu.py --setup-host », which reuses the
existing ensure_* chain, so package names stay defined in one place
(TOOL_PACKAGES / DAEMON_PACKAGES) and keep working for apt, dnf, pacman,
zypper and brew. It fails loudly instead of « || true ».
Verified on the Arch VM that failed: setup-host reports the stale kernel and
exits 1; after a reboot it installs dnsmasq, joins libvirt/kvm, starts the
default network, and reports « Hôte prêt » (exit 0, idempotent on rerun).
The generated command now reads « virt-install --connect qemu:///system ».
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The resume menu exposed internal key names: « Reuse database without state_4 »
tells nobody what will happen, and the most useful option -- keep the prepared
database, redo the version bumps -- was the least readable of the four.
It now prints where the migration actually stands: the zip, the database, the
target, and one line per step with its state. Step 4 reports how many version
bumps are migrated and names them, so « 0/6 · 13 14 15 16 17 18 » replaces a
list index nobody could interpret.
The four numbered options become:
[c] continue where it stopped (was: press enter)
[0-4] replay from that step (new: rewind to any step)
[n] new migration, erase everything (was: [1])
[r] keep the zip, ask everything (was: [4])
Rewinding drops « state_* » from the chosen step onward and keeps the rest:
config_* answers, the zip and the target version are decisions, not progress,
so they are no longer re-asked. state_0_search_missing_module is still forced
back to False, as the previous code did: it fills an in-memory dict the later
steps rely on. Old [2] is now « replay from 0 », old [3] « replay from 4 ».
First use of t() in this file; the new strings are translated in todo_i18n.py.
Verified against the real progression of the 12.0 -> 18.0 migration and a
synthetic one: for every step, the kept keys are exactly those of the earlier
steps, config and target survive, and the module search is forced again.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Measured on erplibre-ubuntu-2404: unattended-upgrades fired in the middle of
an Odoo 12->13 migration and restarted the PostgreSQL cluster three times
(« received fast shutdown request »). OpenUpgrade lost its connection and the
intermediate database was left half migrated.
On a development VM the ERPLibre installer now turns off unattended-upgrades
and the apt-daily timers, and drops an apt.conf.d snippet so they stay off
across reboots. dnf-automatic gets the same treatment on Fedora. It runs right
after the cloud-init wait and before the apt-get calls, so apt-daily can no
longer grab the lock between the two either -- the same contention that made
« apt-get update » fail during deployment.
Production VMs are left untouched: automatic security updates must stay on
there. The switch is the existing dev/prod answer, already threaded down to
_qemu_erplibre_remote_cmd.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When OpenUpgrade fails, the intermediate <db>_upgrade_<N> database is left
half migrated. The clone step had already been recorded as done, so a rerun
skipped it and restarted OpenUpgrade on top of that broken clone.
Clear the clone flag on failure so the replay rebuilds the intermediate
database from the pristine source. Found by a real 12.0 -> 13.0 run: Ubuntu
unattended-upgrades restarted the PostgreSQL cluster mid-migration, Odoo lost
its connection, and the replay would have resumed on the damaged clone.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
key is the only thing pairing a website copy with the module view it came
from. A copy whose key matches nothing is never paired, so it never receives
the new inherit_id, never changes shape, and never joins a view combination.
Renaming the key is therefore enough to take it out of the way:
UPDATE ir_ui_view SET key = '<prefix>.' || key, active = false
active = false alone would NOT work: an inactive copy keeping the same key
still shadows the module view. Nothing is deleted either, so inherit_id
ondelete='restrict' and the website_page foreign keys are never touched, and
the old arch stays in database as a readable archive. --restore undoes it.
Written in plain psql on purpose: it runs on a database not yet migrated,
where starting an Odoo shell of the target version is not guaranteed. It is a
systematic step driven by the detector, valid for every bump, so it lives in
the loop rather than in a per-version fix_migration file.
Offered rather than forced, since those copies carry real customizations.
Verified on a throwaway copy: dry-run writes nothing; apply renames the key and
deactivates while keeping the 4058-char arch; a second apply is a no-op;
restore returns the row to its exact initial state.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A version bump rewrites website views without announcing any of it, so "the
site looks wrong" after a migration is currently unanswerable.
snapshot_cow_views.py records every website_id view (key, mode, inherit_id,
active, arch and its md5) and diffs two snapshots. Rows come back as JSON
straight from Postgres because an arch holds newlines and pipes; the column
list is intersected with information_schema, since ir_ui_view does not expose
the same columns from 12.0 to 18.0. Snapshots hold customer template content,
so they go under private/ and stay out of git.
todo_upgrade.py takes one before and one after each OpenUpgrade run, then
prints the diff. Both are non-blocking: forensic material must never stop an
upgrade.
Measured on the real 12.0 -> 13.0 jump, the diff shows what no log reported:
71 -> 68 copies, 16 deleted and 13 created. portal.frontend_layout is not
converted but DELETED (id 2670) and RECREATED (id 3397), with its children
re-parented from one to the other. The four theme_technolibre copies,
muk_web_branding, project_agile and erplibre_website_snippets_basic_html are
dropped outright, and website_crm.contactus_thanks is renamed to
website_form.contactus_thanks.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two blind spots made the detector miss real breakages.
1. Comparing « mode » is not the right test. What decides the shape an arch
must have is whether the target declares an inherit_id: with one, the arch
must be inheritance specs (<data>, <xpath>, position=); without one, it must
be a standalone template. A view moving from a root template to
« inherit_id + primary="True" » keeps mode='primary' on BOTH sides, so the
old test reported nothing while the copy still broke. The comparison is now
(target declares inherit) vs (stored arch is spec-shaped), and each finding
carries the reason. A mode change with a matching shape is still reported,
as a lesser warning.
arch_db is read as text up to 15.0 and as jsonb from 16.0; both are handled.
2. A module renamed upstream was reported as « module absent », hiding every
view it owns. renamed_modules is now read from the target OpenUpgrade
apriori.py (21 entries for 13.0, 56 for 14.0, 39 for 16.0, 20 for 18.0) and
used before concluding the module is gone.
Also correct the advice printed for a copy at risk: deactivating it is not
enough, an inactive copy keeping the same key still shadows the module view.
Renaming the key is what actually unpairs it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A copy-on-write view freezes the structure of the module view it was copied
from. When that module view changes mode between two versions, the copy keeps
an arch written for the old mode and the upgrade dies on
« Element ... cannot be located in parent view », hours into the migration.
Measured on a real 12.0 database: portal.frontend_layout is declared primary in
12.0 (full QWeb template, so a full-template arch is legitimate) and becomes an
extension in 13.0 (inherit_id + xpath). The 2021 COW copy follows the module,
turns into an extension, and keeps its 12.0 arch. The customization is genuine;
the breakage is produced by the migration.
check_cow_views.py compares the mode stored in database with the mode declared
in the target version sources and sorts the views in three buckets: at risk
(mode change), module absent from the target version, and pages made from the
editor (not at risk). Read-only, and called from step 2 so the list is known
before the long version loop rather than during it.
On the reference database it reports exactly the view that broke the upgrade,
plus 7 views whose module is gone in 13.0 (custom theme included).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Which modules must be uninstalled before a version bump depends on the data of
one specific database, so that list does not belong to a shared versioned file.
- Read private/odoo/migration/<database>/uninstall_module_list_odooXX0_to_odooYY0.txt
first, then the shared versioned defaults under script/odoo/migration/, and
merge them (duplicates dropped). private/ mirrors script/, the convention
already used by script/todo/todo.json -> private/todo/todo_override.json.
- Fix the parser. It was « f.readline().split() »: only the FIRST line was kept,
so a multi-line list was silently truncated, and a comma-separated list turned
into one bogus module name. On a file starting with a comment it returned the
words of that comment as module names. It now reads every line and accepts
commas, several names per line, blank lines and comments.
- Support a « # reason » justification per module, printed before uninstalling;
a module with no reason is flagged. Removing a module must stay reviewable.
- Ignore private/odoo/ in git: these lists describe one database.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The database migration recorded steps as done even when they had failed, so a
resume skipped them and the loop kept going on a broken database.
execute.py
- exec_command_live() left exit_code = None when an exception was raised. None
is falsy, so « if not status: » marked the step done and
« if status and wait_at_error » skipped the error prompt: a crashed command
was reported as a success. Both except blocks now set exit_code = 1.
- The « no Odoo version installed » path returned a bare -1 while callers
unpack a tuple (status, cmd), raising ValueError instead of surfacing the
failure. It now returns the same shape the caller asked for.
todo_upgrade.py
- todo_upgrade_execute(): treat a None status as a failure (defence in depth).
- Database migration (OpenUpgrade): the return code was discarded, with an
explicit « TODO detect error », and the state was written unconditionally.
It is now captured; on failure the loop stops instead of migrating the next
version on top of a half-migrated database.
- Neutralization: the flag was set before the « if not status » test, making
that test dead code. Removed the unconditional assignment.
- Clone and fix-migration hook: their return codes were never captured, so the
steps were marked done whatever happened. Both are now checked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A leftover git daemon from an interrupted run kept port 9418, so the new
daemon failed to bind ("Address already in use") and the final « kill
$DAEMON_PID » failed on the already-dead PID, returning exit 1. That non-zero
status bubbled up through apply_extra_modules() and set exit_code=1, which
silently skipped the post-install steps (add_extra_to_config_conf,
generate_config) -> CybroOdoo never landed in config.conf.
- Kill any leftover daemon before starting a new one. Match the stable
arguments, not "git daemon": « git daemon » execs into « git-daemon »
(hyphen, /usr/lib/git-core/git-daemon), so a "git daemon" (space) pattern
never matched the actual process.
- Move daemon cleanup into an EXIT trap that tolerates an already-dead PID,
so the script never exits non-zero just because the daemon is already gone.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The extra modules (CybroOdoo) were still not applied on an already-installed
environment even after apply_extra_modules() was added. Two root causes:
1. When update_env_version.py runs as a script, sys.path[0] is
script/version/, so « from script.version.erplibre_state import ... »
raised ImportError and _STATE_AVAILABLE fell back to False. The whole
state mechanism was silently disabled: set_version_installed() never
wrote .erplibre-state.json, get_version_extra() stayed False, and the
extra manifest was never merged -> CybroOdoo never cloned. Add a sibling
import fallback so the state module also resolves under script execution.
2. generate_config.sh rewrites config.conf from a fixed addons list that
omits the extra modules. Add add_extra_to_config_conf(), called AFTER
generate_config.sh (last writer wins), which appends the extra addons
paths (derived from the "extra" group of the extra manifest) to
addons_path. Idempotent: already-present paths are skipped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
« Installation avec modules extra » (--with_extra) ne clonait pas CybroOdoo
sur un env déjà installé : validate_environment renvoyait « valide » ->
update_environment (où l'état extra est posé via set_version_installed) était
sauté -> get_version_extra restait False -> git_merge_repo_manifest n'ajoutait
pas manifest/git_manifest_extra_odooXX.xml -> CybroOdoo jamais synchronisé ni
ajouté à config.conf (« Nothing to do »).
Nouvelle étape apply_extra_modules(), appelée dans main() quand --with_extra
est demandé (même si l'env de base est déjà installé) : pose l'état extra,
régénère le manifest local + repo sync (clone CybroOdoo). La post-étape qui
suit régénère config.conf, qui prend alors le nouveau chemin d'addons.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le montage sshfs de « Configurer sshfs » ajoute désormais
« -o follow_symlinks » : sshfs résout les symlinks côté serveur. Sans ça,
git échoue sur les worktrees google-repo d'ERPLibre (leur .git est une
chaîne de symlinks relatifs profonds vers .repo/projects et project-objects)
-> « erreur à la lecture de .git », git status/commit impossibles sur le
montage.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
« Code » (et Update, Git, Database…) apparaissait vide : l'extracteur AST ne
comprenait que le motif « status == "N": self.methode() » (menu QEMU). Les
menus bâtis par « choices = config_file.get_config(...) » + « choices.append(
...) » avec dispatch « str(len(choices)-N) » n'exposaient aucun enfant ->
impossible d'ouvrir leurs commandes.
Ajout de l'extraction de ce motif :
- entrées de CONFIG lues depuis todo.json (rejouées via
execute_from_configuration avec l'entrée en kwargs) ;
- entrées APPENDÉES (dict littéral OU variable menu_entry=… suivie par n° de
ligne) mappées à leur méthode via le dispatch « str(len(choices)-K) » ;
- une entrée qui ouvre un sous-menu (ex. Update) est développée récursivement.
Validé : Code -> 7 commandes (statut/remiser/formater/SHELL/màj module/débogage
+ sous-menu Update), méthodes correctes et existantes, TUI monte sans erreur.
Les dispatches complexes (ex. « Migration BD » via TodoUpgrade) restent
visibles mais non lançables (method None).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
L'entrée « Update - Update all developed staging source code » quitte le
menu Execute (renumérotation 9->8 … 15->14, dispatch elif mis à jour) et
rejoint le sous-menu « Code - Outil pour développeur » (dernière entrée,
appelle prompt_execute_update). Le libellé/i18n (🔃) est réutilisé.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
webdriver.Firefox échouait : « InvalidArgumentException: Argument
--remote-allow-system-access can't be set via capabilities ». Cet argument
(ajouté pour le Firefox snap d'Ubuntu) est désormais REFUSÉ via les
capabilities par les geckodriver récents, qui gèrent seuls l'accès système
du snap. Il cassait donc TOUS les setups sur geckodriver récent (ici Arch
sans snap). On ne le passe plus.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
extraire_svg_graphique_by_3d annote « -> Optional[str] » mais le module
n'importait pas Optional -> « NameError: name 'Optional' is not defined »
au chargement. Ajout de « from typing import Optional ».
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
« Renommer quelles VM ? » devient « Modifier quelles VM ? » : pour chaque
VM choisie, on peut changer le NOM, la taille de DISQUE et la RAM (vide =
garder). _qemu_customize_names -> _qemu_customize_vms renvoie (names,
selected) avec les valeurs mises à jour.
Le multiplicateur de ressources est désormais CUIT dans `selected` (RAM
finale) et le nombre de vCPU fixé une fois (uniforme) -> un override de RAM
par VM est une valeur ABSOLUE, sans ambiguïté avec le multiplicateur. La
fonction eff_res disparaît (plan, dry-run et déploiement lisent les valeurs
finales de `selected`).
Validé : modif VM 1 (nom/disque/RAM) appliquée, VM 2 inchangée.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- État : 10 -> 8. Le libellé « ❌ effacée » (10 cellules) forçait la largeur ;
l'état « effacée » devient l'icône seule 🗑. Reste le plus long « ⏸ pause »
(7), donc 8 suffit.
- VM : 26 -> 22 (tient les noms standards « erplibre-ubuntu-2604 » = 20 ; les
rares noms suffixés d'arch défilent via overflow-x).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
« SELinuxContext=unconfined_u:unconfined_r:unconfined_t:s0 » ne débloque PAS
l'exécution : la transition init_t -> unconfined_t est refusée par la
politique -> le service échouait toujours en 203/EXEC (« Permission denied »
sur run.sh, contexte user_home_t) sur Fedora.
En DEV (VM jetable, « SELinux relâché ») on passe désormais SELinux en
PERMISSIF (setenforce 0 + persistance dans /etc/selinux/config) si actif ->
le service exécute run.sh/venv sous /home sans blocage.
PROD reste confiné (/opt/erplibre hors user_home_t + restorecon) — à
peaufiner quand on reviendra sur la prod.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
La colonne « État » réservait 12 caractères alors que le plus long libellé
(« ❌ effacée ») en fait 10 -> elle paraissait trop large. Ramenée à 10
(sans troncature).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le fix précédent (DPkg::Lock::Timeout) ne suffisait pas : cette option NE
couvre PAS le verrou /var/lib/apt/lists/lock de « apt-get update ». Au 1er
boot, cloud-init/apt-daily tient ce verrou -> update échouait AUSSITÔT
-> lists vides -> « Unable to locate package git » (exit 100), toujours sur
debian-12.
On RÉESSAIE désormais « apt-get update » jusqu'à libération du verrou (et
lists peuplées), borné à ~5 min (30×10s), avant l'install. Une fois update
OK, git/make s'installent normalement.
Note : le remote_cmd est construit côté HÔTE -> relancer le CLI todo pour
que la nouvelle commande soit envoyée aux VM.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Nouvelle 1re colonne « # » : numéro de séquence de chaque VM.
- Sommaire enrichi du « max » de durée : la VM TERMINÉE la plus lente
(pire cas), à côté du temps global et de l'ETA.
Validé headless : # = 1..N, sub_title « … · max MM:SS ».
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Au déploiement (« Déployer une ou plusieurs VM »), après le choix de la
branche, on demande l'ENVIRONNEMENT cible (dev par défaut / prod).
- DEV : ERPLibre dans ~/git/erplibre ; service systemd unconfined si
SELinux actif (comportement inchangé).
- PROD : ERPLibre dans /opt/erplibre (sudo git clone + chown à
l'utilisateur) ; service systemd CONFINÉ par SELinux (pas d'unconfined)
+ restorecon des contextes -> hors user_home_t, un service peut exécuter
le contenu sans lever le confinement.
Le drapeau `prod` est propagé : _qemu_deploy -> _qemu_install_erplibre_
monitored/_vm -> _qemu_erplibre_remote_cmd -> _qemu_odoo_service_cmd.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
« -v » montre l'étape mais pas la sortie des sous-processus. Seul « -vvv »
(debug) affiche git clone/checkout et le build pip -> l'erreur réelle d'une
dépendance VCS/build. Le rejeu de diagnostic passe donc en -vvv.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Corrections issues de la revue complète des 13 VM (6 en échec réel) :
- APT lock (debian-12) : au 1er boot, cloud-init/apt-daily tient le verrou
et « apt-get install » échouait aussitôt. Ajout de
« -o DPkg::Lock::Timeout=600 » (update + install) -> attend le verrou.
- wkhtmltopdf Debian 13 « trixie » (debian-13) : on prenait le build
« bullseye » (dépend de libssl1.1, absent de trixie) -> gdebi échouait.
Désormais bookworm pour bookworm/trixie/+. ET l'échec gdebi est NON
bloquant (avertissement au lieu d'exit 1 : wkhtmltopdf est optionnel,
sinon tout install_os avortait et Odoo n'était jamais installé).
- SELinux 203/EXEC (fedora-41/43/44) : un service système ne peut pas
exécuter run.sh/venv sous /home (contexte user_home_t). Ajout
conditionnel de « SELinuxContext=unconfined_u:unconfined_r:unconfined_t:s0 »
au service (uniquement si getenforce != Disabled).
- poetry status 1 (ubuntu-2604) : « -q » masquait la cause. À l'échec, on
rejoue « poetry install -v » pour capturer l'erreur dans le log.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Nouvelle colonne « Odoo » dans le dashboard : teste par TCP si l'UI web
répond sur le port 8069 de la VM (🟢 up / — sinon). Une fois « up »
détecté, on ne re-teste plus (Odoo ne redescend pas en cours d'install).
Le test (socket, timeout 0,5 s) est fait dans le thread collecteur ->
n'impacte pas la fluidité de l'UI.
Validé headless : port ouvert -> 🟢, fermé -> —.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Déploiement PARALLÈLE de 13 VM -> 2 problèmes observés :
1. « Aucune image Fedora-Cloud-Base-Generic pour Fedora 42 » alors que
l'image EXISTE. download.fedoraproject.org est un redirecteur
(MirrorManager) : sous requêtes parallèles, une VM tombait sur un miroir
incomplet -> index HTML sans correspondance -> sys.exit.
resolve_fedora_url réessaie (2× le redirecteur) puis se replie sur le
serveur MAÎTRE dl.fedoraproject.org (toujours complet). Validé : F41/F42
résolus.
2. Pavé « Fetched capabilities … » (énorme XML) dans la sortie : erreur de
LOGGING de virtinst — sous sudo, HOME/cache inaccessible -> échec
d'écriture du journal debug -> Python déverse l'enregistrement raté.
On préfixe virt-install de « env XDG_CACHE_HOME=… HOME=… » (traverse
sudo) vers un cache écrivable -> journal silencieux.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Cause des GROS ralentissements en changeant de VM / lisant les logs :
- on_data_table_row_highlighted rechargeait le log à CHAQUE mouvement du
curseur (donc en rafale quand on tient la flèche) ;
- _load_selected_log(reset) lisait le fichier ENTIER (offset 0) et écrivait
TOUTES les lignes dans le RichLog, SYNCHRONE sur la boucle d'événements
-> un log de 250 Ko / plusieurs milliers de lignes gelait l'UI, × chaque
pas de navigation.
Corrections :
- _read_tail : au changement de VM on ne lit que la FIN du log (128 Ko /
1000 lignes max) et on cale l'offset sur la taille totale -> le suivi
incrémental continue depuis la fin. Affichage quasi instantané.
- Debounce (0,25 s) : RowHighlighted mémorise seulement la cible ; on ne
charge qu'une fois le curseur stabilisé, et uniquement la DERNIÈRE VM.
Validé headless : 2 flèches rapides -> 1 seul chargement (VM finale),
1000 lignes affichées au lieu de 8000.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
L'auto-refresh coupé par défaut était une mauvaise idée : (1) les lags
persistaient malgré tout ; (2) une fois coupé, on ne voyait plus
l'avancement (et l'activation ne suffisait pas). On revient au
comportement précédent : rafraîchissement automatique TOUJOURS actif
(logs 1s / table 2s / domstate 10s) — au moins on voit la progression.
Revert de 4f0c3cd. L'analyse du VRAI ralentissement (navigation / logs)
est traitée séparément.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
À l'agrandissement, on indique désormais l'espace libre de l'hôte et la
taille virtuelle MAX « soutenable » ≈ (taille réelle actuelle + libre
hôte) — affichée AVANT la saisie pour guider le choix. Si la cible la
dépasse, avertissement NON bloquant : le qcow2 est creux, donc OK tant que
la VM ne remplit pas, mais au-delà l'hôte tomberait à court d'espace.
Ex. : réel 115G + libre 30G -> max soutenable ~145G ; viser 190G affiche
« dépasse de ~45G — surallocation ». L'opération n'est pas bloquée.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sur un hôte en français, « virsh domstate » renvoie « en cours d'exécution »
au lieu de « running ». Le test « state == "running" » échouait donc :
l'agrandissement à chaud tombait dans la branche « qemu-img resize » (au
lieu de « virsh blockresize ») sur une VM ALLUMÉE -> « Failed to get write
lock ».
On force désormais LC_ALL=C / LANG=C (via _qemu_c_env) sur tous les outils
dont on PARSE la sortie, pour avoir l'anglais quelle que soit la locale :
virsh domstate/domname/dominfo/domblklist/list, sgdisk -i, dumpe2fs,
resize2fs -P (todo.py) et virsh list --all (dashboard). Cela répare aussi,
en locale fr, la détection d'état (pause/éteinte), l'analyse des infos
avancées (CPU/RAM) et le parsing de la réduction (partition/GPT).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Trois améliorations autour de la sauvegarde de disque à la réduction :
- La sauvegarde .bak n'est plus systématique : on DEMANDE (défaut OUI)
« Sauvegarder le disque avant réduction ? ». Si non, avertissement (un
échec pourrait casser le disque, pas de restauration possible).
- À la FIN, après avoir proposé de démarrer la VM (donc après test manuel
possible), on propose d'EFFACER la sauvegarde (défaut NON -> on la garde
par prudence).
- « Nettoyer QEMU » détecte désormais les sauvegardes *.qcow2.bak (libellé
« sauvegarde de disque (redim.) ») et permet de les effacer.
_qemu_shrink_revert gère l'absence de sauvegarde (avertit de lancer fsck
au lieu de restaurer). Validé de bout en bout sur image jetable.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
« Partition à réduire introuvable » sur une vraie VM : _qemu_root_part
parsait lsblk en positionnel et EXIGEAIT 4 colonnes. Juste après le connect
nbd, le FSTYPE n'est pas encore en cache -> colonne VIDE -> 3 tokens ->
toutes les partitions étaient ignorées. (Le test initial passait car le FS
avait eu le temps d'être détecté.)
- lsblk -P (paires clé="valeur") : robuste aux colonnes vides ; on repère
la partition par TYPE="part" et on sonde le FSTYPE via blkid si absent.
- _qemu_nbd_connect attend l'APPARITION des sous-périphériques nbdNpM
(jusqu'à ~15 s) avant de rendre la main.
- partprobe silencieux (capture) -> plus de spam « Invalid argument during
seek » pendant la réparation GPT.
Validé sur image jetable au layout cloud (p1 root + p14 bios + p15 ESP),
détection immédiate après connect : 25G -> 15G, GPT « No problems found »,
partitions préservées, fsck propre.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
erplibre-ubuntu-2004 (déployée à 10G) plantait à l'install : « Poetry
installation error status 1 » car le disque était plein (9.7G/10G). Un
ERPLibre + Odoo occupe ~11G, plus le transitoire (caches pip, sync repo,
sources Odoo). Preuve : la VM à 15G réussissait, celle à 10G non.
Plancher relevé à 20G pour TOUTES les versions (ubuntu 20.04/22.04,
debian, fedora, arch). Le qcow2 est CREUX (sparse) : une taille virtuelle
plus grande ne consomme rien tant qu'elle n'est pas remplie -> aucun
surcoût réel, marge confortable pour l'installation.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
La réduction « sûre » via virt-resize ne marchait pas : (1) j'utilisais
« virt-filesystems -b » (option INEXISTANTE -> aucune partition détectée ->
« abandon »), et (2) libguestfs est inutilisable ici (aucun vmlinuz dans
/boot -> l'appliance supermin échoue). La VM restait donc à 25G.
Réécriture SANS libguestfs, avec des outils de base présents et éprouvés
(qemu-nbd, e2fsck, resize2fs, sgdisk, parted/partprobe) :
1. copie .bak AVANT toute modification ;
2. nbd + détection de la racine (plus grosse partition, ext seulement) ;
3. e2fsck -> resize2fs (FS) -> sgdisk réécrit la partition en PRÉSERVANT
type/UUID/nom (PARTUUID intact) -> qemu-img --shrink (conteneur) ->
sgdisk -e (GPT de secours) -> fsck final ;
4. en cas d'échec à N'IMPORTE quelle étape : restauration depuis .bak
-> corruption impossible.
Validé de bout en bout sur une image jetable (GPT + ext4, racine à offset
élevé, 800M de données) : 25G -> 15G, GPT « No problems found », fsck
propre, UUID préservé, données intactes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Si le dashboard se ferme (bug, sortie), on peut désormais le ROUVRIR sur un
run passé pour reprendre l'analyse. Nouvelle entrée [11] du menu QEMU :
« 📈 Rouvrir le suivi d'installation (dernier run / historique) ».
- mon.list_install_runs() : liste les runs (~/.erplibre/qemu-install/*/
session.json) triés du plus récent au plus ancien.
- _qemu_reopen_monitor : affiche l'historique (VM par run), choix (défaut =
le dernier), puis run_monitor(session.json) rouvre le dashboard.
- « Lister les images » passe de [11] à [12] ; config -> 13+.
Validé : 20 runs listés, dashboard reconstruit (titre + 5 VM) depuis un
session.json passé.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sur un gros parc, les ticks périodiques (logs 1s / table 2s / domstate 10s)
causaient du lag à la lecture des logs. Nouveau bouton « a » pour couper/
activer le rafraîchissement automatique, COUPÉ par défaut ; « r » pour
rafraîchir une seule fois à la demande.
- Chaque tick est scindé en wrapper _tick_* (ne fait rien si auto coupé) +
corps _do_* (le travail réel), réutilisable pour les rafraîchissements
forcés (init à l'ouverture, toggle, refresh manuel, après pause/reprise).
- Un relevé initial unique (table + domstate) à l'ouverture affiche l'état
courant même auto coupé ; changer de VM recharge son log (synchrone).
- Indicateur #autobar : « ⏸ Rafraîchissement auto : COUPÉ (a=activer ·
r=rafraîchir) » / « ▶ … : ACTIVÉ (a=couper) ».
Validé headless : défaut OFF, toggle a, refresh r, table peuplée à l'init.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
« Suivi interactif (dashboard) ? » répondait NON par défaut (réponse vide).
Le défaut est désormais OUI : nouveau helper _is_yes_default_yes (vide = oui)
et libellé « (O/n, défaut : oui) » / « (Y/n, default: yes) ».
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
virt-resize/virt-filesystems (requis pour la réduction SÛRE de disque) sont
des binaires SYSTÈME (OCaml), pas des paquets Python : ils ne peuvent pas
vivre dans .venv.erplibre. On les ajoute donc au jeu de paquets système du
profil « ERPLibre Déploiement » (à côté de qemu/libvirt/virtinst) :
apt libguestfs-tools · dnf guestfs-tools · pacman libguestfs.
Ainsi tout hôte provisionné avec ce profil peut réduire un disque sans
risque. (La commande de redimensionnement propose déjà d'installer
libguestfs à la demande si absent.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Réduire avec « qemu-img resize --shrink » tronquait le conteneur qcow2 SANS
réduire le FS/partition/GPT invités : partition racine tronquée + GPT de
secours perdue -> dracut-initqueue en échec, OS non bootable.
Désormais la réduction passe par virt-resize (libguestfs) qui réduit
proprement système de fichiers + partition + GPT :
- _qemu_safe_shrink : écrit dans une NOUVELLE image (qemu-img create +
virt-resize --shrink <partition>), et ne remplace l'originale (mv) QUE si
virt-resize réussit. En cas d'échec/refus, le disque d'origine reste
INTACT (impossible de corrompre). Sauvegarde conservée en .bak.
- _qemu_largest_partition : détecte la partition racine (la plus grosse) via
virt-filesystems.
- _qemu_install_libguestfs : propose d'installer libguestfs-tools si absent ;
sinon on ABANDONNE (on ne tronque jamais).
- _qemu_offer_start extrait (redémarrage après extinction).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sur erplibre-ubuntu-2604 (Ubuntu 26.04), wkhtmltopdf ne s'installait pas :
la liste de versions s'arrêtait à 25.10, donc WKHTMLTOX_X64 restait VIDE
-> « sudo gdebi $(basename '') » -> « Usage: gdebi… » -> erreur -> Odoo :
« You need Wkhtmltopdf to print a pdf ».
- Bloc Ubuntu réécrit : 18.04->bionic, 20.04->focal, ELSE->jammy (build le
plus récent publié par wkhtmltopdf, valable 22.04..26.04 et au-delà).
Un « else » évite toute URL vide sur une future version.
- Garde-fou : si l'URL est vide malgré tout, on saute proprement
(wkhtmltopdf est optionnel) au lieu d'appeler gdebi sans fichier.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sur erplibre-ubuntu-2004, l'installation ERPLibre plantait : pyenv ne
pouvait pas compiler Python 3.12.10 (« no acceptable C compiler found »),
car build-essential n'était jamais installé.
Cause : shfmt figurait dans le MÊME « apt-get install » que build-essential
(git, cmake, …). shfmt est absent des dépôts Ubuntu < 22.04 -> apt refuse
le lot ENTIER (« Unable to locate package shfmt ») -> exit 1 avant même les
dépendances de compilation de pyenv.
Fix : shfmt (simple formateur shell, non requis pour exécuter ERPLibre) est
retiré du lot critique et installé SÉPARÉMENT en best-effort (jamais fatal).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Après un shrink, « Démarrer la VM » faisait « virsh start <id> » : une fois
la VM éteinte, l'ID numérique disparaît -> « failed to get domain 32 ».
On résout désormais le NOM canonique dès le début de _qemu_resize_disk
(VM encore allumée, ID résoluble) et on l'utilise partout, y compris pour
le redémarrage final. Plus de dépendance à l'ID une fois la VM arrêtée.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le choix du navigateur propose désormais « [i] Installer un autre
navigateur » en plus des navigateurs déjà installés. L'option ouvre le
sous-menu d'installation (choix w3m/lynx/links/elinks, commande adaptée à
l'OS, validation), même s'il existe déjà un navigateur.
Le flux d'installation est extrait dans _qemu_install_cli_browser
(réutilisé quand aucun navigateur n'est présent ET via l'option [i]).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le navigateur ne faisait qu'imprimer sans réagir au clavier : la commande
passait par exec_command_live, qui exécute avec stdout=PIPE (et sans stdin
terminal) -> un navigateur texte (w3m/elinks) n'a pas de TTY interactif.
On lance désormais le navigateur avec os.system(), qui hérite du vrai
terminal (stdin/stdout/stderr) — même principe que le suivi d'installation
(TUI) qui l'appelle dans self.suspend(). Ici, en CLI simple, aucun suspend
n'est nécessaire.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Nouvelle entrée [10] dans « Gérer » du menu QEMU/KVM : 🧪 Tester une VM.
Elle liste les VM, demande laquelle, résout son IP puis ouvre
http://IP:8069 (Odoo) dans un navigateur web EN LIGNE DE COMMANDE choisi
par l'utilisateur.
- _qemu_test_vm : résolution d'IP (_qemu_vm_ip), lancement du navigateur,
message d'aide si la page ne s'affiche pas (Odoo pas démarré / réseau).
- _qemu_choose_cli_browser : liste les navigateurs CLI installés et laisse
choisir ; si aucun, propose d'en installer un (réutilise CLI_BROWSERS /
INSTALLABLE_BROWSERS / browser_install_command de qemu_install_monitor).
- « Lister les images » passe de [10] à [11] ; les entrées de config
suivent à 12+ (branche else inchangée via la liste `real`).
Validé : rendu du menu, choix du navigateur ([2]->elinks, vide->w3m).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
La colonne ⚠ donne le NOMBRE ; on peut désormais LIRE les lignes associées.
Touche « d » (ou indice dans la barre) : ouvre une petite fenêtre modale
(ErrorLinesScreen) listant, pour la VM sélectionnée, les lignes d'erreurs
puis d'avertissements (numéro de ligne + texte), défilable. Échap/q ferme.
- scan_log_error_lines() : mêmes détection + listes d'ignore que la suite de
tests ERPLibre, mais retient les lignes (bornées à 500) avec leur numéro.
- Scan à la demande à l'ouverture -> toujours à jour (marche même pendant
l'installation).
Validé headless : modale s'ouvre avec les lignes, se ferme à Échap.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sur les images cloud, /var/lib/apt/lists est vide (package_update désactivé)
-> « Unable to locate package qemu-guest-agent » et agent inactif. On
rafraîchit l'index (timeout 120, || true) AVANT l'install dans le runcmd
(après sshd, sans bloquer SSH). Idem pacman -Sy. Sous-shell ( ) et non
accolades { } (indicateur YAML).
Validé sur une VM de test Ubuntu 24.04 (redéployée) :
- cloud-init status: done ; user erplibre créé + mot de passe (login console
erplibre/erplibre OK) ;
- qemu-guest-agent: active ;
- guest-ping OK ; guest-exec autorisé (echo -> code 0, tourne en root).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
La ligne runcmd Fedora commençait par « - [ -f … ] » : en YAML, « - [ »
démarre une SÉQUENCE EN FLUX, donc « [ -f … ] && sed … » rendait tout le
user-data invalide. cloud-init rejetait alors la config ENTIÈRE : aucun
utilisateur créé, pas de SSH, login console erplibre/erplibre impossible,
VM « en attente de démarrage ».
- « test -f … && … » au lieu de « [ -f … ] && … » (ne démarre pas de
séquence YAML).
- timeout 300 sur l'installation de qemu-guest-agent : un miroir lent ne
bloque plus cloud-init.
- Validé : le user-data se parse en YAML (tous les runcmd sont des chaînes),
avec et sans mot de passe.
Les VM déjà déployées avec le mauvais seed doivent être redéployées ; les
prochaines repartent correctement.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Suivi d'installation (qemu_install_monitor.py) :
- Table VM/État LISIBLE : largeurs de colonnes fixes (VM 26, ⚠ 4, État 12,
Durée 7, Disque 8) -> l'État n'est plus tronqué à 3-4 caractères
(« ❌ effacée », « ⏸ en pause » lisibles) ; table défilable (overflow-x,
height 1fr) pour les longs noms et les gros parcs.
- « w » web : offre désormais la LISTE des navigateurs CLI installés
(_choose_browser) pour choisir lequel utiliser, + option [i] installer.
Déploiement (deploy_qemu.py) :
- guest-exec AUTORISÉ : on vide la liste de blocage de qemu-ga
(block-rpcs/blacklist vides — pas allow-rpcs qui est une liste BLANCHE et
casserait les autres RPC), + neutralise /etc/sysconfig/qemu-ga (Fedora).
Installation (todo.py) :
- Profils AVEC Odoo (install_odoo*) UNIQUEMENT : Odoo est enregistré comme
service systemd (erplibre.service, inspiré de script/systemd/
install_daemon.sh) puis enable --now. Pas pour ERPLibre seul / mobile /
Déploiement.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Provisioning (deploy_qemu.py) :
- qemu-guest-agent installé + activé dans le runcmd cloud-init (APRÈS
sshd, « || true » : ne bloque pas le boot si le réseau est lent).
- Canal virtio org.qemu.guest_agent.0 ajouté à virt-install : virsh peut
piloter la VM SANS réseau.
- Extension du FS invité (todo.py) : nouveau tier AGENT INVITÉ
(_qemu_guest_exec via qemu-agent-command guest-exec) entre SSH et la
console série -> étend le FS même sans IP.
Suivi d'installation (qemu_install_monitor.py) :
- Détection d'erreurs dans le log à la complétion (succès OU échec) en
réutilisant la logique de script/test/run_parallel_test.py (sous-chaîne
error/warning + listes d'ignore). Nouvelle colonne « ⚠ » À GAUCHE d'État
(⚠N erreurs / ⚡N avert. / ✓ propre).
- Sommaire de stats EN CHIFFRES (📊 total · ✅ · ❌ · ⏳ · ⏸ · 🗑 · ⚠ · ⚡) ;
CLIC pour déplier le détail (VM en erreur + durées).
- Boutons « p » Pause tout (virsh suspend des VM running) et « o »
Reprendre tout (virsh resume) — les logs continuent (offsets conservés).
Validé headless : succès-avec-erreur -> ⚠1, stats/détail, pause du parc.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Trois améliorations QEMU :
- Extension du FS invité robuste : résolution d'IP avec BATTEMENT
(_qemu_resolve_ips, parallèle, boot émulé lent) au lieu d'un timeout
court ; en cas d'absence d'IP ou d'échec SSH, repli sur la CONSOLE
SÉRIE avec la commande growpart/resize prête à coller (login
erplibre/erplibre). Commande factorisée dans _GROW_FS_REMOTE.
- Temps estimés PAR VERSION : nouveau mon.avg_by_version(distro, version)
et _qemu_stat_avg("version", v, distro) -> suffixe « · ~46s moy (1) »
dans les listes de versions (prompt simple + granulaire).
- i18n : « Versions for » n'avait AUCUNE traduction (restait en anglais).
Ajout FR « Versions des » (+ « Version des ») et capitalisation de la
distro -> « Versions des Ubuntu : ».
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Deux points :
- Fil d'Ariane depuis la télémétrie : une commande lancée DEPUIS le TUI de
télémétrie ne passait par aucun menu, donc son chemin n'était jamais
affiché. On imprime désormais « 📍 TODO › … › <commande> » (dernier
segment traduit + icône) avant l'exécution, et on enregistre le chemin.
- Arrêt de VM (réduction disque) par SIGNAL : virsh shutdown --mode
acpi,agent (bouton ACPI puis agent invité) au lieu d'un arrêt implicite.
Pendant l'attente, on affiche un compte à rebours du timeout
(« ⏳ arrêt en cours… NNN s restantes ») et le délai max au départ.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- F3 fait maintenant défiler Arbre -> Kanban -> Liste -> Arbre. Le résumé
et le footer annoncent la PROCHAINE vue (« F3 → Liste »).
- Nouvelle vue Liste : tous les menus empilés verticalement, chacun avec
ses sections et ses commandes exécutables.
- Les colonnes Kanban et la vue Liste affichent les SECTIONS (── … ──)
pour guider le choix, et chaque commande porte son ICÔNE.
L'icône vient de t() : l'AST extrait la clé i18n (prompt_description /
prompt_description_key), et _disp() la résout en libellé traduit + icône.
Les sections sont capturées via _choice_entries (marqueurs {"section": …}).
Validé headless : cycle F3, 11 en-têtes de menu, sections en Liste/Kanban.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Après « virsh list --all », le menu propose désormais :
[1] Infos avancées (vCPU, RAM, disque)
[2] Changer l'état d'une ou plusieurs VM
[Entrée] Rien
Changement d'état (_qemu_change_state) :
- Saisie d'une liste de VM séparée par des virgules (noms ou ID, résolus
via _qemu_domname et validés contre les VM existantes).
- Choix de l'état cible : Ouvrir (start) ou Fermer (shutdown).
- DOUBLE validation (« Appliquer : … ? » puis « Confirmer pour de vrai ? »)
avant d'exécuter virsh start/shutdown sur chaque VM.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Si la VM a dû être éteinte pour la réduction, on le NOTE (« La VM a été
éteinte pour le redimensionnement. ») puis on demande (o/N) si on veut la
redémarrer (virsh start sur le nom canonique). Si la VM était déjà éteinte
ou n'a pas eu besoin de l'être (agrandissement à chaud), on ne demande
rien : drapeau was_shut_down positionné uniquement après un arrêt effectif.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Avant, réduire le disque d'une VM allumée affichait seulement « Éteignez
la VM » puis abandonnait. Désormais on DEMANDE (o/N) si on veut l'éteindre
et réessayer :
- _qemu_shutdown_wait : arrêt ACPI gracieux (virsh shutdown), attente
jusqu'à « shut off » (timeout 120s), puis propose un arrêt forcé
(virsh destroy) si l'arrêt traîne.
- _qemu_domname : résout un ID numérique en nom canonique — l'ID
disparaît une fois la VM éteinte, le polling doit utiliser le nom.
Validé : domname(14) -> erplibre-ubuntu-2004.
- Tous les nouveaux prompts affichent la valeur par défaut (o/N).
Une fois éteinte, la réduction (qemu-img resize --shrink) se poursuit
automatiquement après confirmation du risque.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Après « sudo virsh list --all », le menu propose désormais (o/N) un
tableau détaillé par VM : état, vCPU, RAM allouée (Max memory), taille
disque virtuelle et réelle (qemu-img info -U, lit même VM allumée), plus
l'espace total/libre/utilisé du stockage des images (shutil.disk_usage).
- _qemu_list_vms(ask_advanced=False) : seul le menu [4] prompte ;
les appels internes (IP, console, resize, delete) restent inchangés.
- Helpers _qemu_dominfo (vcpu + Max memory) et _qemu_disk_sizes
(virtuel + réel, -U). Validé en réel : 8 vCPU / 8.0G / 30.0G / 2.5G.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Icônes devant chaque entrée du menu Execute (Code 💻, Config 🔧, Run 🏃,
Test 🧪, Process 📟, Database 💾, Git 🌿, Mise à jour 🔃, Doc 📖, GPT code
🤖, Automatisation 🦾, Déploiement 🚀, Réseau 📡, Sécurité 🔒, Langue 🌍).
- Icônes sur les titres de section (Développement 🧰, Données 📊, Sources &
documentation 📚, IA & automatisation 🧠, Déploiement/réseau/sécurité 🌐,
Préférences 🎨).
- « Back » (Retour) reçoit 🔙 — icône « Précédent » partagée par tous les
sous-menus via fill_help_info.
Emojis larges (2 cellules) -> 1 espace, alignement de la colonne label
préservé. Modifs uniquement dans les VALEURS de todo_i18n.py.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
qemu-img info échouait (« Failed to get shared write lock ») sur une VM
running car libvirt tient le lock d'écriture : la taille virtuelle lue
tombait à 0.0 G, et « -10G » donnait alors 0-10 = -10 → « Taille invalide ».
- Ajout de -U (--force-share) aux deux appels qemu-img info (affichage +
_qemu_disk_virtual_bytes) : lecture seule sûre même VM allumée.
- Garde-fou : si la taille reste illisible (0), on abandonne avec un
message clair au lieu de calculer une cible négative.
Le redimensionnement lui-même était déjà correct (virsh blockresize à
chaud si running, qemu-img resize si éteinte) — pas besoin d'éteindre la
VM pour AGRANDIR. Validé : lecture -U renvoie 30.0 G sur une VM running.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Deux problèmes du bouton « w » (ouvrir l'UI web de la VM) :
- L'installation prenait w3m sans demander. Désormais on CHOISIT le
navigateur CLI à installer (w3m / lynx / links / elinks) via _install_cli_
browser ; la commande (adaptée à l'OS) est affichée puis validée.
browser_install_command prend un paramètre `browser`.
- Au lancement, le navigateur « clignotait » et revenait au TUI sans qu'on
voie l'erreur (souvent Odoo pas démarré sur :8069). On affiche maintenant
la commande, le CODE DE SORTIE, un indice (Odoo/réseau) et une PAUSE
« Entrée pour revenir » pour diagnostiquer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
L'historique d'installation est désormais enregistré dans un fichier DÉDIÉ
.venv.erplibre/qemu_install_stats.json (repli ~/.erplibre) : chaque run garde
distro + version + architecture + durée + horodatage (500 derniers).
- record_duration(distro, version, arch, secs) enregistre le run (appelé par
le dashboard à la complétion d'une VM).
- Menus de sélection enrichis : chaque architecture et chaque distribution
affiche la DURÉE MOYENNE d'install historique « · ~5m moy (3) » quand la
donnée existe. La DERNIÈRE install (distro version [arch] — durée) est
rappelée en tête du déploiement.
- eta_reference lit désormais les runs (médiane par arch, repli global).
Nouveaux helpers : avg_by_arch, avg_by_distro, last_run.
Validé : enregistrement + moyennes (amd64 ~5m (2), ubuntu ~13m (3)),
dernière install affichée, distros sans données masquées.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Quand on installe ERPLibre sur une/des VM, on choisit désormais CE qu'on
installe (au lieu de forcer install_odoo_18) :
ERPLibre + Odoo 18/17/16/15/14/13/12 (make install_os && make install_odoo_X)
ERPLibre + toutes les versions Odoo (make install_odoo_all_version)
ERPLibre seulement (sans Odoo) (./script/install/install_erplibre.sh)
ERPLibre mobile (home) (./mobile/install_and_run.sh)
ERPLibre Déploiement (+ QEMU + dev) (install_dev + qemu/libvirt/virtinst)
_qemu_pick_install_profile renvoie la commande finale ; _qemu_erplibre_remote_cmd
l'exécute dans ~/git/erplibre ; le profil est threadé aux installeurs
(monitoré + streamé). Cibles make vérifiées dans les Makefiles.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Après « i » (installer lm-sensors) dans la vue système F2, l'affichage
n'était pas garanti d'être remis à jour : le message « lm-sensors absent »
pouvait rester même si les capteurs devenaient lisibles.
Fix : à la fin de action_install_sensors, on ré-échantillonne le système
(température incluse, _sys_prev remis à zéro) et on réécrit la case
SYNCHRONEMENT + self.refresh(), puis on notifie le résultat (« capteurs
désormais disponibles » ou « toujours pas de température — reboot/modprobe
requis ? »).
Validé headless : la ligne Température passe de « lm-sensors absent… » à
« 48°C (max) » une fois la température lisible.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ajoute une icône descriptive devant chaque section et entrée du menu QEMU :
🚀 Déploiement / Déployer, 🔍 Prévisualiser, ⬇ Télécharger, 🛠 Gérer,
📋 Lister, 🌐 IP, 🖥 Console, 📐 Redimensionner, 🗑 Effacer, 🧹 Nettoyer,
📚 Catalogue, 🗂 Lister images.
Uniquement les VALEURS i18n (fr+en) changent -> aucune modif de todo.py, les
clés (et donc les appels t(...) + la télémétrie) restent inchangés.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- F2 : nouvelle vue SYSTÈME (état/uptime + charge, CPU %, mémoire, disque,
réseau ↓/↑, batterie, température). Les I/O sont déportées en thread
(asyncio.to_thread) ; rafraîchissement 2 s (deltas CPU/réseau). Si la
température n'est pas lisible (ni /sys/class/thermal ni lm-sensors), on
propose d'installer lm-sensors avec la touche « i » (commande selon l'OS —
apt/dnf/pacman — affichée puis validée). read_temperature essaie
/sys/class/thermal d'abord (sans dépendance).
- Vue Kanban GRILLE : clic sur le TITRE d'une case -> elle s'AGRANDIT
(row-span, prend toute la hauteur d'une colonne) ; re-clic -> taille
normale (classe kbig).
Validé headless : vue système peuplée (mém/disque/temp), F2/F3 basculent,
clic titre en grille -> kbig on/off.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Après exécution d'une commande choisie dans la télémétrie, on propose de
REVENIR (r) ou de quitter. En revenant, la vue ET la position du curseur
sont RESTAURÉES (chemin porté par chaque nœud/carte ; run_tui prend/rend un
`state`).
- Vue Kanban : F4 fait défiler la DISPOSITION —
columns (une rangée de colonnes) / swimlanes (une rangée par menu de
niveau 1) / grid (grille 3 colonnes). F3 bascule Arbre/Kanban.
Validé headless : F3/F4 cyclent les dispositions, sélection -> action+état,
relance avec `state` -> curseur restauré sur le bon nœud.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Ticks ASYNCHRONES : les I/O bloquantes (lecture des logs, stat disque,
subprocess « virsh list », /proc) sont déportées en THREAD via
asyncio.to_thread ; seules les mises à jour d'UI restent sur la boucle
d'événements Textual -> plus de gel, même sous forte charge ou disque lent.
_collect_table (thread) rassemble statut+disque+télémétrie, l'application
au tableau se fait ensuite sur la boucle. _tick_log/_tick_domstate idem.
Helper _read_new (lecture incrémentale) réutilisé.
- Touche « w » : si aucun navigateur CLI n'est présent, on PROPOSE de
l'installer selon l'OS (apt Ubuntu/Debian, dnf Fedora, pacman Arch — nos 4
systèmes) : la commande est AFFICHÉE puis exécutée après validation (o/N).
Validé headless : ticks async mettent l'UI à jour ; browser_install_command
renvoie la bonne commande selon l'OS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le TUI de télémétrie devient aussi un LANCEUR, avec deux vues :
- Sélectionner une COMMANDE (feuille de l'arbre, ou carte du Kanban) +
Entrée -> le TUI se ferme et la commande est EXÉCUTÉE (getattr(self, méthode)
(**kwargs)). Les kwargs littéraux sont extraits du code (ex. Aperçu ->
_qemu_deploy(dry_run=True)), donc la commande est rejouée à l'identique.
- Vue KANBAN : une colonne par menu contenant des commandes (issu du code),
cartes = commandes exécutables + compteur de visites du menu.
- F3 bascule entre vue Arbre et vue Kanban. Résumé mis à jour (Entrée =
exécuter, F3 = vue).
_dispatch capture désormais (méthode, kwargs) ; les feuilles de l'arbre
portent la méthode en data ; run_tui renvoie (méthode, kwargs) à exécuter.
Validé headless : 11 colonnes Kanban, kwargs (dry_run/production_ready)
capturés, bascule F3, capture de l'action à exécuter.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
La télémétrie n'affichait que les chemins effectivement visités. Elle
observe désormais les MÉTADONNÉES DE NAVIGATION issues du CODE : todo.py est
analysé (AST) pour reconstruire l'arbre RÉEL des menus et commandes.
- build_code_tree() lit _MENU_LABELS + les listes « choices » + le dispatch
« status == N: self.X() » de chaque menu pour bâtir l'arborescence complète
(sous-menus = cibles présentes dans _MENU_LABELS ; sinon commandes-feuilles,
libellées par leur prompt_description).
- Le TUI affiche cet arbre COMPLET (tous les menus, même jamais visités) avec
le compteur de visites en surimpression sur chaque menu (via la télémétrie
persistée). Repli sur l'arbre des seuls chemins visités si l'analyse échoue.
Validé : l'arbre reconstruit couvre Execute -> Deploy -> SSH/QEMU (+ toutes
leurs commandes : Resize, Delete, …), Git -> Git local server, GPT code ->
RTK, etc. ; compteurs corrects (menus visités > 0, autres 0).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Nouvelle entrée « [5] Télémétrie de navigation (TUI) » dans le menu principal.
- Enregistrement : Todo._menu_header (seul constructeur du fil d'Ariane)
appelle todo_telemetry.record(fil) à chaque affichage de menu. Dédup des
ré-affichages consécutifs -> on ne compte que les TRANSITIONS de
navigation. Persistant dans ~/.erplibre/todo_telemetry.json. Best-effort :
ne casse jamais la navigation.
- Visualisation : todo_telemetry.run_tui() ouvre un TUI Textual affichant
l'ARBRE des fonctionnalités visitées (chaque nœud = un menu, avec son
nombre de visites), trié par usage décroissant. Touches : q (quitter),
r (réinitialiser), e (tout déplier). Sommaire : total navigations + menus.
Validé headless : enregistrement + dédup, arbre imbriqué à compteurs,
montage du TUI, peuplement, touches.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pendant le suivi, une VM pouvait « stagner » sans qu'on sache pourquoi
(mise en pause, ou effacée à côté). Le dashboard interroge désormais l'état
libvirt (« virsh list --all ») à INTERVALLE LENT (toutes les 10 s, un seul
appel pour tout le parc ; le tableau applique le cache à chaque tick de 2 s) :
- VM en pause (virsh suspend) -> État « ⏸ pause » (non terminal, peut
reprendre).
- VM absente de virsh (effacée pendant l'attente) -> État « ❌ effacée »,
terminal (on cesse de lire son log).
Sinon on garde le statut basé sur le log (⏳ / ✅ / ❌).
Validé headless : running -> ⏳, paused -> ⏸ pause, absente -> ❌ effacée.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Nouvelle entrée dans le menu QEMU/KVM (section Gérer). Elle :
- affiche le disque principal (qcow2 via domblklist) + « qemu-img info »
(taille virtuelle + réelle) et l'état de la VM ;
- demande +NG (agrandir), -NG (réduire) ou NG (taille cible) ;
- applique : agrandissement À CHAUD si la VM tourne (virsh blockresize),
sinon qemu-img resize ; réduction via qemu-img resize --shrink (VM éteinte
obligatoire + avertissement fort : le FS invité n'est PAS réduit, risque
de perte de données) ;
- propose ensuite d'étendre le FS invité via SSH (growpart + resize2fs /
xfs_growfs / btrfs, device et type de FS détectés).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ajouts au suivi d'installation (dashboard Textual) :
- Télémétrie hôte (barre dédiée, MAJ 2 s) : CPU % = charge / nb CPU, et
disque du dossier des images (utilisé/total/libre) — utile pour anticiper
le remplissage du disque (cas vécu).
- Colonne « Disque » par VM : taille RÉELLE du qcow2 (st_blocks) -> on voit
quelle VM grossit.
- ETA « hypothèse » : les durées d'install sont MÉMORISÉES par architecture
dans ~/.erplibre/qemu-install/stats.json ; l'ETA du parc = médiane
historique (par arch, repli global) - temps écoulé, affichée au sous-titre.
- Touche « w » : ouvre l'UI web de la VM (Odoo :8069) dans un navigateur CLI
(browsh/carbonyl/w3m/links/elinks/lynx, le 1er trouvé ; sinon notifie quoi
installer). Surtout utile quand l'install est terminée.
Helpers module : load_stats/record_duration/eta_reference, _fmt_size/_fmt_secs,
vm_disk_path/disk_actual_size, cli_browser. run_monitor(run_app=False) expose
l'app pour test headless.
Validé headless (run_test) : montage, ticks, enregistrement des durées, ETA,
télémétrie, touches f/w — sans crash.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Encore du lag à 16 VM : update_cell était appelé pour CHAQUE ligne à chaque
tick (état + « Durée »), or la Durée — dérivée d'un `started` global — est
identique sur toutes les lignes et se réécrivait sans cesse -> re-render
permanent de la DataTable.
- _set_cell : n'appelle update_cell (donc ne re-render) que si la valeur
CHANGE. Mesuré : 312 -> 18 update_cell sur 10 ticks × 16 VM.
- La durée « live » passe dans le SOUS-TITRE (une seule mise à jour) ; la
colonne Durée n'est écrite qu'à la complétion (durée finale figée).
- Le NOMBRE de VM est affiché : titre « … (N VM) », sous-titre
« k/N terminées · mm:ss ».
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le TUI de suivi laggait à 30 VM et se figeait quand l'I/O ralentissait
(disque plein / logs effacés). Cause : à chaque tick (1 s) on lisait le
fichier log ENTIER de CHAQUE VM pour le statut + on relisait tout le log
sélectionné, en synchrone sur la boucle d'événements Textual.
Corrections :
- read_status ne lit que les 4 derniers Ko (le marqueur de sortie est sur
la dernière ligne) : un log de 2 Mo passe de « tout lire » à 0,1 ms.
- _load_selected_log lit de façon INCRÉMENTALE (seek à l'offset) au lieu de
relire tout le fichier.
- Les VM TERMINÉES ne sont plus relues (cache _final).
- _tick_table (2 s) et _tick_log (1 s) séparés et PROTÉGÉS par try/except :
une erreur I/O transitoire (disque plein, log supprimé) ne tue plus la
boucle -> l'interface reste réactive aux touches.
- RichLog borné (max_lines=5000).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sur un gros parc (30 VM, ~18 émulées), la résolution d'IP « bloquait » 10 min
en silence sur les VM lentes. Trois causes traitées :
- Joignabilité par PING (ICMP) d'abord, TCP:22 en repli. Le ping répond dès
que le réseau de la VM est up, BIEN AVANT sshd : on ne retenait pas une VM
qui a déjà son IP juste parce que sshd (lent en émulation) n'était pas prêt.
Le ping distingue toujours le bail actif (répond) du bail périmé (non).
- IP cherchée sur PLUSIEURS sources : lease (dnsmasq), agent (qemu-guest-
agent DANS la VM) et arp (table ARP hôte). Le bail dnsmasq peut être vide
sous forte charge alors que la VM a une IP (constaté sur ce parc).
- BATTEMENT toutes les 30 s listant les VM encore en attente + timeout PAR VM
ramené à 5 min (au lieu de 10) dans la phase de résolution -> plus de
silence prolongé, et on n'attend pas indéfiniment une VM sans IP.
NB : déployer la matrice complète (30 VM dont ~18 émulées TCG) sature un seul
hôte ; certaines VM émulées n'obtiennent pas de bail DHCP à temps (limite de
capacité, pas un bug). La résolution dégrade désormais proprement.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Améliore le suivi du déploiement multi-VM :
- Compteur de COMPLÉTION devant chaque résultat (« [1/26] », « [2/26] »… =
ordre où les tâches TERMINENT), en plus de l'ID de préparation stable
(« [7/26] ») — utile car l'exécution parallèle rend les résultats
désordonnés.
- Temps d'exécution PAR ÉTAPE : durée de chaque VM (déploiement, résolution
d'IP) + bilan de phase (« Bilan déploiement : 24 OK, 2 en échec, 26 VM,
3m10s » ; « IP résolues : 25/26 (2m). »).
- SOMMAIRE TOTAL encadré en fin de flux (VM déployées, total avec
existantes, temps total).
Helper _fmt_dur (« 45s » / « 2m05s »).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Chaque question oui/non affiche désormais explicitement la valeur par défaut
appliquée si on laisse vide, ex. :
« Suivi interactif (dashboard) ? (o/N, défaut : non) : »
Les 18 prompts passent par _is_yes(input(...)) : une réponse vide vaut NON.
On l'indique donc dans les traductions (fr « (o/N, défaut : non) », en
« (y/N, default: no) »). Seules les VALEURS i18n changent ; les clés (donc
les appels t(...) dans todo.py) sont inchangées.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sur un gros parc (ex. 26 VM), on ne voyait ni combien de VM allaient être
déployées ni lesquelles étaient en cours (résultats parallèles dans le
désordre). Deux améliorations :
- Le prompt de parallélisme affiche le NOMBRE de VM à déployer :
« Déploiements en parallèle (défaut : 24, 26 VM) : ». Pour cela, la
séparation à-créer / déjà-existantes est faite AVANT le prompt (on connaît
le vrai compte).
- Chaque VM reçoit un ID « k/N » suivant l'ordre de préparation, affiché
dans les résultats de déploiement (« ✅ [3/26] nom ») et dans la résolution
d'IP (« [5/30] nom : ip »). L'ID reste stable même si les résultats
reviennent dans le désordre -> on suit lesquelles sont traitées.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le menu QEMU avait deux commandes de déploiement redondantes : [1] « Déployer
une VM » (mono, noms/ressources personnalisés) et [4] « Déployer l'infra »
(multi, matrice distro×version×archi). Elles sont FUSIONNÉES en une seule,
[1] « Déployer une ou plusieurs VM » (_qemu_deploy), qui gère aussi bien 1 VM
que N, avec :
- renommage GRANULAIRE des VM à la demande (_qemu_customize_names) : noms
auto par défaut, on en renomme certains par numéros séparés de virgules ;
- multiplicateur de ressources (x1..x4) déjà présent ;
- aperçu dry-run ([2]) qui passe par le même flux et imprime les commandes
deploy_qemu (--dry-run) via le helper partagé _qemu_build_deploy_parts.
Menu simplifié : [1] Déployer une/plusieurs VM, [2] Aperçu (dry-run),
[3] Télécharger une image, puis Gérer/Catalogue renumérotés. Code mort retiré
(_qemu_deploy_vm, _qemu_prompt_arch, _normalize_disk_size,
_qemu_offer_ssh_config).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Quand le téléchargement échoue sur tous les miroirs, le message indique
désormais le CHEMIN de destination visé et une commande prête à copier pour
reprendre manuellement :
Destination : /var/lib/.../fedora-cloud-41-aarch64.qcow2
Reprendre le téléchargement manuellement :
sudo curl -fL -C - -o <dest> \
<url>
puis relancez le déploiement (l'image en cache sera réutilisée).
curl -C - reprend un .part partiel (ou repart de zéro), -f échoue proprement
sur une erreur HTTP. Complète la détection de téléchargement incomplet.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le menu Deploy listait 11 entrées « SSH - … » à plat, ce qui le surchargeait.
Elles sont déplacées dans un sous-menu « SSH (hôte distant)… »
(prompt_execute_deploy_ssh), comme le sous-menu QEMU. Le menu Deploy tient
désormais en 5 entrées :
── Local ──
[1] Cloner ERPLibre localement
[2] Configurer sshfs
── Distant & services ──
[3] SSH (hôte distant)… -> sous-menu (11 opérations SSH)
[4] NTFY
[5] QEMU/KVM…
Le fil d'Ariane affiche « TODO › Execute › Deploy › SSH » dans le sous-menu.
Les libellés et actions SSH sont inchangés (traductions réutilisées).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Avant le déploiement infra, on demande un multiplicateur de ressources :
x1 (minimum catalogue) .. x4. Il multiplie la RAM (base = minimum de la
version) et les vCPU (base 2), en bornant les vCPU au nombre de cœurs de
l'hôte et en signalant si la RAM totale dépasse la RAM libre.
- Le prompt affiche les ressources de l'hôte (vCPU, RAM libre) et, pour
chaque multiplicateur, le total RAM + vCPU/VM avec un ⚠ si ça dépasse.
- Le plan reflète les ressources effectives (vCPU + RAM×mult par VM).
- Les jobs passent --memory et --vcpus effectifs à deploy_qemu (avant,
l'infra laissait le minimum catalogue + 2 vCPU par défaut).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Après un déploiement infra, la résolution d'IP tournait EN SÉRIE (boucle
ssh_config puis install), jusqu'à 10 min par VM émulée, SANS aucune sortie :
le dashboard « n'ouvrait jamais » (impression de blocage), d'autant qu'une
VM qui ne boote pas (IP jamais attribuée) bloquait tout le reste.
Fix : _qemu_resolve_ips résout les IP de toutes les VM EN PARALLÈLE, avec
progression affichée par VM, une seule fois, réutilisée pour ~/.ssh/config
ET l'installation (plus de double résolution). _qemu_install_erplibre_*
acceptent une IP/ip_map déjà résolue.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Une VM arm64 (Fedora aarch64 émulée) restait FIGÉE au firmware (CPU ~2 %,
aucune IP, seed non appliqué). Deux causes :
1) Téléchargement tronqué non détecté. _download_one ne vérifiait pas que le
nombre d'octets reçus == Content-Length. Une connexion coupée en cours
laissait un .part TRONQUÉ, validé comme « complet » (tmp.replace) : un
qcow2 valide mais VIDE (128 Kio, juste l'en-tête, 5 Gio virtuels) ->
disque sans OS -> pas de boot. Pire, ce cache corrompu était réutilisé.
Fix : contrôle de complétude (done < total -> exception -> miroir suivant
ou échec net, jamais de commit d'un fichier tronqué).
2) arm64 : « --boot uefi » simple laissait libvirt choisir l'AAVMF « secure »
(clés Microsoft enrôlées) -> pas de boot. On désactive Secure Boot comme
pour x86 (secure-boot=no -> AAVMF_CODE.no-secboot.fd).
Validé : redéploiement Fedora 42 aarch64 avec image complète + firmware
non-SB -> CPU 111 % (boot actif), IP en ~80 s, SSH OK (erplibre, aarch64,
Fedora 42 Cloud, sudo OK).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Le log d'installation commence désormais par un en-tête identifiant
l'installation : date, VM, distribution + version, architecture, branche, IP.
Écrit dès la création du log (plus jamais vide au démarrage).
_qemu_install_erplibre_monitored déduit distro/version/arch de chaque VM
(nom via _qemu_infra_name + arch via virsh dumpxml) et les passe à
launch_installs, qui écrit l'en-tête et les stocke dans le manifeste.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
L'install ERPLibre échouait sur une VM Ubuntu 26.04 :
« configure: error: no acceptable C compiler found in $PATH » -> pyenv ne
compile pas Python -> pas de venv -> pip/poetry/repo absents.
Cause : install_dev.sh ne listait Ubuntu que jusqu'à 25.10. Sur 26.04 il
tombait dans le « else » (« Your version is not supported… : 26.04 ») et
n'exécutait donc PAS install_debian_dependency.sh -> ni build-essential ni
gcc installés. Le catalogue de déploiement propose pourtant 26.04.
Fix : ajout de 26.04 à la liste des versions Ubuntu supportées.
Note : Ubuntu 26.04 utilise uutils (rust-coreutils), bogué en big-endian ;
sur s390x un « sort » peut paniquer (non bloquant ici) — problème amont.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Menu « Déployer l'infra ERPLibre » : la sélection d'architecture propose
désormais [all] = toutes les architectures supportées. L'architecture
devient une dimension à part entière de la matrice de déploiement.
- [all] archis + [all] catalogue : une VM par (distro, version, archi
publiée par cette distro). Ex. ubuntu -> amd64/arm64/s390x, debian/fedora
-> amd64/arm64, arch -> amd64 (30 VM au total avec le catalogue courant).
- [all] archis + [granulaire] : la liste à plat inclut l'archi ([amd64]/
[arm64]/[s390x]) ; on choisit des combinaisons précises par virgules.
- [all] archis + [principal] : version par défaut de chaque distro × chaque
archi supportée.
- Une archi précise (arm64/s390x) restreint le catalogue aux distros qui la
publient (inchangé).
selected porte l'archi par élément ; les noms sont suffixés par archi pour
les non-natives (erplibre-ubuntu-2604-s390x) -> pas de collision. --arch
transmis par VM. Chaque distro ne reçoit que les archis QU'ELLE publie.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Symptôme : ~/.ssh/config (et l'install) pointaient sur .31 alors que la VM
était joignable sur .32 -> SSH KO (« No route to host »), console OK.
Cause : au boot, une image cloud demande d'abord une IP DHCP avec son
hostname par défaut (« ubuntu ») -> 1er bail ; puis cloud-init fixe le vrai
hostname et le client redemande -> 2e bail (IP différente). La MÊME MAC a
donc DEUX baux ; le 1er (« ubuntu ») devient périmé. Le code retenait
aveuglément le PREMIER IPv4 de « virsh domifaddr » = le bail périmé. Plus
visible sur s390x/arm64 émulé (boot lent -> les deux baux coexistent).
Fix :
- todo._qemu_vm_ip : renvoie en priorité l'IP dont le bail dnsmasq porte le
hostname == nom de la VM (bail définitif), sinon une IP JOIGNABLE (sshd up,
test TCP:22), sinon le dernier bail — jamais le 1er au hasard. Attente
portée à 10 min (boot émulé lent). Helpers _qemu_lease_candidates /
_qemu_ip_reachable / _qemu_lease_ip_for_host.
- deploy_qemu.wait_for_ip : même logique (IP joignable, sinon la plus
récente) pour l'IP affichée en fin de déploiement standalone.
Validé sur la VM s390x réelle : candidats [.31, .32] -> choisit .32
(hostname-match + seule joignable) ; ssh via alias OK.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1) Nommage : _qemu_infra_name ajoute un suffixe d'architecture quand elle
diffère de la native de l'hôte -> déployer un s390x sur un hôte amd64
donne « erplibre-ubuntu-2604-s390x » (au lieu de « erplibre-ubuntu-2604 »).
Évite les collisions de noms entre archis et rend l'archi visible. Passé
dans les 3 appelants (VM unique + plan + jobs de l'infra).
2) Log d'installation vide : sur une archi ÉMULÉE (s390x/arm64 sur x86) le
boot prend plusieurs minutes ; pendant ce temps l'attente n'écrivait rien
-> le log restait VIDE et paraissait « bloqué ». _launch_one écrit
désormais un en-tête d'attente immédiat + un battement toutes les ~30 s,
et l'attente sshd/cloud-init passe de ~12 à ~20 min (boot émulé lent).
_qemu_wait_ssh (chemin streamé) : timeout porté à 20 min également.
Validé : nommage (amd64/natif sans suffixe, s390x/arm64 suffixés) ; le log
se remplit dès le départ (en-tête + « ... 30s »).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ajoute l'architecture arm64 (aarch64) au déploiement, en plus d'amd64
(x86_64) et s390x. Disponibilité réelle vérifiée (juillet 2026) :
- arm64 : Ubuntu, Debian, Fedora (Arch : pas d'image cloud aarch64 -> rejet).
- s390x : Ubuntu seulement (inchangé).
deploy_qemu.py :
- host_arch() : arch native de l'hôte ; toute arch différente est ÉMULÉE
(TCG, --virt-type qemu) — plus de liste FOREIGN_ARCHES figée.
- virt_install : arm64 -> --arch aarch64 --machine virt --boot uefi (firmware
AAVMF résolu par libvirt) ; s390x inchangé ; x86 UEFI/OVMF inchangé.
- ensure_emulator(arch) généralisé : installe qemu-system-aarch64 + firmware
UEFI AAVMF (arm64) ou qemu-system-s390x (s390x), selon le gestionnaire de
paquets, avec vérif de présence (binaire + firmware pour arm64).
- Validation --arch tôt via ARCH_DISTRO_SUPPORT (message clair si la distro
ne publie pas l'arch). Les URLs d'images arm64 marchent déjà via les alias
(fedora/arch aarch64 ; ubuntu/debian arm64).
todo.py : menus d'architecture (VM unique + infra) proposent amd64/arm64/
s390x selon la distro ; l'architecture NATIVE de l'hôte est marquée d'un *
(défaut) ; les autres sont signalées « émulé, lent ». Le catalogue infra est
restreint aux distros publiant l'arch choisie (arm64 -> ubuntu/debian/fedora,
s390x -> ubuntu).
install_debian_dependency.sh : wkhtmltopdf (deb amd64 en dur) ignoré
best-effort sur toute arch non-x86_64 (au lieu d'avorter l'install).
Validé : dry-runs Ubuntu/Debian/Fedora arm64 (bonnes URLs + virt-install
aarch64/virt/uefi/qemu), rejet Arch arm64, non-régression s390x, menus et
filtrage simulés (natif *, arm64 par distro).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
repo sync échouait sur Fedora :
git: 'daemon' is not a git command.
fatal: cannot obtain manifest git://127.0.0.1:9418/ -> Connection refused
Unable to sync manifest ./manifest/git_manifest_erplibre.xml
Cause : sur Fedora la sous-commande « git daemon » N'EST PAS fournie par le
paquet « git » de base (contrairement à Debian/Ubuntu/Arch) mais par le
paquet séparé « git-daemon ». ERPLibre sert son manifeste via un « git
daemon » local (git://127.0.0.1:9418/) pendant repo sync ; sans le binaire,
le daemon ne démarre pas -> connexion refusée -> synchro du manifeste KO.
Fix : ajout de git-daemon aux dépendances Fedora.
Validé sur VM Fedora 42 réelle : avant, « git daemon » rc=1 ; après
« dnf install git-daemon », rc=0, et un git daemon local sert bien
git://127.0.0.1:9418/ (git ls-remote renvoie les refs).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Menu « Déployer l'infra ERPLibre » :
- Nouveau choix [granulaire] dans la sélection des distributions : affiche
la liste À PLAT de TOUTES les versions du catalogue (distro + version,
défauts marqués d'un *) et permet d'en choisir plusieurs par leurs numéros
séparés par des virgules (ex. « 1,3,12 »). Complète [all] et [principal].
- Demande l'ARCHITECTURE du parc en début de flux, par défaut l'architecture
NATIVE de l'hôte (amd64/arm64/s390x via uname -m). s390x est proposé mais
émulé (lent) et n'a d'images cloud que pour Ubuntu -> le catalogue est
alors restreint à Ubuntu (les autres distros sont ignorées avec un
avertissement). --arch est transmis à deploy_qemu.py quand ≠ amd64.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Après le passage à pacman, l'init PostgreSQL échouait :
« /usr/bin/postgres: /usr/lib/libm.so.6: version GLIBC_2.44 not found ».
Cause : Arch est en rolling release et NE SUPPORTE PAS les mises à jour
partielles. Sur une image cloud dont la glibc date du build de l'image, un
« pacman -S <paquet récent> » (postgresql 18.4) installe un binaire lié à
une glibc plus récente que celle du système -> symbole introuvable.
Fix : mise à jour COMPLÈTE (pacman -Syu) avant d'installer, dans les deux
endroits : install_arch_linux.sh (deps OS) et le bootstrap Arch de
_qemu_erplibre_remote_cmd (todo.py, remplace -Syy par -Syu).
Validé sur VM Arch réelle : après -Syu, « postgres (PostgreSQL) 18.4 » OK,
initdb OK, service postgresql actif, superuser erplibre créé, psql connecte.
Chaîne complète confirmée : gcc 16.1.1, make, node, psql tous présents.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Deux échecs d'install ERPLibre sur VM constatés dans les logs :
Arch — install_arch_linux.sh utilisait « yay » (helper AUR ABSENT d'une
image cloud et qui refuse de tourner en root) et n'installait JAMAIS
base-devel : sur une image cloud fraîche rien ne s'installait, donc pas de
compilateur C -> pyenv : « no acceptable C compiler found » -> build de
Python 3.12 échoué -> pas de venv -> pip/poetry/repo absents. Réécrit avec
pacman (dépôts officiels) : base-devel (gcc/make), deps de build pyenv
(openssl/zlib/xz/tk/…), PostgreSQL (+ initdb, Arch ne l'initialise pas),
Node/npm, deps Odoo en best-effort paquet par paquet (pacman refuse toute
la transaction sur un seul nom inconnu). wkhtmltopdf (AUR) : ignoré
proprement.
Fedora — « Connection closed by remote host » (exit 255) dès le début :
« cloud-init status --wait » tournait dans une session SSH UNIQUE tuée
quand cloud-init régénère les clés d'hôte et REDÉMARRE sshd au 1er boot.
On attend désormais la FIN de cloud-init via des connexions COURTES
successives (chaque tentative survit à un redémarrage de sshd), AVANT de
lancer l'install — dans les deux chemins : _qemu_wait_ssh (streamé) et
_launch_one (détaché/monitoré). On matche sur le TEXTE de « cloud-init
status » (done/disabled/error/degraded) car son code de sortie n'est pas
fiable selon la version.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sur apt, qemu-system-misc NE contient PAS l'émulateur s390x (il fournit
alpha/avr/hppa/… mais pas s390x) : ensure_s390x_emulator installait donc
un paquet inutile puis échouait (« qemu-system-s390x toujours absent »).
Le binaire est fourni par le paquet dédié qemu-system-s390x.
Validé de bout en bout : déploiement réel Ubuntu 24.04 s390x émulé (TCG) ->
domaine type=qemu, machine s390-ccw-virtio, emulator qemu-system-s390x ->
boot OK, IP DHCP, SSH par clé (erplibre), sudo OK, uname -m = s390x
(IBM z15, /proc/sysinfo Type: 8561).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ajoute --arch s390x au déploiement QEMU. Réalité constatée (juillet 2026) :
seul Ubuntu publie des images cloud s390x (cloud-images.ubuntu.com : OK) ;
Debian et Fedora n'en publient pas (404), Arch ne cible que x86_64/aarch64.
s390x est donc scopé à Ubuntu et rejeté proprement pour les autres distros.
deploy_qemu.py :
- FOREIGN_ARCHES / S390X_DISTROS : s390x n'est valide que pour Ubuntu.
- virt_install : sur s390x -> --arch s390x, --machine s390-ccw-virtio,
--virt-type qemu (émulation TCG, pas de KVM sur hôte x86), console SCLP
(pas de série ISA), amorçage IPL/zipl (aucun --boot UEFI/BIOS). Les
disques/réseau bus=virtio sont mappés en virtio-ccw par libvirt.
- ensure_s390x_emulator : installe qemu-system-s390x au besoin
(apt: qemu-system-misc, dnf: qemu-system-s390x, pacman:
qemu-emulators-full, zypper: qemu-s390).
- Avertit que s390x est ÉMULÉ (lent) sur un hôte x86.
todo.py : le menu « Déployer une VM » demande l'architecture (s390x proposé
uniquement pour Ubuntu, avec avertissement lenteur) et passe --arch.
install_debian_dependency.sh : wkhtmltopdf n'a pas de build s390x -> skip
best-effort au lieu d'avorter l'install (la branche s390x rust-all existait
déjà pour compiler les wheels).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fedora install stalled: the log stopped right after boot with no exit
marker. Cause: cloud-init was running "dnf -y install qemu-guest-agent
openssh-server" at first boot; on this host's slow VM network that dnf
took many minutes (metadata refresh + download), so cloud-init stayed
"running" and the install (which waits for cloud-init) never proceeded.
Fix: do NOT install any package via cloud-init. sshd is already present
in every cloud image (verified Ubuntu/Debian/Fedora/Arch) and is just
enabled via runcmd; the tools (curl/git/make…) are installed by the
ERPLibre bootstrap with optimised mirrors. cloud-init now finishes in
seconds (Fedora: "status: done" at 10s instead of blocking 5+ min).
--package still adds a cloud-init packages block on demand.
Also: when ERPLibre is installed via the infra menu, add
ERPLIBRE_EXTRA_DISK_GB (5G) to each VM's disk, because the image minimum
left only ~97 MB free after installation. Drop the now-useless
"--package git" (the bootstrap installs git).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Arch install failed at once: "error: failed to synchronize all
databases (unable to lock database)" then "reflector: command not
found". Cause: cloud-init was still running its own pacman (installing
qemu-guest-agent/openssh) and held /var/lib/pacman/db.lck when our
bootstrap started, so every pacman call failed and reflector never got
installed.
Fix: the remote install now waits for cloud-init to finish first
("cloud-init status --wait", 15 min timeout) before touching any package
manager — this releases the apt/dnf/pacman lock cloud-init holds during
its package stage (helps all distros). For Arch, also drop a stale
db.lck when no pacman is actually running. Verified live: cloud-init
done, no lock, reflector + refresh + curl/git/make all succeed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Dashboard: the info bar under the log now shows the selected VM's log
file path (in addition to its ssh command), so it can be opened/shared.
Install bootstrap: each package-manager branch first refreshes the repos
so the VM is as fast as possible, and Arch (pacman) is now supported:
- apt: apt-get update
- dnf: dnf makecache (dnf5 picks the fastest mirrors)
- pacman (Arch): reflector selects the 20 fastest HTTPS mirrors ->
pacman -Syy -> install curl/git/make (Arch was previously unsupported
by the bootstrap and would abort with "no package manager").
Validated live on the Arch VM: reflector + refresh + install all OK.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
cloud-init reported "status: error" on Arch because the seed's packages
list requested "openssh-server", which does not exist on Arch (the
package is "openssh"). Pick the SSH server package per distro: "openssh"
on Arch, "openssh-server" elsewhere. Verified on Arch: cloud-init now
finishes with "status: done", openssh installed, sshd active.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Arch VMs booted (after the Secure Boot fix) but eth0 stayed DOWN and
cloud-init never finished -> no IP, no SSH. Cause found in the guest:
cloud-init's Arch renderer writes the systemd-networkd [Match] Name from
the ethernet KEY of the v2 config (ignoring "match:"), so our key
"primary" produced "Name=primary" which matches no interface (Arch's NIC
is eth0). Rename the key to "eth0": Arch now renders "Name=eth0" and
DHCPs; Debian/Fedora/Ubuntu still use "match: name: e*" (verified Debian
redeploy: enp1s0 UP with an IP), so no regression.
Arch is now fully working end to end: boot, DHCP IP, SSH by key, user
erplibre in sudo, passwordless sudo, console login erplibre/erplibre.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
OVMF Secure Boot rejected Arch's unsigned GRUB with "Access Denied" ->
"No bootable option" -> dropped to the firmware menu, never booted.
Disable Secure Boot in the UEFI boot (firmware.feature secure-boot=no):
Arch now boots and cloud-init applies user/password/hostname/ssh-key
(console login erplibre/erplibre works). Ubuntu/Debian/Fedora also boot
without Secure Boot (their signed shim is not required), no regression.
Note: Arch networking (eth0 stays down, cloud-init does not finish) is
still under investigation and tracked separately.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Arch (rolling release, single "latest" version) joins ubuntu/debian/
fedora: official cloud image geo.mirror.pkgbuild.com/images/latest/
Arch-Linux-x86_64-cloudimg.qcow2 (cloud-init included), osinfo=archlinux
(known locally), UEFI boot + virtio seed like the others. ERPLibre
already has install_arch_linux.sh. Exposed in the infra/deploy menus.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
wkhtmltopdf 0.12.6.1-3 no longer ships a fedora-* RPM (my URL 404'd).
The AlmaLinux 9 (EL9) RPM is compatible with Fedora — tested on the
Fedora 42 VM: "dnf install <almalinux9 rpm>" resolves the deps and
"wkhtmltopdf --version" reports 0.12.6.1 (with patched qt). Use it, with
an AlmaLinux 8 fallback.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Tested make install_os on a real Fedora 42 VM (exit 0: gcc 15, psql,
node 22 installed). Two fixes from the run:
- postgresql-setup --initdb failed with "invalid locale settings"
because Fedora cloud images ship no LANG; force a valid locale via
PGSETUP_INITDB_OPTIONS=--locale=C.UTF-8 and wipe a partial data dir
first (a failed init leaves /var/lib/pgsql/data/log behind and blocks
the retry). Verified: cluster PG 16 initialised, service active,
erplibre superuser created.
- the dev-tools group was referenced by display name ("C Development
Tools and Libraries", "No match"); use the group id c-development.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
From the QEMU install logs: Debian 12 succeeded, Ubuntu 24.04 and
Fedora 42 failed.
Ubuntu (install_debian_dependency.sh):
- apt failed with "Could not get lock" (cloud-init/unattended-upgrades
hold it on first boot) -> all apt-get calls now use
DPkg::Lock::Timeout=600 so apt waits for the lock.
- "postgis" is not a package on Ubuntu 24.04, so the postgresql line
exited 1 before installing build-essential -> no C compiler -> pyenv
could not build Python 3.12.10. PostGIS is now best-effort (tries
postgis, then postgresql-postgis) and never aborts; postgresql-contrib
is added.
Fedora (new install_fedora_dependency.sh, wired into install_dev.sh):
- install_dev.sh dispatched Fedora to the apt script; it now has a
fedora / ID_LIKE branch calling a dnf-based dependency installer
(dev tools, postgresql-server + initdb, pyenv build deps, node, etc.),
using --refresh --skip-unavailable to tolerate mirror/name issues.
Bootstrap (todo.py): the dnf "curl git make" install hit a GPG/checksum
failure on a fresh image; it now uses --refresh and retries after
"dnf clean all".
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
After "Opening the interactive monitor..." the install now lists every
VM's log file path, so they can be read or shared even when leaving the
dashboard before the installs finish.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Root cause of "git/make: command not found": the bootstrap ran
"apt-get update -qq && apt-get install -y $PKGS". Inside an && list,
set -e does NOT abort on failure, so when apt-get update failed (flaky
VM network) the install was silently skipped and the script marched on
without git/make. Now: "apt-get update || true; apt-get install -y
$PKGS" (update best-effort, install mandatory), followed by an explicit
"command -v curl git make" check that exits with a clear message
("Outil manquant ... (reseau de la VM ?)") instead of a cryptic failure
later. Shared by the streamed and the monitored install paths.
Dashboard: add "c" to copy the selected VM's full log to the clipboard
(OSC 52, works over SSH) with a notification; the ssh bar documents
Shift+drag for native terminal selection; on close, print a ready-to-
copy "tail -n +1 <logdir>/*.log" to share the logs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Deploying ERPLibre (single VM [1] and infra [4]) now asks "Interactive
monitoring dashboard?". When yes, the installs run DETACHED (setsid, one
log file + exit marker per VM) and a Textual dashboard opens:
- left: a table of VMs with live status (⏳ running / ✅ done / ❌ failed
with exit code) and elapsed time;
- right: the selected VM's log, tailed live (f toggles follow);
- footer: the VM's ssh command; s suspends the dashboard and SSHes in;
- q quits back to the menu — the installs keep running detached, so you
can leave before they finish and reopen later on the same log dir.
New module qemu_install_monitor.py holds the detached launcher and the
Textual app. The remote install script is factored into
_qemu_erplibre_remote_cmd. Falls back gracefully if textual is missing.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds textual to the ERPLibre tooling requirements (installed into
.venv.erplibre by install_locally.sh) for the upcoming interactive
dashboard that monitors parallel ERPLibre installs across VMs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Right after boot the VM's sshd may not be up yet (notably on Fedora),
so the install failed with "Connection refused". _qemu_install_erplibre_vm
now polls SSH (ssh ... true, BatchMode) for up to 3 minutes and skips
with a clear message if it never answers, instead of failing instantly.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The infra distribution prompt now accepts "principal" (alongside
numbers and "all"): it deploys the default version of each distro —
ubuntu 24.04, debian 12, fedora 42 — one VM each. The default version is
also marked with a * in the distro listing so it is visible.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Minimal cloud images ship without make (and often without curl), so the
ERPLibre install cloned the repo then died on "make: command not found".
The remote bootstrap now installs curl, git and make via apt (Debian/
Ubuntu) or dnf/yum (Fedora) — running apt-get update first since
cloud-init did not refresh the lists — before cloning and running
make install_os / make install_odoo_18.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Installation was very noisy: poetry ran "install -vvv" and repo sync /
git daemon ran with -v/--verbose. They are now quiet by default
(poetry -q, repo sync -q, no git daemon --verbose) and the detailed
logs come back only when EL_VERBOSE=1.
Applied to install_locally.sh (poetry) and every manifest script (repo
sync + git daemon). env_var.sh documents EL_VERBOSE and respects a value
already set in the environment, so "EL_VERBOSE=1 make install_odoo_18"
works.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
"Deploy a new VM" now asks, after a successful deploy, whether to
install ERPLibre into ~/git/erplibre (pick a branch), just like the
infra command.
Installing ERPLibre no longer only clones the repo: after the clone it
runs "make install_os" then "make install_odoo_18" over SSH (the
erplibre user has passwordless sudo from cloud-init). The Odoo target is
a constant (ERPLIBRE_ODOO_TARGET) and both the single-VM and infra flows
use the same install path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The cleanup command now runs, each listed and confirmed independently:
1. orphan files (disks/seeds/.part/nvram) — as before;
2. ghost domains: libvirt domains whose disk no longer exists ->
offer destroy + undefine --nvram;
3. stale codename-named Ubuntu images (noble/resolute/... left by the
move to /releases/ version-named images) -> targeted delete;
4. orphan ~/.ssh/config entries: "Host erplibre-*" blocks with no
matching VM (personal SSH hosts are never touched);
5. stale libvirt DHCP leases whose MAC belongs to no VM (best-effort
rewrite of the dnsmasq status file + SIGHUP; they also self-expire);
6. the full base-image cache remains an explicit opt-out at the end.
Detections verified read-only on the host.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
New Manage entry that finds and lists QEMU leftovers with no matching
libvirt domain, then asks before deleting:
- working disks /var/lib/libvirt/images/<name>.qcow2
- cloud-init seeds .../iso/<name>-seed.iso
- interrupted downloads (*.part)
- orphan UEFI nvram files
Sizes are shown human-readable with a total. Cached base cloud images
(reusable) are listed and offered separately, since removing them forces
a re-download; this also cleans the stale codename-named duplicates left
by the switch to /releases/ version-named images.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Three papercuts hit when creating a VM from the menu:
1. Disk size "30" was passed verbatim to qemu-img resize, which read it
as 30 BYTES and failed ("use --shrink"). The prompt now shows units
(e.g. 30G, 1T) and a bare number is normalised to GB ("30" -> "30G"),
in the menu and in deploy_qemu.py itself.
2. The VM name is no longer required: leaving it blank uses the default
erplibre-<distro>-<version> (e.g. erplibre-ubuntu-2604). Distro and
version are now asked first so the default can be offered.
3. The menu ignored the deploy exit code and carried on to "wait for the
VM IP" even after a failure, hanging forever with the error scrolled
off. It now checks the return code and stops with a clear message.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Debian 13 (trixie) genericcloud dropped the BIOS/GRUB-pc bootloader, so
SeaBIOS looped forever on "Booting Debian GNU/Linux" and the install
never completed. Boot via UEFI (OVMF) by default: trixie boots, and
Ubuntu/Debian 12/Fedora keep working (verified Debian 13 and Ubuntu
24.04 end to end). Add a --bios opt-out for hosts without OVMF, and pull
the UEFI firmware (ovmf / edk2-ovmf) with the libvirt/QEMU stack. VM
deletion already removes the nvram (undefine --nvram).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When the exact osinfo id is missing from the local osinfo-db (e.g.
ubuntu26.04, fedora43/44), the fallback was detect=on,require=off, which
made virt-install fall back to "generic" and warn "VM performance may
suffer". Fall back instead to the latest KNOWN osinfo of the same distro
(ubuntu26.04 -> ubuntu25.10, fedora44 -> fedora42): proper virtio/OS
defaults, no warning. detect=on,require=off remains only if nothing of
the family is known.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Debian VMs came up with no configured user, no hostname and no SSH key,
so neither the console (erplibre/erplibre) nor SSH worked. Two causes,
found via the guest cloud-init.log:
1. The NoCloud seed was attached as a CD-ROM. Debian's initramfs does
not load the CD driver (sr_mod) at the init-local stage, so the
"cidata" volume was invisible and cloud-init fell back to an empty
DMI seed (Ubuntu tolerates the CD). Attach the seed as a read-only
virtio disk instead: virtio-blk is in the initramfs, the label is
seen immediately, and the user-data is applied. Works for Ubuntu too.
2. The user was added to group "admin", which does not exist on Debian
(it does on Ubuntu) -> useradd failed and the user was never created.
Use "sudo" (present on both) instead.
Validated end to end: Debian 12 and Ubuntu 24.04 both get the erplibre
user (in sudo), the hostname, console login erplibre/erplibre and SSH by
key.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The infra "Parallel deployments" prompt defaulted to a fixed 4. It now
defaults to the host CPU count (os.cpu_count()), still capped by the
number of selected VMs, with a fallback of 4 when the count is unknown.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
_write_ssh_config_entry used a DOTALL (?ms) regex to remove an existing
"Host <name>" block; with DOTALL, ".*" crossed newlines and matched from
the block to the end of the file, so re-adding an already-present host
deleted every entry after it. Dropping DOTALL (keep MULTILINE, match
indented lines with [^\n]*) removes only the target block and preserves
the rest. Not a parallelism issue — the writes are sequential.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Interim Ubuntu releases (24.10, 25.04) were removed from cloud-images
/current/ once EOL, so their image URL 404'd. Switch the Ubuntu image
URL to /releases/<version>/release/ubuntu-<version>-server-cloudimg-
<arch>.img, which stays available for every published release, and drop
the EOL interim versions from the catalogue (keep supported LTS + 25.10
+ 26.04).
Also:
- silence the "crypt is deprecated" DeprecationWarning (openssl fallback
already covers Python 3.13+);
- a 404 during download now reports "image not found (EOL/removed), pick
a supported LTS" instead of blaming connectivity;
- add a NoCloud network-config (DHCP on e*) to the seed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The infra fleet was deployed one VM at a time, each blocking up to 90s
waiting for its DHCP lease. Deploys now run concurrently through a
ThreadPoolExecutor (default min(count, 4), promptable), with each job's
output captured and printed per VM as it finishes (no interleaving). VMs
are created with --no-wait-ip so workers return quickly; IPs are then
collected once for the ~/.ssh/config step. Preferred over shelling out
to GNU parallel: no external dependency, grouped output, and it stays
integrated with the ssh-config and ERPLibre-clone steps.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
cloud-init's package_update/package_upgrade forced apt to hit the distro
mirrors on first boot; on a slow or unreachable network that stalls
cloud-init and delays SSH availability. They are now off by default
(SSH is already present in the cloud images and enabled via runcmd);
--apt-update opts back in (and runs upgrade too unless --no-upgrade).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
VMs were created with user "erplibre" but NO password (lock_passwd:
true), so the serial console refused every login and only SSH-by-key
worked — hence "can't connect".
- The todo deploy flow (single VM and infra) now sets a default console
password "erplibre" for the "erplibre" user, in addition to the SSH
key, so virsh console and password SSH both work. It is shown as
"Console/SSH login: erplibre / erplibre" and can be changed at the
single-VM prompt.
- The console entry now prints the default login and the Ctrl+] hint.
- After creating a VM, offer to add it to ~/.ssh/config (Host <name>,
HostName <ip>, User erplibre) so "ssh <name>" just works; the infra
flow asks once and adds every VM. Existing blocks for the same Host
are replaced, the rest of the file is preserved, mode kept 600.
Note: VMs created before this change have no console password — delete
and redeploy them to get the erplibre/erplibre login.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Group the longer menus with section headers (reordering items so the
numbering stays sequential within each section):
- Execute: Development / Data / Sources & documentation / AI &
automation / Deployment, network & security / Preferences.
- QEMU/KVM: Deployment / Manage / Catalog.
- Database: Backup / Restore / Danger zone.
- RTK: Setup / Status / Optimize.
- Config: Generate / Advanced.
The QEMU config-entry lookup now skips section rows so appended
makefile entries still map to the right number.
Also in the QEMU menu:
- new "Open the console on a VM" entry (lists VMs, asks which, reminds
Ctrl+] to quit, runs virsh console);
- "Show a VM IP address" accepts "all" to print every VM's IP.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
fill_help_info() now renders {"section": "..."} entries as a header line
without consuming a number, so numbering stays continuous over the real
commands (the hardcoded elif chains keep matching). The Deploy menu is
split into "Local", "SSH (remote host)" and "Virtualization &
notifications", making the long SSH block easy to scan.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
New QEMU/KVM menu entry to remove one, several or all VMs. It lists the
defined domains, lets you pick by number or "all", asks whether to also
delete the disk images (the qcow2 working disk + the seed ISO), shows
what will be removed and asks for confirmation. Each VM is powered off
(destroy) then undefined (--nvram, with a fallback for older virsh);
disks are only deleted when explicitly requested.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Debian deploy appeared frozen: cloud.debian.org sent this host to an
unreachable mirror and urlretrieve had no timeout, so it hung forever;
worse, output was block-buffered under the todo pipe so nothing showed.
- Stream downloads via urlopen with a 30s per-operation timeout: a dead
mirror now fails fast instead of hanging.
- Try several Debian mirrors in order (cloud.debian.org, then two
acc.umu.se mirrors) — first responsive one wins.
- Reconfigure stdout to line-buffered in main() so headers and progress
appear live even when captured by the menu.
- Add openssh-server to the cloud-init packages and a runcmd enabling
ssh/sshd, so every VM (Debian genericcloud included) is SSH-reachable.
- Add a timeout to the SHA256SUMS fetch too.
Validated: debian-11 fails over from cloud.debian.org to gemmei and
completes; debian-12 downloads fully.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Generalise the _is_yes() helper (y/yes/o/oui) across the whole file so
French answers work everywhere, and add _is_no() (n/no/non) for the
default-yes prompts. Converted: system-install, Pycharm, SSH-password
(default yes via _is_no), template overwrite, keep-temp-database (was
locale-gated, now accepts both), git-repo fetch and the mobile
personalize/debug/picture prompts. Drop the now-unused get_lang import.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The confirmations only matched "y", but the French prompts show "(o/N)",
so answering "o" (oui) was treated as no — the infra deployment aborted
with "Annulé." right after the user confirmed. Add a _is_yes() helper
(y/yes/o/oui) and use it for the infra deploy confirmation, the ERPLibre
install prompt, the disk-overwrite prompt and the SHA256 verify prompt.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
New QEMU/KVM entry that stands up a fleet of minimal VMs, one per
selected cloud image. It:
- lets you pick distros then versions (multi-select, "all" per level, or
the whole catalogue), reading the specs straight from deploy_qemu.py so
there is no duplication;
- prints a plan with each VM's minimum RAM/disk, the total concurrent RAM
and virtual disk, and the host's available RAM, warning when the fleet
cannot all run at once;
- deploys sequentially after confirmation (minimum sizing per version),
skipping VMs that already exist;
- optionally clones ERPLibre into ~/git/erplibre on each VM, asking which
branch (list fetched via git ls-remote) and pulling git in through
cloud-init.
Naming is erplibre-<distro>-<version> (e.g. erplibre-ubuntu-2404).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Every menu now prints a breadcrumb line (e.g. "📍 TODO › Execute ›
Deploy › QEMU/KVM") right above "Command:", so it is always clear where
you are and the path can be copied to describe a menu unambiguously.
The trail is derived from the call stack via a method-name -> label map,
so no menu method had to change: fill_help_info and the three inline
menus just render self._menu_header() instead of t("Command:").
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Generalise the deployer beyond Ubuntu with a distro registry and a new
--distro flag (ubuntu default, debian, fedora). Each distro keeps its
own codenames, osinfo ids and per-version minimum RAM/disk. Image URLs
are built per distro (Ubuntu current/, Debian latest/, Fedora resolved
from the release index since it has no "latest" link).
Add --list-images to print the whole catalogue with specs, exposed in
the todo menu ("List available images and specs"); the deploy/download
flows now prompt for the distro first.
Resilient osinfo: when the local osinfo-db does not know an id (e.g.
ubuntu26.04, fedora43+), fall back to virt-install detect=on,require=off
instead of failing. Ubuntu 26.04 (resolute) is now enabled.
Also: throttle the download progress to integer-percent steps (single
updating line on a TTY, at most 101 lines when captured by the menu),
and skip the spurious "network is already active" error by checking the
libvirt network state before net-start.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The "Show a VM IP address" entry prompted for a VM name with no visible
list, so the user had to guess the name/ID. It now runs virsh list --all
first and asks for a "VM name or ID", making the expected input obvious.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replace the fixed 8192 MB / 4 vCPU / 20G defaults with the minimum
resources required by the chosen Ubuntu version (libosinfo/osinfo-db
values): 20.04 -> 2048 MB/5G, 22.04 -> 2048 MB/10G, 24.04+ -> 3072
MB/20G. --vcpus now defaults to 2. This stops a small host from being
starved by an oversized default (an 8 GB VM failed to allocate on a
6.7 GB host) and silences the libvirt "less than recommended" warning.
In the todo menu, leaving the RAM/vCPU/disk fields blank now means
"version minimum" (the flag is simply not passed) instead of forcing
8192/4/20G. Any explicit value still overrides. README regenerated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Expose the QEMU deploy script from the interactive assistant so users
can create, preview, download and manage Ubuntu VMs without memorising
CLI flags. Adds a QEMU/KVM entry under Deploy, its fr/en translations,
and an extensible qemu_from_makefile section in todo.json.
Generated by Claude Code 2.1.210 claude-opus-4-8
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
The script previously required a manually supplied image path and only
installed the client tools, so a bare host aborted at virt-install with
a missing libvirt-sock. It now derives and downloads the Ubuntu cloud
image from --version, detects and installs the full libvirt+QEMU stack
(daemon and system emulator included) through the host package manager,
then enables libvirtd. Adds a --download-only mode, a hypervisor
preflight check, clean error messages instead of Python tracebacks, and
a bilingual usage guide (README).
Generated by Claude Code 2.1.210 claude-opus-4-8
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
Give users a guided, confirmation-gated way to drop one or all
databases from the interactive CLI. Previously this meant running
make db_drop_all or odoo_bin db --drop by hand, which is easy to
mistype and offers no safeguard. The new entry requires an explicit
'oui'/'yes' (default no) before any irreversible deletion.
Generated by Claude Code 2.1.191 claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
Provides Claude Code with structured context about the
ERPLibre Home Mobile project (OWL 2, Capacitor 8, Vite,
Vitest, SQLite) including stack, conventions, commands,
and migration patterns.
Generated by Claude Code 2.1.108 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
sentencepiece provides the SentencePiece tokenizer (C++) needed by the
MarianMT Android plugin. Cloned shallow (depth=1) alongside whisper.cpp.
Generated by Claude Code 2.1.105 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
Add a one-command installer for the ntfy push notification server
(Ubuntu/Debian and Arch Linux), wired into the todo.py Deploy menu.
Users can now deploy a local ntfy server from the CLI and subscribe
to topics from their mobile device (ntfy app) to receive push
notifications from ERPLibre.
Generated by Claude Code 2.1.101 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
Ajout remote ggerganov et projet whisper.cpp (clone shallow, revision
master) dans le manifest Google Repo mobile. Chemin cible :
mobile/erplibre_home_mobile/android/app/src/main/cpp/whisper
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Enable deploying and managing ERPLibre on remote servers via SSH
directly from make and the interactive todo.py CLI, since only
local deployment was previously supported.
- New conf/make.ssh.Makefile with 11 targets: ssh_check, ssh_push,
ssh_install, ssh_run, ssh_stop, ssh_restart, ssh_status, ssh_logs,
ssh_make, ssh_install_systemd, ssh_install_nginx
- Variables: SSH_HOST (required), SSH_USER, SSH_PORT, SSH_KEY,
SSH_PATH, SSH_TARGET, SSH_DOMAIN, SSH_ADMIN_EMAIL
- Execute > Deploy menu extended with 11 SSH options in todo.py
- All strings translated fr/en in todo_i18n.py
Generated by Claude Code 2.1.101 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
@capacitor/cli v8.x requires Node.js >=22.0.0. The install script
was pinning NODE_MAJOR=20, causing a fatal error when installing
the mobile app via todo.py.
Generated by Claude Code 2.1.101 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
Poetry install and git-repo sync are independent (different write
paths). Running them sequentially wastes time. Split install_locally.sh
into EL_PHASE=setup|poetry|all phases so install_locally_dev.sh can
background the repo sync while poetry runs in the foreground, reducing
total install time by up to 50% on slow connections.
Set EL_PARALLEL_INSTALL=0 to restore sequential behavior for debugging.
Generated by Claude Code 2.1.88 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
CybroOdoo repos are large and slow to clone, making them unsuitable
for default installation. Moves them to opt-in per-version extra
manifests, introduces .erplibre-state.json to track installation
options per Odoo version, and surfaces the choice in the TODO CLI
sub-menu. Switch auto-detects extra from state and warns when no
state is recorded.
Generated by Claude Code 2.1.87 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
make version had no visibility into whether the mobile project was
active. Add detection via presence of mobile/erplibre_home_mobile
directory (cloned by repo sync --with_mobile) and display its status
alongside the existing Odoo/Python/Poetry version info.
Generated by Claude Code 2.1.87 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
Define the integration contract between erplibre_mobile and ERPLibre
platform: JSON-RPC 2.0 call specs, Note→project.task field mapping,
GeoMultiPoint format for geolocation entries, conflict resolution
strategy, re-auth flow, and version compatibility matrix.
Establishes the shared source of truth before implementation begins.
Generated by Claude Code 2.1.87 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
Add a multi-agent orchestrator that coordinates all 25 specialist agents
through 5 phases (analysis, design, implementation, verification, release)
with agent-to-agent communication via Agent Teams. Add /feature slash
command as the entry point. Enable CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS
in project settings to allow direct inter-agent messaging.
Generated by Claude Code 2.1.81 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>
Add a full catalog of Claude Code subagents covering all development
disciplines needed for a banking-grade open-source mobile app:
code quality, QA, backend, frontend, UX, architecture, security,
docs, community, product, ethics, DevOps/SRE, release, incident
response, performance, pentest, accessibility, compliance, risk,
data governance, legal/license, support, localization, and AI
agent engineering.
Generated by Claude Code 2.1.81 model claude-sonnet-4-6
Co-Authored-By: Mathieu Benoit <mathben@technolibre.ca>