Commit Graph
109 Commits
Author SHA1 Message Date
Bohan Jiangandmultica-agent 4f5fbc9216 fix(vcs): explain untrusted TLS certificates and trust private CAs via Helm (MUL-7639) (#8761)
* fix(vcs): report untrusted provider TLS certificates on connect (MUL-7639)

ConnectVCS reported every non-token failure as "could not reach the
provider instance" and logged nothing, so a Gitea behind a private CA
looked like a network problem. Log the underlying validation error and
tell certificate failures (untrusted CA, host name mismatch, other
verification failures such as expiry) apart from unreachable instances.
Status codes are unchanged.

Co-authored-by: multica-agent <github@multica.ai>

* feat(helm): trust extra CA certificates in the backend (MUL-7639)

backend.extraCACerts.configMap mounts an existing ConfigMap of PEM
certificates read-only and points SSL_CERT_DIR at the system directory
plus the mount, so the backend trusts an internal CA without dropping
public CAs or disabling TLS verification. Unset renders unchanged.

Co-authored-by: multica-agent <github@multica.ai>

* docs(self-host): document trusting a private CA (MUL-7639)

Co-authored-by: multica-agent <github@multica.ai>

* fix(helm): render when values predate extraCACerts (MUL-7639)

helm upgrade --reuse-values from an older chart carries no
backend.extraCACerts key, and reading .configMap on it failed the whole
render with a nil pointer even when no CA was wanted. Fall back to an
empty dict, and cover the missing-key, default and enabled renders in
the chart test.

Co-authored-by: multica-agent <github@multica.ai>

* docs(self-host): restart Compose backend after replacing a CA (MUL-7639)

up -d keeps the running container when only a mounted CA file changed,
so the backend kept its old trust store. Use restart instead.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: multica-agent <github@multica.ai>
2026-09-24 13:04:01 +08:00
ZIceandmultica-agent 7a3f29b4be fix(installer): preserve Windows PowerShell 5.1 disk parsing (#8663)
Co-authored-by: multica-agent <github@multica.ai>
2026-09-22 15:49:49 +08:00
Bohan Jiangandmultica-agent 53578958bc fix(selfhost): pass MULTICA_DINGTALK_SECRET_KEY through and guard integration keys (#8617)
The backend service environment in docker-compose.selfhost.yml is an
explicit allowlist, and MULTICA_DINGTALK_SECRET_KEY was never on it:
Docker self-hosters who set the key in .env still got a disabled DingTalk
integration. Telegram had the same gap until #8611.

- Map MULTICA_DINGTALK_SECRET_KEY into the backend service.
- Document MULTICA_SLACK_SECRET_KEY and MULTICA_TELEGRAM_SECRET_KEY in
  .env.example next to the other channel keys.
- Extend selfhost-config.test.sh: every key the server loads through
  secretbox.LoadKey must be mapped in docker-compose.selfhost.yml and
  listed in .env.example. The list comes from the server, not the docs,
  because Telegram's key was missing from both.
- Run script-checks when server/cmd/server/router.go changes, where the
  integration gates live, so a new channel key trips the guard in the PR
  that adds it.

Co-authored-by: multica-agent <github@multica.ai>
2026-09-21 12:07:23 +08:00
0cb64a77aa feat(ui-lab): add design system workbench (MUL-7425) (#8388)
* feat(ui-lab): add shared component design workbench

* feat(ui-lab): support English and Chinese

* feat(ui-lab): scaffold design system catalog

* fix(ui-lab): cover narrow headers and CI path filtering

* feat(ui-lab): add dialog motion and component usage rules

* feat(ui-lab): add semantic color workbench

* fix(ui): update brand color to #0070E3

* feat(ui-lab): add color usage and live contrast checks

* feat(ui-lab): support collapsible sidebars

* fix(ui-lab): place collapse controls in sidebar headers

* fix(ui-lab): avoid repeating page titles in breadcrumbs

* test(agent): synchronize Cursor background finalization fixture

* fix(ui-lab): preserve contrast and list scroll restoration (MUL-7425)

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-09-16 14:39:55 +08:00
d7f8c3f2b6 MUL-7404: ci: scope validation and remove duplicate lifecycle runs (#8451)
* ci: scope validation and remove duplicate lifecycle runs

Co-authored-by: multica-agent <github@multica.ai>

* fix(migrations): renumber the Triage validate migration to 489

#8449 and #8436 merged minutes apart and both took 485, so
TestMigrationNumericPrefixesAreUnique fails on main and every PR branched from
it. #8436 owns 485-488, so the single validate file moves to 489.

Renaming an applied migration makes databases that already recorded
485_issue_triage_state_validate run it once more under the new version string.
That is safe here: VALIDATE CONSTRAINT on an already-validated constraint is a
no-op, verified against PostgreSQL 17.

Co-authored-by: multica-agent <github@multica.ai>
(cherry picked from commit 09140a04ae)

* fix(ci): gate installers and reuse frontend quality setup

Co-authored-by: multica-agent <github@multica.ai>

* docs(ci): remove redundant CI overview

Co-authored-by: multica-agent <github@multica.ai>

* refactor(ci): remove unused output and duplicate test fixtures

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: J <bohan@devv.ai>
2026-09-15 18:40:14 +08:00
02a8ff3e95 feat(maintenance): add resumable internal backfill API (MUL-7365) (#8436)
* feat(maintenance): add resumable internal API backfills (MUL-7365)

Co-authored-by: multica-agent <github@multica.ai>

* fix(maintenance): ship bounded container operations driver

Co-authored-by: multica-agent <github@multica.ai>

* test(cursor): synchronize ownership before finalization

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-09-15 17:47:54 +08:00
b54e665295 MUL-7280: ci(backend): cache race builds per commit and shard pkg/agent onto its own runner (#8302)
* MUL-7280: ci(backend): cache race builds per commit and shard pkg/agent onto its own runner

backend-tests never warmed its Go build cache: setup-go shares one go.sum-keyed entry across every Linux job, go-vulnerability-scan saved a 95 MB one first, and the test job race-compiled the whole module from scratch on every run. Own the caches explicitly (modules by go.sum, race build objects by commit SHA with prefix fallback, published from main only), compile build and vet in race mode so the objects are shared, and pin -count=1 so a rolling cache cannot replay results against a changed schema.

pkg/agent ran as a 139s serial tail at -parallel 2; it is timer-bound and reads no database, so it moves to its own runner without service containers. test-go.sh gains --only regular|agent for that split; the default still runs both halves.

Co-authored-by: multica-agent <github@multica.ai>

* MUL-7280: ci(backend): make backend-tests the sole Go cache writer and key build objects by image

backend-agent-tests restored the module cache through the combined actions/cache action with the same go.sum key as backend-tests. Being faster, it saved first with only pkg/agent's slice of the module graph, and backend-tests then failed to reserve the immutable key: the same first-writer-wins problem this change set out to fix. The agent job now restores both caches read-only, leaving backend-tests as the only writer.

The build-cache key also gains ImageOS. Cached cgo and race objects are linked against the runner image's glibc, so an entry must not cross an Ubuntu release when ubuntu-latest rolls forward; setup-go keys on ImageOS for the same reason (actions/setup-go#368). ImageOS is a runner environment variable rather than a context value, so each job resolves the key prefix once in a shell step and the restore and save steps reuse it.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-09-11 15:57:26 +08:00
6af2268bd9 MUL-7227 test(perf): measure comment typing while agent runs stream (#8276)
* test(perf): measure comment typing while runs stream (MUL-7227)

The regression that made typing a comment take tens of seconds was a summary
animation replacing a DOM node per streamed message, and nothing we had could
see it. The component suite runs in jsdom, which has no style recalculation, no
layout and no main-thread contention, so several thousand passing tests said
nothing about the thing users felt. #8260 pinned the DOM node reuse; this pins
what the user actually waits for.

One browser scenario, on the real page: app shell, shared CSS, the ProseMirror
composer, inline runs, a production Next.js build. Only HTTP and the WebSocket
are synthetic, so no API server, database, daemon, agent or account is
involved. Nothing disables animations or trims the DOM — the cost being
measured lives in exactly that machinery.

Fixed workload, identical for both builds under comparison: 80 comments with
prose and code, three running agents, 1000 seeded transcript messages each, 60
ticks of live messages 100ms apart, and 172 keystrokes at 25ms. Messages are
scheduled from Node rather than the page, so a stalling build cannot quietly
measure less work than the one it is being compared against.

`typing_elapsed_ms` is what the user waits for: first keystroke to the editor
holding the whole string, main-thread queueing included. Style recalculation,
layout, task time and long tasks say where it went. An empty page types very
fast, so a run whose messages never reached the UI, or whose composer did not
end up holding every character, is reported as `invalid` or `timeout` and never
as a duration — including when the scenario blows its budget, which is the case
whose numbers matter most.

Reports rather than gates. `scripts/perf-compare.mjs` builds and measures two
refs in sequence on one machine, each installed against its own lockfile, with
the spec, fixture and browser always taken from the running tree so the product
is the only difference. The workflow is separate and not required: one sample
per ref cannot separate a small regression from machine noise, and the noise
floor has not been calibrated on CI hardware yet.

Verified against the known regression: restoring the pre-#8260 summary
animation moves style recalculation from ~130ms to 1577ms and typing from
~4.9s to 6.2s, against a spread of ~1% across three runs of the fixed build.
Disconnecting the message injection reports `timeout: run 0 end marker`, and
typing fewer characters reports `timeout: editor content` — neither can pass as
fast. #8260's DOM reuse test is untouched and still green.

Tests and run configuration only: no product change.

Co-authored-by: multica-agent <github@multica.ai>

* test(perf): always build for real when comparing two refs (MUL-7227)

The comparison reports a build time, and turbo's cache made it read as ~1.7s
when the real build is ~60s — understating what this costs CI by a factor of
thirty. Worse, a run where one ref hit the cache and the other missed would put
that difference in the timing split as if it belonged to the products.

`--force` on both sides: two real builds, two comparable numbers.

Co-authored-by: multica-agent <github@multica.ai>

* test(perf): report cached builds instead of forcing real ones (MUL-7227)

Forcing a real build on both sides cost minutes on every local run to remove an
ambiguity that a single line of output removes instead. A restored build is
byte-identical to the one that produced it, so the cache cannot move the numbers
this script collects — only the build time it reports, which is now labelled
when it came from the cache.

It buys nothing on CI either: the runner is cold every time, so both builds are
real whether or not they are forced. Locally a repeat comparison drops from
roughly three minutes to under one.

Co-authored-by: multica-agent <github@multica.ai>

* test(perf): run the comparison on demand, and let it exit (MUL-7227)

The comparison never exited on success. The frontend server is spawned
detached, and its teardown hung off `process.on("exit")` — but a live child
keeps the event loop running, so `exit` never fired and the child was never
killed. The report was written and the process then waited on the server it
was supposed to stop. On CI that ran until the job was cancelled at sixty
minutes; locally it left the servers running after being killed. The failure
path only terminated because it happened to call `process.exit(1)`.

Teardown is explicit now. Each side stops its server, waits until the process
has exited and the port has stopped answering, and removes its worktree before
the next side starts — so the second build is also measured on a machine with
nothing of the first one running. The script then exits explicitly, and the
`exit` and signal handlers remain only as a last resort for cancellation.

Checked on every path, with a harness that records the gap between the report
and the exit and looks for anything left behind: success, a scenario failure
with a server running, an unknown ref, and SIGTERM mid-run all exit promptly
with no server, worktree or temp directory remaining. Run through the same
harness, the previous script wrote its report, then was still running 247s
later and left three servers and two worktrees behind.

The workflow is manual only, capped at fifteen minutes. On CI one comparison
took 340s — 292s of it the two production builds, 26s the two measurements —
which is a lot to spend on every frontend PR for a report that gates nothing.
The routine guard stays the component test from #8260, which pins the exact
failure mode and runs with the normal suite; this is for changes that touch
summary animation, long lists, transcript rendering or global styles.

No change to the scenario or its metrics.

Co-authored-by: multica-agent <github@multica.ai>

* test(perf): only trust this run's report, and wait for the whole server (MUL-7227)

Two ways the comparison could report something that was not true.

A report left in the output directory by an earlier run was read back when
this run's scenario failed before writing one, so two failed scenarios showed
the previous run's timings, were marked a usable comparison, and exited 0. And
a scenario that wrote `ok` and then failed was passed through too: the exit
code was recorded as `spec_failed` and then ignored. Each side's report is now
deleted before the scenario runs, the files this script writes are cleared up
front, and a failing scenario process overrides an `ok` report — whatever
failed after the numbers were written is exactly what nobody has looked at.

And a server that stopped answering was taken for one that had stopped. Any
fetch error counted as a closed port, a timeout included, and the server was
dropped from the last-resort cleanup before anything had been confirmed. With
the pnpm wrapper gone but a child still holding the port, the script exited 0
and left two servers running. Teardown now waits for the whole process group to
exit — escalating to SIGKILL if it outlives SIGTERM — keeps the server
registered until then, and treats the port as free only when a connection is
refused, since a hung server still completes the handshake.

`scripts/perf-compare.test.sh` drives the real script with a fake pnpm — no
build, no browser, 9 seconds — through success, a scenario that writes nothing,
a stale report, `ok` followed by failure, a child that ignores SIGTERM, and an
unknown ref. Each of the three new cases fails against the previous script for
the reason above. It runs in `frontend-build` like the other script tests;
the browser comparison itself stays manual.

`PERF_STOP_GRACE_MS` sets the SIGTERM grace (default 10s) so the test can
exercise the SIGKILL path without waiting on it.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-09-11 13:30:21 +08:00
378e3bd24b MUL-7154: unify the global radius system (#8175)
* fix(ui): unify global radius tokens

Co-authored-by: multica-agent <github@multica.ai>

* fix(ui): reduce global radius scale

Co-authored-by: multica-agent <github@multica.ai>

* fix(ui): address radius review findings

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-09-10 21:45:46 +08:00
2ed2432a2e fix(dev): recognize nested listener process groups (#8176)
Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-09-08 18:54:25 +08:00
6a65922ef7 MUL-7016: Add read replica infrastructure (#7979)
* feat(server): add read replica infrastructure (MUL-7016)

Co-authored-by: multica-agent <github@multica.ai>

* fix(server): make replica fallback passive (MUL-7016)

Co-authored-by: multica-agent <github@multica.ai>

* fix(server): address replica routing nits (MUL-7016)

Co-authored-by: multica-agent <github@multica.ai>

* refactor(server): tighten replica routing API (MUL-7016)

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-09-03 18:54:53 +08:00
7c1d826a98 chore(metrics): remove database-sampled metrics (#7819)
Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-31 17:06:17 +08:00
mikyan fe59262188 feat(agent): add Huawei Cloud CodeArts support (#6985)
* feat(agent): add Huawei Cloud CodeArts support

* fix(migrations): resolve CodeArts prefix collision

* fix(agent): address CodeArts review feedback

* fix(agent): tighten CodeArts integration scope

* fix(agent): resolve CodeArts follow-up review blockers

* fix(agent): redact CodeArts launch arguments

* docs: sync supported runtime counts

* fix(agent): own CodeArts model probe process tree
2026-08-31 12:30:59 +08:00
64ec7f5416 MUL-6737: fix(timeouts): close the two gaps the 2h inactivity budget left open (#7699)
* fix(timeouts): close the two gaps the 2h inactivity budget left open (MUL-6737)

Two independent places where a timeout still answered the wrong question.

Codex runs a semantic-inactivity timer inside the app-server protocol, and
unlike the daemon's watchdog it cannot see that a tool is in flight: a
commandExecution that prints nothing for the window trips it anyway. At a
fixed 10m ceiling that timer killed a quiet test suite or an
output-buffering `docker build` long before the 2h daemon budget meant to
protect it, which made "we raised the budget to 2h" false for Codex users.
It now derives from the daemon budget, taking the larger of the idle and
tool windows because it stands in for both. A tool budget of 0 means "never
force-stop during a tool", which this timer cannot express, so it falls back
to the idle budget rather than running unbounded; with the watchdog suite
disabled entirely Codex keeps its own built-in default, as it always has.
The env override and flag are unchanged.

Queued tasks expired on a pure wall clock: queued longer than a TTL meant
failed. That conflated "nobody is coming for this" with "the queue ahead is
long", and only the first is a failure. MUL-6558 was the second — a
self-hosted runtime with low concurrency held its own queue past 2h and
healthy work died as queued_expired. The TTL knob added then only moved the
cliff. Expiry now keys on whether the owning runtime still proves liveness,
the same signal FailTasksForOfflineRuntimes already uses for
dispatched/running rows, so a daemon going down retires its queued and its
in-flight work on one clock instead of two. A busy runtime keeps its
backlog for as long as it takes. MULTICA_TASK_QUEUED_TTL is removed with
the wall clock it configured.

Heartbeat age is read directly instead of gating on runtime.status='online'
so a row stuck at 'online' with a dead heartbeat still releases its queue,
and rows with no usable runtime binding — runtime_id IS NULL from before
migration 251, whose CHECK landed NOT VALID, or a runtime row that is gone —
are expired on sight rather than stranded forever by a JOIN that matches
nothing.

Tests: the queued sweep test now runs both phases against the same rows,
asserting the 5h-old task survives while its runtime heartbeats (the
MUL-6558 regression) and only expires once that runtime goes silent.
Assertions are scoped to the test's own task ids, because a runtime-keyed
sweep legitimately also collects queued rows other tests left on the shared
fixture runtime.

Co-authored-by: multica-agent <github@multica.ai>

* fix(sweeper): give a queued task its own reconnect grace before expiring it

Review catch: keying expiry on runtime liveness alone was too aggressive in
one direction. Enqueue binds a task to agent.runtime_id without checking
that the runtime is up (task.go CreateAgentTask call sites), so for a
runtime that has already been dark for longer than the grace, the liveness
clause is satisfied the moment a new task is created — assigning an issue
to a laptop closed overnight would fail it inside one 30s sweep tick
instead of waiting for the machine to come back. That is the opposite of
what "reconnect grace" promises.

Require the row's own age as well: a queued task now waits a full grace,
counted from when it started waiting, before it can be given up on. A
heartbeating runtime still never expires its backlog, so the MUL-6558 fix
is unchanged. The age bound also stops the sweep re-evaluating the runtime
subquery against every queued row on every tick.

Also remove the deploy surfaces the previous commit missed — the knob was
deleted from the server but helm (values + configmap), the self-host
compose file, and the CI-run helm render assertions still advertised it, so
operators would have kept setting a variable nothing reads. Same for five
orphaned comment lines in .env.example describing the removed variable.

Correct the fail-closed comment on the unbound-runtime arms: they are not
reachable for "rows predating migration 251" — runtime_id was NOT NULL from
migration 004 until 251 replaced it with a CHECK that every insert and
update is verified against, and the migration-004 FK is ON DELETE RESTRICT.
The arms stay as defence against a future schema change, which is what the
comment now says.

Tests: a third phase covers the reported case — a task enqueued against an
already-dead runtime survives, then expires once it has waited a full grace
of its own. Phase 2's expectation is corrected accordingly: the row created
at now() no longer expires when its runtime dies, because it has not waited
its own grace yet.

Co-authored-by: multica-agent <github@multica.ai>

* docs(timeouts): state both queued-expiry conditions and correct the FK claim

Review follow-up on b567511.

The user docs still described queued expiry as a single condition — "fails
once the runtime has been silent past the reconnect grace" — which is the
pre-b567511 behaviour and skips exactly the case that commit fixed. A
reader assigning work to a machine that went offline last night would
conclude the task already qualifies to fail. All four locales of tasks,
daemon-runtimes and troubleshooting now state both conditions and why the
second one exists.

The environment-variables pages contradicted themselves: the env table had
been updated to say the Codex semantic budget is derived, while the
`multica config` supported-keys table further down still advertised a 10m
default. That table now describes the resolution rule instead of a value.

The schema comment added in b567511 named the wrong constraint.
agent_task_queue_runtime_id_fkey (migration 004) is ON DELETE CASCADE; the
ON DELETE RESTRICT one is agent_runtime_id_fkey on agent.runtime_id. The
conclusion is unchanged — a queued row cannot reference a deleted runtime —
but the mechanism is a cascade plus migration 251's explicit unbind of
history rows, not a restrict.

Also log the resolved codex_semantic_inactivity at daemon startup, for the
same reason tool_watchdog is logged: it is now a derived value, so without
it the effective Codex budget cannot be read anywhere until a timeout
fires. And correct a test comment that still claimed both non-exempt rows
fail — the fresh row now stays queued, which is the assertion right below it.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-28 18:00:35 +08:00
LinYushen 55fe792483 refactor(cloud): use URL-only internal connection (#7566)
* refactor(cloud): use URL-only internal connection

* refactor(entitlement): remove Cloud rollout action

* refactor(cloud): remove unused client tuning
2026-08-26 14:40:19 +08:00
LinYushen baf1bbf340 fix(billing): make Cloud seat enforcement automatic (#7542)
Replace the separate Fleet/capacity URLs and seat-capacity switches with the single managed MULTICA_CLOUD_URL connection. Strict seat admission turns on whenever Multica is Cloud-connected and fails closed when the capacity machine token is absent or invalid; the recovery worker starts only for a valid executor so live intents keep their retry budget during credential outages.

Split invitation throttling into a non-consuming precheck and post-reservation consumption. Persistent capacity_full and capacity_overcommitted rejections charge only the actor budget; transient failures charge none. Preserve every Cloud/proxy HTTP 429 as retryable with both Retry-After forms, defer outbox rows without spending attempts, and use Cloud's rate-limit scope so a workspace 429 delays only that tenant while global or unscoped 429s stop the current batch.

BREAKING: delete/unset MULTICA_FLEET_URL, MULTICA_CLOUD_FLEET_URL, MULTICA_CLOUD_FLEET_TIMEOUT, MULTICA_SUBSCRIPTION_CAPACITY_ENABLED, MULTICA_SUBSCRIPTION_CAPACITY_URL and MULTICA_SUBSCRIPTION_CAPACITY_WORKER_ENABLED before rollout. The startup guard rejects any non-empty value, so MULTICA_SUBSCRIPTION_CAPACITY_ENABLED=false still blocks boot. Configure MULTICA_CLOUD_URL, MULTICA_CLOUD_TIMEOUT and MULTICA_SUBSCRIPTION_CAPACITY_SERVICE_TOKEN instead; entitlements keep their independent MULTICA_ENTITLEMENT_POLICY_* configuration.

Companion: multica-ai/multica-cloud#57
2026-08-26 12:42:42 +08:00
LinYushen 60a55ec57b feat(billing): purchase a seat from the invite flow (#7538)
* feat(billing): purchase seats from full invite flow

* fix(billing): harden full-capacity invite flow

* fix(billing): explain overcommit with cloud seat occupancy

Cloud refuses a reserve on used + reserved, but the overcommit toast filled
members_over_capacity_description with the product member count. The most
common trigger is a first ledger snapshot whose members alone still fit
inside the purchased seats, so the message read "3 members but only 3
purchased seats" and did not explain the rejection.

Add occupancy_over_capacity_description, which names occupied seats,
purchased seats, members and pending invitations, and feed it from the
Cloud summary. members_over_capacity_description stays as-is: the billing
tab banner is genuinely member-count driven.
2026-08-25 17:19:46 +08:00
da7131ab59 MUL-6625: enforce prepaid member seat capacity (#7491)
* MUL-6625: enforce prepaid member seat capacity

Co-authored-by: multica-agent <github@multica.ai>

* fix(seat-capacity): address review findings

Co-authored-by: multica-agent <github@multica.ai>

* fix(seat-capacity): close retry and worker races

Co-authored-by: multica-agent <github@multica.ai>

* fix(seat-capacity): bound lock waits during rollback

Co-authored-by: multica-agent <github@multica.ai>

* fix(seat-capacity): pin the cross-repo confirm settlement budget

Multica Cloud must wait longer than this worker's whole retry budget before it
treats a member row as one its ledger never recorded; adopting a member whose
confirm is still retrying counts one person as two seats, and an over-estimate
is what asks a customer to buy seats they do not owe.

That budget was only implicit here — 10 attempts of exponential backoff plus a
tick and a lock wait each — and Cloud's horizon had been set from a single
retry. Extract the backoff into retryBackoff, expose SettlementBudget, and
assert the worst case (29m15s) stays inside Cloud's 45m horizon, so raising
maxAttempts or the backoff cap fails review here instead of widening the window
in which Cloud cannot quantify a ledger.

Co-authored-by: multica-agent <github@multica.ai>

* Revert "fix(seat-capacity): pin the cross-repo confirm settlement budget"

This reverts commit d0ec36f647.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-25 12:51:44 +08:00
8a5a6adbf7 feat(agent): add ZeroClaw as a native ACP runtime (MUL-6511) (#7351)
* feat(agent): add ZeroClaw as a generic ACP agent runtime

Add ZeroClaw (a Rust-based, single-binary generic agent CLI) as a
first-party runtime, driven over the standard ACP JSON-RPC transport via
`zeroclaw acp`, mirroring the Dim/QwenPaw/Grok ACP backend shape. Only the
binary, launch args, and blocked-flags policy are ZeroClaw-specific; the
backend reuses the shared hermesClient ACP transport and deliverable
tracker.

No provider-specific quirks (permission-preset injection, minimum-version
gate, or a separate authenticate step) are assumed without evidence — those
were verified against real Dim/Grok binaries and would need the same
hands-on verification against a real ZeroClaw binary before being added
here.

Backend: SupportedTypes/New/launchHeaders wiring, model discovery via
session/new's advertised catalog, daemon probe with
MULTICA_ZEROCLAW_PATH/MODEL overrides, AGENTS.md context delivery,
default-command-name registration, metrics label, and a migration widening
the runtime_profile.protocol_family whitelist.

Frontend: provider union, MCP-config support flag, and a placeholder
provider-logo mark (no official ZeroClaw brand asset exists yet).

Tests mirror the existing ACP backend conventions: fresh-session,
resume-via-session/load, resume-not-found, transient-load-error, blocked
launch args, timeout, and the shared ACP deliverable-boundary regression
suite.

* fix(agent): renumber ZeroClaw's migration to 378 to avoid a same-day collision

Both this branch and a concurrently-implemented Junie runtime branch
independently picked migration 377 off the same origin/main tip.
Renumbering ZeroClaw to 378 up front avoids opening two PRs from the
same author that claim the same migration number.

* fix(agent): align the ZeroClaw ACP backend with the real 0.8.4 runtime

The backend was written against an assumed vanilla ACP handshake and never
checked against a ZeroClaw binary. Driving a real zeroclaw 0.8.4 over stdio
shows five of those assumptions are wrong, and the test fake encoded the same
assumptions, so the suite stayed green against a contract the runtime does not
have.

- Drop the session/set_model request. It is not in ZeroClaw's dispatch table
  (-32601) and no ACP handler reads a model param at all, so it could never
  succeed; because its failure was fatal, any agent with a model configured
  failed before session/prompt was reached. The model comes from the ZeroClaw
  agent profile, so opt zeroclaw out of ModelSelectionSupported and stop
  reading MULTICA_ZEROCLAW_MODEL rather than leaving the UI offering an
  override that is silently ignored.

- Label token usage from initialize's _meta.zeroclaw.defaultModel, the only
  model identity the ACP surface exposes. session/new returns just
  {sessionId, workspaceDir}, so the two extractACPCurrentModelID reads were
  dead code and the model-catalog discovery they fed could only ever return an
  empty fallback while paying an ACP subprocess per catalog read. Return an
  empty catalog directly, as qwenpaw/mcode do.

- Resume through session/resume instead of session/load. Both restore the
  transcript into the agent, but load also replays every retained message back
  as session/update notifications, which fed the previous answer into this
  turn's deliverable. Add the acceptNotification turn gate every other ACP
  backend carries, so out-of-turn chunks are dropped even if a runtime flushes
  them early (#1997).

- Send agentAlias on session/new when the operator names one. ZeroClaw
  auto-selects an agent only when its config holds exactly one entry, so
  multi-agent installs failed with -32602 and had no way out: `zeroclaw acp`
  has no --agent flag, and passing one kills the process at argument parsing.
  The alias is consumed from custom_args and travels as a param; it is omitted
  when unset, because a hardcoded guess breaks the single-agent case that
  works today. The refusal now names both remedies.

- Stop advertising MCP support. No ZeroClaw handler reads params.mcpServers;
  acp_enable_mcp gates that agent's own mcp_bundles, not the client list, so
  MCP is operator-side only. Hide the MCP tab and warn on a stale mcp_config
  instead of failing the task.

Also teach isACPSessionNotFound about ZeroClaw's custom -32000
SESSION_NOT_FOUND, without which a stale session id was retried forever
instead of restarting fresh. The wording check still decides, so a transient
error under the same code stays non-rejecting.

Tests now encode the observed handshake — _meta default model, session/new
without a catalog, -32000 for an unknown session, -32601 for set_model — and
cover the replayed-history, alias and MCP paths.

* fix(agent): make ZeroClaw integration merge-safe

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): tighten ZeroClaw ACP fallbacks

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): preserve ZeroClaw retry semantics

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-24 14:39:08 +08:00
8d2da6b8db MUL-6479 feat(dev): make a local environment a named object with one verb per lifecycle step (#7325)
* MUL-6479 feat(dev): make a local environment a named object with one verb per lifecycle step

Local development had six entry points, thirty-four make targets and no way to
see or delete what they created. `make up` / `status` / `list` / `down` /
`destroy` / `gc` / `env-exec` replace that with one set of verbs, shared by
humans and agents, with `C=api,web,daemon,desktop` selecting components.

Three defects made the old flow untrustworthy rather than merely verbose.

1. The database the tooling created was not the one the application reached.
   `ensure-postgres.sh` creates through `docker exec`; when a native PostgreSQL
   owns 5432 the container never binds the host port, so the create lands in a
   server the backend never talks to. It printed "✓ PostgreSQL ready" and the
   next step died with `SQLSTATE 3D000`. On the machine this was written on the
   two servers had already diverged: 94 same-named databases, 69 more in the
   container with no counterpart. `up` now creates through DATABASE_URL and
   treats a successful `migrate up` — which pings that same string — as the only
   proof; failure diagnoses the port owner instead of reporting a missing
   database.

2. Identity was computed, not allocated. Ports, database names and CLI profiles
   all came from `cksum($PWD) % 1000`, and a collision was silent: the loser
   fails to bind while the winner keeps answering 200. Allocation now happens
   under a lock, probing from the same path hash so a checkout keeps its
   familiar numbers, and requires the slot to be both unregistered and free of a
   live listener. The manifest in ~/.multica/dev/envs/<name>/ is what every
   later command reads.

3. Creating an environment left no record of how to destroy it, so there was no
   destroy — only archaeology. `down` stops processes and keeps the data;
   `destroy` consumes the manifest: processes, database, profile, slot. `list`
   makes leaks visible and `gc` collects environments whose checkout is gone or
   whose TTL passed.

/health now reports pid, commit and started_at. A 200 alone proves something is
listening, which is exactly what let a failed restart keep serving the previous
build; `up` compares started_at against its own launch and refuses an answer
that predates it.

Two constraints are enforced rather than documented: the daemon is started from
a built server/bin/multica, never `go run` (the daemon re-execs its own path for
every task, and `go run` deletes that binary when the launcher exits), and
`C=daemon` under a .multica/daemon_task_context.json marker stops with that
explanation up front instead of spending a login on a refusal.

scripts/dev-env.test.sh covers the registry verbs against a throwaway
MULTICA_DEV_HOME with no services running, and pins the regression that made
`down` exit 1 after reporting success: on bash 3.2 a command substitution whose
function ends in a failing command aborts the script under `set -e`, and "no
process is listening" is that function's normal answer.

`make dev` is unchanged and still runs in the foreground.

Co-authored-by: multica-agent <github@multica.ai>

* fix(dev): close local environment lifecycle gaps

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-23 16:43:58 +08:00
Martin MontesandMartin Montes 5a80f802d3 MUL-6558 feat(server): make task queue expiry configurable via MULTICA_TASK_QUEUED_TTL (#7418)
Lets self-hosted deployments raise the queued-task expiry window above the built-in 2h default, so low-concurrency runtimes stop losing legitimately-pending work to queued_expired. Default behavior is unchanged. Covers .env.example, Compose, Helm values/ConfigMap, and the behavior docs in all four locales.

Co-authored-by: Martin Montes <mmontes11@users.noreply.github.com>
2026-08-23 12:56:06 +08:00
53e15b35c6 MUL-6502: fix startup recovery during PostgreSQL outages (#7372)
* fix(startup): retry transient database outages

Co-authored-by: multica-agent <github@multica.ai>

* fix(startup): address database retry review

Co-authored-by: multica-agent <github@multica.ai>

* test(startup): address retry review nits

Co-authored-by: multica-agent <github@multica.ai>

* fix(startup): preserve native pgx connect timeout

Co-authored-by: multica-agent <github@multica.ai>

* fix(startup): apply timeout fallback to services

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-21 14:33:04 +08:00
814ee0cbf7 docs: remove six finished implementation records (#7258)
* docs: remove the two implementation docs under docs/

The database index audit and the Telegram channel review were one-off
implementation records for MUL-6108 and the Telegram contribution; both
are finished work and neither is referenced by the product docs. Drops
the now-dangling review pointer from the Telegram handoff notes.

docs/assets/ stays — the READMEs load the logos from it.

Co-authored-by: multica-agent <github@multica.ai>

* docs: remove the Telegram integration handoff notes

HANDOFF-telegram.md was the implementation record for the Telegram
channel contribution and is referenced by nothing. Operator-facing setup
lives in the localized Telegram guides under apps/docs/content/docs/.

Co-authored-by: multica-agent <github@multica.ai>

* docs: remove three finished implementation leftovers

- apps/mobile/docs/markdown-renderer-research.md: 2026-05 renderer
  selection research and incident log. Its conclusions already live in
  markdown-rendering-adr.md and in the MD_FONT/MD_LINE comments; the HIG
  calibration source is folded into markdown-style.ts so the numbers keep
  their rationale.
- packages/views/locales/glossary.md: a redirect stub. The content moved
  to the docs-site conventions pages, which now say the file is gone
  instead of claiming a stub still points at them.
- scripts/audit-redundant-indexes.sql: the MUL-6108 one-off audit query,
  orphaned once its report was deleted — no caller in Makefile, check.sh
  or CI.

All pointers into the deleted files are updated; nothing dangles.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-20 17:33:29 +08:00
1ea40db1c4 fix(build): honor GOOS as a Make variable when naming build outputs (#7309)
#7271 derives the Windows `.exe` suffix from a parse-time `go env GOOS`,
which cannot see GOOS passed as a Make variable: `make build GOOS=windows`
still emits PE binaries named `bin/server`, `bin/multica`, `bin/migrate` —
the exact artifacts #7255 reported as unable to re-exec themselves. The
top-level `export` puts the Make variable in the recipe's environment, so
`go build` targets Windows while the name says otherwise.

Reading `$(GOOS)` first covers both invocation forms, and making the
assignment target-specific keeps the toolchain probe on `build` alone: a
global assignment is expanded on every make invocation — `export` expands
even a recursive one — so `make help` / `make clean` printed
`go: Command not found` on checkouts without Go.

Adds scripts/makefile-build.test.sh to pin both invocation forms, the
non-Windows names, and the no-Go case, wired into the backend CI job that
already runs the other Makefile-adjacent shell tests.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-20 16:24:22 +08:00
e1be1625cb MUL-5959: feat(agent): add Dim (DimCode) ACP runtime (#6675)
* feat(agent): add Dim (DimCode) ACP runtime

Add Dim (dimcode, the `dim` CLI) as a first-party agent runtime, driven
over the ACP (Agent Client Protocol) transport via `dim acp`.

Key integration points:
- New dim backend (pkg/agent/dim.go): spawns `dim acp`, performs the
  ACP initialize/session/new handshake, and raises the runtime's
  hardcoded read-only permission preset to full-access (plus agent mode)
  via session/set_config_option before the first prompt — without it every
  file write is denied by a capability rule. Model override uses
  session/set_model; the model catalog is read from session/new.
- Session resume is intentionally skipped: Dim binds sessions to the
  creating process, so a later process's session/load is rejected with
  "held by another process" even after a clean session/close. The backend
  always starts a fresh session and reports ResumeRejected so the daemon
  classifies the run correctly. A best-effort session/close keeps Dim's
  own session list free of orphaned entries.
- Registration: SupportedTypes whitelist, New() factory, launch header,
  daemon probe (MULTICA_DIM_PATH / MULTICA_DIM_MODEL), default agent
  command name, protocol_family CHECK migration (255), model discovery,
  MCP config support, metrics label, and runtime display.
- Docs/UI: provider logo, landing i18n (en/zh/ja/ko), providers and
  install-agent-runtime docs, README, CLI_AND_DAEMON, SELF_HOSTING.
- Tests: dim backend unit tests (fresh session + resume behavior), shared
  ACP deliverable case, logo and MCP-support tests.

* feat(views): use official DimCode app icon for Dim provider logo

Replace the placeholder inline SVG with the official DimCode desktop
client app icon (from /opt/DimAgent/resources/build/icon.png, resized
to 128px), imported as a static asset the same way the Qwen Code mark
is handled.

* test(agent): add real dim ACP smoke tests (agentintegration)

Two gated tests behind MULTICA_RUN_REAL_AGENT_SMOKE=1:

- TestDimRealACPSmoke: full end-to-end against real `dim acp` —
  initialize, session/new, set_config_option (permission/mode), prompt,
  and assert completed output.
- TestDimRealResumeRejected: pass a fake ResumeSessionID, assert the
  backend starts a fresh session (no session/load), and reports
  ResumeRejected=true so the daemon classifies the run correctly.

Verified against dimcode 0.3.2 and 0.3.8.

* fix: renumber dim migration 265 → 271 (upstream added 265-270)

* fix: renumber dim migration to 272 (upstream added 271)

* fix(agent): dim cross-run resume via session/load + review fixes

Address review feedback on the Dim ACP runtime:

- #1 Deliver runtime brief to Dim: add dim to the AGENTS.md provider
  list (verified dim 0.3.8+ reads AGENTS.md from the session cwd).
- #2 Cross-run session continuity: dim 0.3.10+ releases its per-process
  session lock ~5s after the owning process exits, so resume now goes
  through the standard ACP session/load (same as traecli/kiro/grok)
  instead of always starting fresh. A loaded session retains its
  permission/mode, so set_config_option runs only on fresh sessions.
  Add a two-run regression (TestDimRealCrossRunResume) proving run B
  recalls context established only in run A.
- #3 Process lifecycle races: the deferred cleanup now cancels the run
  context before cmd.Wait (a child ignoring stdin EOF no longer hangs
  Result delivery) with a bounded force-kill fallback; the success path
  waits for the final prompt notification with a grace window instead
  of a non-blocking read that could miss the last message.
- #4 Isolate model discovery by executable path (discoveryCacheKey).
- #6 Real smoke test now writes a sentinel file to prove full-access is
  effective; a best-effort session/close is sent before teardown.
- #7 gofmt.

* test(agent): add dim process-lifecycle regressions (review #3)

- TestDimCleanupKillsHangingChild: a child that ignores stdin EOF and
  SIGTERM after an early failure is force-killed within
  dimProcessWaitTimeout, so Result still closes.
- TestDimPromptMissingNotificationStillCompletes: when session/prompt
  returns without stopReason (onPromptDone never fires), the bounded
  final-notification wait falls through and Result completes instead of
  hanging.

* fix(migrations): renumber dim runtime_profile migration 272 → 273 (upstream added 272)

* fix(agent): review round 2 — process tree, quiescence, fail-closed, version check

Address all remaining blockers from the second review:

- #1 Require dim >= 0.3.10 (version check at initialize); bounded retry
  on session/load when the lock is not yet released instead of silently
  starting fresh; integration test no longer sleeps before resume.
- #2 Process-tree cleanup: configureProcessGroup + signalProcessGroup +
  waitProcessGroupGone; cmd.Cancel returns nil so we own all signalling;
  regression asserts the descendant PID is gone.
- #3 Notification quiescence: onActivity + waitForACPNotificationQuiescence
  (same as hermes); Dim equivalent of the late-final-notification test.
- #4 Permission fail-closed: set_config_option runs on BOTH fresh and
  resumed sessions; on config failure, session/close is sent before
  returning so a partially configured session is not resumed; regression.
- #5 Migration renumbered 273 → 274 (upstream added 273).
- #6 AGENTS.md mapping regression, two-executable cache isolation test,
  session/close unit assertion, smoke test now executes a command.
- #7 README.md + CLI_AND_DAEMON.md updated with 22 CLIs including Dim.

* chore: remove temporary review actions doc (not for PR)

* fix(migrations): renumber dim migration 274 → 310 (upstream added 274-309)

* review: fix retry break scope, dedup isACPSessionNotFound comment, update stale notes

* fix(agent): pass *exec.Cmd to signalProcessGroup/waitProcessGroupGone (upstream signature change)

* fix(agent): close session on set_model failure + resume regression test

Audit-found gaps from self-review:
- set_model failure now sends session/close before returning (same as
  set_config_option failure) — reviewer #4 said 'permission, mode, or
  model'.
- Add TestDimConfigFailThenResumeReestablishes: run A config fails →
  session/close sent → run B resumes and re-applies set_config_option
  (fail-closed), completing successfully.
- Fix two stale comments that contradicted the code (config block now
  runs on both fresh and resumed sessions).

* fix(agent): add WaitDelay, use labeled break in retry loop

Self-audit improvements:
- Add cmd.WaitDelay=10s for consistency with claude.go (hard backstop
  if a process somehow survives SIGKILL).
- Use labeled break (break loadRetry) so runCtx cancellation during the
  retry delay exits the for loop directly, not just the select.
- Update PR description: migration 273→310, set_config_option now
  re-applied on both fresh and resumed sessions.

* fix(migrations): renumber dim migration 310 → 313 (upstream added 310-312)

* fix(migrations): remove stale 313 dim migration (replaced by 314 after upstream added dsh at 313)

* fix: DSH compatibility + thinking levels + reviewer round 3

Systematic fix of all 39 items from the compatibility gap analysis:

Rebase + migration:
- Rebase to latest upstream/main (resolves all DSH conflicts)
- Migration 314_runtime_profile_add_dim (whitelist includes both dsh + dim)
- Remove stale 313 dim migration

DSH + Dim coexistence (provider lists, counts, docs):
- SupportedTypes, config.go, metrics labels: both dsh + dim
- i18n (4 files): count 23, lists include DSH + Dim + Oh-My-Pi
- README.md/README.zh.md: count 23
- CLI_AND_DAEMON.md: 23-row table
- environment-variables.mdx (4 langs): MULTICA_DIM_PATH/MODEL callout
- display.ts, mcp-support.ts, types/agent.ts, provider-logo.tsx: both dsh+dim

Thinking levels (was MISSING — dim supports thought_level):
- dim.go: call applyACPEffortOption after set_model
- thinking.go: add dim to acpCatalogThinkingProviders
- dim.go: retain sessionResult for effort option

Reviewer round 3:
- Version check fail-closed: empty/malformed → reject (not allow)
- Cache test through ListModels call site with fake executables
- set_model failure regression test
- Test companions: sidecar_manifest_test, runtime_config_test add dim

Code quality:
- agent.go: fix unreachable duplicate return (rebase artifact)
- dim.go: labeled break in retry loop, WaitDelay, session/close on set_model fail

* fix: deliverable test fake version + provider-logo rebase artifact

* fix(migrations): renumber dim migration 314 → 315 (upstream added 314_workspace_mcp_config)

* fix(migrations): renumber dim migration 315 → 319 (upstream added 315-318)

* fix: Windows process tree, MinVersions, retry tests, migration 327

Address all 5 remaining blockers from review round 4:

1. Windows process-tree ownership: use startOwnedProcessTree instead of
   cmd.Start(), releaseProcessGroup in cleanup defer. This attaches the
   child to a Job Object on Windows (no-op on Unix), ensuring descendants
   are captured and terminated.

2. Migration renumbered 319 → 327 (upstream added 319-326).

3. Add dim: 0.3.10 to MinVersions in version.go so the daemon registers
   old Dim binaries as offline and refuses triggers (defense in depth on
   top of the ACP agentInfo.version check in dim.go).

4. Session-lock retry contract tests: success after bounded retries,
   exhaustion without falling through to session/new. Fake script
   supports DIM_LOAD_HELD_N (held for N calls then succeed) and
   DIM_LOAD_HELD_ALWAYS (always held).

5. README.md:206 20→23 runtimes. PR description updated.

* test: add missing reviewer-requested regressions (round 4)

- Windows Job Object regression: dim_windows_test.go proves the Dim
  backend captures and terminates descendants via startOwnedProcessTree
  (Windows-only build tag; companion to TestStartOwnedProcessTree).
- MinVersions registration: add dim test cases to TestCheckMinVersion
  (0.3.10 ok, 0.3.9 rejected, invalid rejected).
- Retry cancellation: TestDimSessionLoadRetryCancelled cancels the
  context during the retry delay, asserts no session/new fallback.

* fix: sync upstream Command signature changes + restore dim additions

Upstream changed ListModels/discoverXxxModels/detectCLIVersion to accept
Command (struct with Path+Prefix) instead of string. The previous rebase
kept our old-signature versions, causing widespread compile failures.

Restored all affected files from upstream, then re-applied dim-specific
additions:
- agent.go: dim in SupportedTypes, New(), launchHeaders
- hermes.go: isACPHeldByProcess + -32002 in isACPSessionNotFound
- thinking.go: dim in acpCatalogThinkingProviders
- models.go: dim case in ListModels + discoverDimModels
- version.go: dim in MinVersions
- All dim test files preserved (dim_test.go, dim_integration_test.go,
  dim_windows_test.go, models_test.go dim cache test)

* fix: rebase to latest main, migration 327→341, resolve DetectVersion conflict

* fix: remove unused os/exec import in dim_windows_test.go

* fix(migrations): renumber dim migration 341 → 342 (upstream added 341)

* fix: use Command.exec instead of exec.CommandContext (upstream GH #7046)

Upstream's TestOnlyLaunchGoSpawnsRuntimeProcesses requires all backends
to build processes through Command.exec (launch.go), so a custom runtime's
fixed_args are carried into the subprocess. Dim was the only backend still
using exec.CommandContext directly.

* fix: launch prefix policy + remove empty test files + deterministic retry tests

Address all 4 findings from review round 5:

1. Add "dim": dimBlockedArgs to launchPrefixBlockedArgs so protocol-breaking
   flags (--help/--auth-setup/--remote) are filtered from fixed_args.
2. Remove 77 zero-byte *_test.go files accidentally added at repo root.
3. Make retry tests deterministic: TestDimSessionLoadRetryCancelled waits
   for the first session/load request before cancelling (no fixed sleep);
   TestDimSessionLoadRetryExhausted asserts exactly 4 load attempts.
4. PR description will be updated separately.

* fix: rebase to latest main, migration 342→343 (upstream added mcode at 342)

* chore: trigger CI

* chore: refresh PR

* fix: restore mcode that was lost during merge conflict resolution

Merge took 'ours' side which predated mcode. Restore mcode in:
- SupportedTypes, New() factory, launchHeaders
- agent_supported_types_test.go want map
- metrics/labels.go
- config.go defaultAgentCommandNames
- migration 343 up/down whitelists

* fix: restore mcode probe in agents_probe.go + agent-cli-command-names.txt

* fix: restore all remaining mcode references lost during merge

Files fixed:
- README.md: mcode row in runtimes table
- mcp-support.test.ts: mcode assertion
- config.go: mcode in Agents comment + error message
- runtime_config.go: mcode in AGENTS.md case
- version.go: mcode in MinVersions
- version_test.go: mcode test cases

Verified: no file has fewer mcode references than upstream.

* fix: restore mcode in environment-variables docs + CLI_AND_DAEMON ACP list

- environment-variables.{mdx,ja,ko,zh}: restored 'MiniMax Code' in the
  QwenPaw model-variable sentence
- CLI_AND_DAEMON.md: added MiniMax Code to ACP-family list + loadSession
  fallback description

Verified: every changed file now has >= mcode references vs upstream.

* fix(migrations): renumber dim migration 343 → 344 (upstream added 343)

* fix(migrations): renumber dim migration 344 → 348 (upstream added 344-347 plugin migrations)

* fix(migrations): update down migration comment 344 → 348

* chore: retrigger CI (flaky TestOpenclawDiscoveryCacheConcurrentPreparations — upstream test, passes locally x5)

* fix(migrations): renumber dim migration 348 → 352 (upstream added 348-351)

* fix: Dim launch-prefix regression test + version.go comment repair (review #6)

- Add TestDimLaunchPrefixFiltersBlockedFlags: proves allowed prefix
  reaches command before acp, and --help/--auth-setup/--remote/-h are
  stripped from Dim's launch prefix.
- Add dim to TestLaunchPrefixReachesACPFamilies family list.
- Repair version.go: dim's cross-run session/load comment was appended
  to mcode entry during a rebase conflict; restore two independent
  accurate comments.

* fix(migrations): update down migration comment 348 → 352

* chore: retrigger CI (backend-tests stuck 50+ min)

* fix(migrations): renumber dim migration 352 → 362 (upstream added 352-361)

* fix: use logAgentCommand for dim argv logging (upstream redaction requirement)

Upstream's TestOnlyLaunchGoLogsAgentCommandArgs enforces that all
runtime argv logging goes through Config.logAgentCommand so sensitive
args are redacted. dim.go was still using Logger.Info directly.

* fix(migrations): renumber dim migration to 370

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-20 12:54:46 +08:00
658b0b7d9f MUL-6406: fix(auth): fail fast in production on insecure JWT secret defaults (#7212)
* fix(auth): fail fast in production on insecure JWT secret defaults

* chore(deploy): require a strong JWT_SECRET in compose and docs

* fix(ci): make selfhost config test pass with required JWT_SECRET

The compose file's JWT_SECRET:? error message contained a colon-space
sequence that YAML parsed as a mapping indicator, breaking docker compose
config. Quote the openssl hint instead.

.env.example now ships an empty JWT_SECRET, so the test's docker compose
config calls fail interpolation; supply a throwaway secret from the
environment, which Compose lets outrank the env file.

* fix(ci): seed JWT_SECRET into the recipe .env in selfhost config test

The Makefile includes .env and bare-exports every variable to the recipe
environment, so the make-driven recipe cases clobbered the test's
JWT_SECRET export with the empty value .env.example ships. The stub's
`docker compose config` then failed interpolation with stderr discarded,
produced empty JSON, and crashed the node parser. Seed the throwaway
secret into the recipe .env, which both make and Compose load.

* docs(env): use a CSPRNG for the Windows JWT_SECRET generation hint

Get-Random is not cryptographically secure, so the Windows hint could
leave self-hosters with a predictable signing key. Point at
RandomNumberGenerator instead, per review.

* docs(selfhost): fix stale JWT secret comments

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-20 12:54:19 +08:00
801c671973 chore(assets): convert bitmaps to WebP, dedupe, and gate image size (MUL-6352) (#7149)
* chore(assets): convert web and docs bitmaps to WebP

The repo carried 21.7MB of PNG/JPG and zero modern formats. next/image
already serves visitors AVIF/WebP, but the raw masters are still paid for
on every clone, every Docker build context, and every deploy upload.

Re-encode with cwebp and cap the source width at the largest size any
layout can request: 1920px for the decorative landing backgrounds (they
render behind content under object-cover), 2640px for the landing hero
(2x its 1320px slot), 1600px for screenshots (2x the ~720px docs and
use-case column). EXIF goes with it — the four feature backgrounds were
144dpi print exports.

Encoder choice is per image rather than uniform. Most compress best as
lossy; two of the docs screenshots are flat palettized PNGs that encode
smaller with `-lossless -z 9`, so they use that instead. Screenshots are
q80 to keep UI text crisp, decorative art q72.

  apps/web/public    10.70MB -> 1.57MB
  apps/docs/public    7.47MB -> 2.44MB

Verified: 1:1 crops against losslessly-downscaled references show no
visible difference; all 266 image references in the tree resolve; docs
prerenders all 177 pages and the landing/use-case pages render with no
broken images.

Co-authored-by: multica-agent <github@multica.ai>

* chore(assets): drop the duplicated README hero image

docs/assets/hero-board.png was byte-for-byte identical to the docs site's
workspace-overview screenshot — 633KB stored twice. Only the two READMEs
referenced the copy, so point them at the canonical file instead.

Co-authored-by: multica-agent <github@multica.ai>

* chore(assets): losslessly recompress app icons

oxipng -o max --strip safe over the app icon PNGs: -640KB with pixels
untouched (verified by comparing lossless-WebP encodes of the before and
after of every file — all 16 identical).

Deliberately not switched to WebP and not regenerated: the prerendered
sizes under apps/desktop/build/icons/ exist because electron-builder's
auto-derivation silently shipped only the 1024x1024 source in the v0.2.31
.deb (#2515), and the PWA manifest and .icns/.ico containers want PNG.
Filenames, dimensions, and the electron-builder config are unchanged, so
packaging behaves exactly as before.

Co-authored-by: multica-agent <github@multica.ai>

* ci: gate pull requests on a 300KB image budget

Keeps the previous three commits from being undone one screenshot at a
time. A bitmap added — or grown — past 300KB fails the check until the PR
description carries an `Oversized image exemption:` line saying why.
Shrinking an oversized file always passes: the check compares against the
size at the base ref, not just the absolute size.

Path-filtered to PRs that touch a bitmap, so the full-history checkout it
needs to diff against the base is only paid for when it is relevant.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-19 21:27:07 +08:00
3460f9aeaf MUL-6342: enforce entitlement-backed autopilot quotas (#7194)
* feat(autopilot): enforce entitlement run quotas

Co-authored-by: multica-agent <github@multica.ai>

* chore(deploy): expose entitlement policy config (#7195)

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>

* fix(autopilot): address quota review findings

Co-authored-by: multica-agent <github@multica.ai>

* fix(autopilot): finish quota review nits

Co-authored-by: multica-agent <github@multica.ai>

* refactor(autopilot): simplify quota schema

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-19 15:51:26 +08:00
24d8c521a6 MUL-6364: make the LLM client retry budget explicit, wired, and validated (#7201)
* feat(llm): make the LLM client retry budget explicit, wired, and validated

Config.MaxRetries was a bare int that nothing outside tests ever set. The
production wiring in handler.New never passed it and no MULTICA_LLM_* env var
fed it, so the field was configurable in name only and every deployment ran on
the SDK default.

Its semantics were also inverted. Because the zero value and an explicit 0 are
the same int, "0" fell into the unset branch and silently produced 2 retries,
while a negative — clamped to 0 to dodge the panic in option.WithMaxRetries —
was the only value that actually disabled retries.

- MaxRetries becomes *int so unset, 0 (disabled) and N (exact) stay distinct.
  New always passes the budget to the SDK explicitly, including the default, so
  the enforced policy and the reported one cannot drift apart.
- Add MULTICA_LLM_MAX_RETRIES as the single configuration source. Non-numeric,
  negative and above-ceiling values fail the boot instead of being corrected
  into something that looks configured. The ceiling of 5 is a latency budget:
  backoff reaches ~21s at 6 retries while the internal callers of this layer
  time out after 8s and 20s.
- Report the effective policy at startup via Client.RetryBudget, whose fields
  are counts and a fixed enum so the line cannot carry a key or gateway URL.
- Document the four states in .env.example and the environment-variable docs,
  and record what the budget does not cover: the GenerateJSON parameter
  negotiation and a stream that breaks after ChatStream returns.

Tests pin every state by upstream request count, so the unset-versus-zero
collapse cannot come back unnoticed.

Refs MUL-6364, closes #7154

Co-authored-by: multica-agent <github@multica.ai>

* fix(llm): reject invalid retry budgets at the client boundary and wire self-host

Review found the retry budget still reachable in two invalid shapes, and never
reachable at all on the main self-hosted deployment path.

- docker-compose.selfhost.yml declares the backend environment as an explicit
  allowlist and mapped none of the MULTICA_LLM_* variables, so every value
  .env.example and the docs told a self-hoster to set was dropped before the
  container: the LLM layer read as unconfigured no matter what they wrote. Map
  all four, and assert propagation in scripts/selfhost-config.test.sh, plus a
  drift guard so a future MULTICA_LLM_* documented without a mapping fails the
  script instead of silently going nowhere.
- Config.MaxRetries took a *int, so a negative still reached llm.New, which
  warned and used the default instead. Warning is not silence, but substituting
  a different budget is still the implicit correction #7154 asked to remove.
  It now takes *RetryOverride, whose field is unexported and whose only
  constructor rejects negatives — the invalid state is unrepresentable, so New
  has no correction branch left to take.
- The budget is a ceiling, not a quota: only retryable failures consume it, and
  a success or the caller's deadline can end a call sooner. Say "at most N"
  rather than "exactly N" in the field docs, .env.example and all four
  environment-variable docs.

The negative-value test now asserts rejection instead of pinning the fallback
as contract.

Refs MUL-6364, closes #7154

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-19 15:00:56 +08:00
41ddf6c0a1 refactor(frontend): drop dead UI components and phantom deps, add knip (MUL-6353) (#7147)
* refactor(frontend): drop dead UI components and phantom deps, add knip

`.npmrc` sets shamefully-hoist=true, which hides both halves of dependency
drift: an undeclared import still resolves, and an obsolete declaration still
installs. Neither build, typecheck, nor lint reads the import graph, so
package.json stopped being the source of truth for what is actually used.

Removed 16 zero-reference components from packages/ui (1,907 lines, all
`pnpm ui:add` output that never got wired up) plus mention-hover-card.tsx,
which was written but never imported — member and agent mentions render as a
plain `.mention` span. The comment in issue-hover-card.tsx claiming otherwise
was stale; corrected to match rich-content.tsx.

Deleting pagination.tsx orphaned its two i18n keys, dropped from all four
locales and from the `ui` namespace type.

Dependency declarations removed where nothing imports them:
- vaul, embla-carousel-react — only the deleted drawer/carousel used them
- date-fns, @tanstack/react-query-devtools — zero imports repo-wide
- motion, tw-animate-css from packages/ui — real users are views/desktop/web
- 44 of apps/web's 58 declarations, including all 15 @tiptap/* entries.
  One had already drifted: apps/web pinned tailwind-merge ^3.5.0 against the
  catalog's ^3.4.0, a second source of truth for the same package.

Added @types/hast to packages/views, which imports it directly.

knip enforces this going forward, scoped to files/dependencies/unlisted —
`exports` is unusable here because ui and views are consumed through per-file
exports maps. The CI step runs in warn mode (continue-on-error) and reports 9
pre-existing findings left out of this cleanup; clear those before promoting
it to blocking.

Verified: build, typecheck, lint green across web/desktop/ui/core/views.
6,503 tests pass; the 9 failures are pre-existing on a clean tree (a local
jsdom localStorage issue). Web dev server boots and serves /login 200.

Co-authored-by: multica-agent <github@multica.ai>

* fix(ci): cover the packages/ui blind spot knip cannot see

Review caught that knip provided no regression protection for the exact
thing the preceding cleanup was about. knip derives entry points from a
workspace's package.json `exports` — graph/build.js calls
`getEntrySpecifiersFromManifest` unconditionally, with no config flag to opt
out — and packages/ui exports four wildcards. Each expands to a glob over its
whole directory, so every file under components/ui, components/common,
markdown, and hooks registers as an entry point and can never be reported
unused. A zero-reference probe component confirmed it: knip's output was
byte-identical with and without it.

`entry: []`, negated `entry` patterns, `--production`, and `--strict` were all
tried. Manifest entries come from a separate code path that none of them
subtract from. packages/views is unaffected because 44 of its 45 exports name
a specific file, which is why knip does flag dead files there.

Added scripts/check-ui-wildcard-exports.mjs to cover those four directories,
blocking from the start since it reports nothing once the one file it found is
gone. It reads the wildcards from the manifest, so narrowing an export hands
that directory back to knip without the two overlapping. Its test asserts all
three states: clean tree passes, a stranded component fails and is named, and
importing that component clears it again.

The check found packages/ui/hooks/use-auto-scroll.ts — 80 lines whose only
occurrence repo-wide was its own definition. Same category as the 16
components already removed, and invisible to knip for the same reason.

Also from review:
- The path filter missed knip.jsonc and apps/docs/**, so config-only and
  docs-only PRs skipped knip even though it analyzes docs.
- The knip step paired `continue-on-error` with `|| true`, which zeroed the
  exit code before Actions saw it — no warning surfaced, only log text.
  Dropping `|| true` makes the warn phase visible and the later promotion to
  blocking a one-line change.
- Corrected the knip.jsonc comment that claimed apps were the only entry
  points for all three shared packages. True for views and core, never true
  for packages/ui.

Verified: build, typecheck, lint green (12/12, 0 errors). 6,503 tests pass;
the same 9 pre-existing failures as before, reproduced on a clean tree.

Co-authored-by: multica-agent <github@multica.ai>

* fix(ci): resolve import paths in the ui export check instead of matching names

Review found the check passed a dead component whenever any file elsewhere in
the repo imported a same-named module. The relative-import branch was a
whole-repo text regex on the basename, so `../components/option-card` in
packages/views vouched for an untouched packages/ui/components/ui/option-card.

That is not a corner case: 17 of the 59 guarded files already share a basename
with another file in the repo — `button`, `card`, `input`, `label`, `avatar`,
`tabs`, `switch`, `skeleton`, `separator`, `theme-provider` among them. shadcn
names are generic by construction, so the hole covered the most likely names a
future `pnpm ui:add` would produce.

Specifiers are now resolved against each importing file's own directory and
compared against real paths, with `@multica/ui/...` mapped through the same
manifest wildcards the candidate list is built from.

Liveness is now reachability from outside the guarded directories rather than
a reference count, so two dead components importing only each other no longer
vouch for one another.

Also implements the exemption the failure message promises: a file named by a
non-wildcard export is an entry point in its own right and is skipped. The
candidate loop previously read the wildcard directories unconditionally, so
following that advice would not have cleared the check.

Tests cover all five states: clean tree, stranded component named, cleared by a
real import, the option-card basename collision, and a mutually-importing dead
pair reported as two files.

Verified: check and its tests pass, knip unchanged at 8 files + 1 dependency,
build/typecheck/lint 12/12 green.

Co-authored-by: multica-agent <github@multica.ai>

* fix(ci): parse imports with the TypeScript parser, not a source regex

Review found the scanner counted commented-out imports as real edges. It ran
a regex over raw source, so a component kept its liveness from a line that
does not execute — and "delete the last real import, leave the comment
behind" is a normal step in a refactor, which makes the false edge appear at
exactly the moment the component stops being used.

Edges now come from real syntax nodes: static import/export declarations,
`import()`, `require()`, and `import("x")` type nodes. `vi.mock("...")` is
excluded on purpose — it names a module without importing it, so a component
whose only remaining mention is a test mock is dead, and counting it would let
one outlive its last real use.

Scope, measured rather than assumed: 52 specifiers in this repo appear only in
comments or prose, spread over 43 files, so the mechanism is live. None of
them currently resolve into packages/ui, so no dead component is being missed
today — this closes a latent hole, not an active miss. (An earlier count of 3
was wrong: those were `lazy(() => import(...))` calls that my throwaway
measurement script failed to treat as real.)

Test gains a sixth state: a component referenced only by a commented-out
import and a string containing one must still be reported.

Verified: six-state test passes, clean tree passes in 0.9s over 2,089 files,
knip unchanged at 8 files + 1 dependency, build/typecheck/lint 12/12 green.
`typescript` is already a root devDependency, so no new declaration.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-18 22:29:33 +08:00
a1cbd7c2ef MUL-6354: tighten build and generation guardrails (#7150)
* chore(build): tighten repository hygiene guardrails

Co-authored-by: multica-agent <github@multica.ai>

* chore(ci): remove snapshot ignore check

Co-authored-by: multica-agent <github@multica.ai>

* chore(ci): address repository hygiene review nits

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-18 18:32:19 +08:00
de902a4aab MUL-6273: feat(agent): add MiniMax Code ACP runtime (#7035)
* feat(agent): add MiniMax Code ACP runtime

* fix(agent): terminate mcode process trees

* fix(agent): address mcode review blockers

* test(agent): keep mcode model contract API-stable

* fix(migrations): resolve mcode prefix collision

* fix(agent): launch mcode through runtime command

Assisted-by: MiniMax Code

* fix(migrations): renumber mcode migration

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-18 12:44:27 +08:00
d0bd6683f1 MUL-5733: keep Go security patches current (#6409)
* fix(release): keep Go security patches current

Co-authored-by: multica-agent <github@multica.ai>

* fix(release): address vulnerability review

Co-authored-by: multica-agent <github@multica.ai>

* fix(release): restrict vulnerability bypass

Co-authored-by: multica-agent <github@multica.ai>

* fix(release): require Go 1.26.6

Co-authored-by: multica-agent <github@multica.ai>

* docs: align Go prerequisite with security floor

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
2026-08-17 16:18:53 +08:00
ZIce 3be6bccec7 MUL-6189 docs: align contributor Node version guidance with CI (#7003)
Aligns CONTRIBUTING.md, README.md, README.zh.md, SELF_HOSTING_ADVANCED.md and scripts/dev.sh with the Node 22 / pnpm 10.28.2 / Go 1.26.1 versions used by CI, and adds .nvmrc plus engines.node so version managers and package tooling flag Node 20 early.

Closes #6972
2026-08-17 15:25:31 +08:00
YYClaw e9528c722a MUL-5855: feat(dev): add worktree database cleanup (#6560)
* feat(dev): add worktree database cleanup

* fix(dev): handle database cleanup cancellation

* fix(dev): protect nested current worktree
2026-08-14 15:40:34 +08:00
Jiayuan Zhang 4f10a944b1 feat(runtime): add DeepSeek Harness support (#6923)
* feat(runtime): add DeepSeek Harness support

* docs(runtime): reflect public DSH release
2026-08-13 23:18:17 +08:00
cb8ce085ab MUL-6108: 清理首批冗余数据库索引 (#6878)
* perf(db): drop redundant indexes (MUL-6108)

Co-authored-by: multica-agent <github@multica.ai>

* fix(db): harden concurrent rollback retries (MUL-6108)

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Sol-Boy <sol-boy@multica-ai.local>
2026-08-13 13:51:15 +08:00
mxbao063-baoandmichelle 02b73d5874 fix(selfhost): forward SMTP_FROM_EMAIL to backend (#6653)
Co-authored-by: michelle <michelle@michelledeMacBook-Pro.local>
2026-08-11 15:54:51 +08:00
feigandBohan-J 3dd128532c feat(agent): add oh-my-pi (omp) as a supported agent runtime (#6514)
* feat(agent): add oh-my-pi (omp) as a supported agent runtime

oh-my-pi (omp, https://omp.sh) is a separate CLI that is a drop-in fork
of pi and speaks the same JSON event protocol. It is registered as an
independent provider so a host with both pi and omp installed gets two
runtimes instead of one.

omp is a runtime identity on the pi protocol, NOT a new protocol family.
It is declared in a single BuiltinRuntimes descriptor
(server/pkg/agent/builtin_runtimes.go) that carries id, protocol_family,
default_command, env_prefix, display_name, skills_dir, user_skills_dir,
launch_header, backend overrides, and model_discovery strategy. Every
consumer (agents_probe.go, config.go, daemon.go, execenv/context.go,
local_skills.go, agent.go New/LaunchHeader, models.go ListModels) derives
from this descriptor — no per-file hardcoded omp special cases.

NewRuntime() is the typed entry point for runtime identities: it resolves
id → descriptor → protocol family backend → applies overrides through a
backendOverrideApplicator interface (piBackend implements it). New()
delegates to NewRuntime() for built-in runtime identities, keeping New()
meaning exactly one thing: the protocol-family factory.

Model discovery: parseOmpModels reads the real {"models":[...]} wrapper
with separate provider/id/selector/name fields. Model.ID is the selector
(provider/id), matching parsePiModels' convention so buildPiArgs emits
both --provider and --model — not just --model. A regression test
(TestOmpSelectorSurvivesToBuildPiArgs) pins this end-to-end.

Frontend: provider-logo.tsx maps omp to PiLogo. The transcript dialog uses
the shared providerDisplayName helper. Landing copy updated to 21 tools
in all 4 locales.

Tests (omp_test.go): descriptor dispatch, NewRuntime entry point, binary
resolution, omp-labeled errors, event-stream completion, pi+omp
side-by-side registration, selector→buildPiArgs regression, model parser
coverage (real JSON/empty wrapper/invalid stderr/duplicate), and a guard
that every descriptor field is non-empty and consumed.

Closes #3989

* fix(views): keep Claude Code label in the transcript run details

Switching the provider row to the shared runtime formatter renamed every
Claude run: the daemon has no display-name override for claude, so the
shared formatter answers "Claude", and the legacy claude-code value
title-cased into "Claude-code". Both read as "Claude Code" before.

Keep those two aliases in a table local to this view and defer everything
else to the shared formatter, so the row stays in lockstep with the
runtime list (#5260) without renaming Claude runs.

* test(daemon): assert pi and omp both reach the register payload

The existing omp test stopped at probeAgentCLIs, so nothing covered the
part a user actually sees: whether both runtimes are registered with the
server, each under its own display name. Discovery finding two entries and
New() building two backends both stop short of version detection and the
registration payload.

Add a daemon-level test that runs real discovery off a fake PATH, drives
syncWorkspacesFromAPI, and asserts the payload carries type=pi and
type=omp with names "Pi" and "Oh-My-Pi". Rename the discovery test to
say what it covers, and record the registered display name in the batch
fixture so the name can be asserted.

* docs(agent): correct descriptor comments that describe the old behaviour

Three doc comments still described behaviour the fail-closed rework
replaced:

- ModelDiscovery claimed a nil strategy falls back to the family's
  discovery; ListModels deliberately returns an empty catalog instead,
  because omp exits non-zero on pi's --list-models.
- backendOverrideApplicator claimed backends without it are returned
  unchanged; NewRuntime returns an error rather than dropping the
  descriptor's executable and label.
- ProtocolFamily credited New() with applying the ID-specific defaults;
  that moved to NewRuntime.

Also repair the truncated first sentence of the NewRuntime doc, and fix
discoverOmpModels still calling the output a JSON array when
parseOmpModels right below it documents the {"models":[...]} wrapper.

* style(server): gofmt the files this branch touched

Four files were left unformatted: the descriptor literal and the omp test
struct lost their key alignment, execenv/context.go had its new import out
of order, and local_skills.go kept the old indentation after its switch
moved inside an if/else. Whitespace only — no behaviour change.

---------

Co-authored-by: Bohan-J <bhjiang@outlook.com>
2026-08-10 16:34:07 +08:00
b98f760f1f MUL-5604: 支持 Reasonix runtime (#6370)
* feat(agent): add Reasonix runtime

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): make Reasonix sandbox host-adaptive

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 14:19:28 +08:00
6b5d26b235 feat: add QwenPaw ACP backend support (MUL-5355) (#5986)
* feat: add QwenPaw ACP backend support

Add qwenpaw as a supported agent backend. QwenPaw runs via `qwenpaw acp`
over stdio using the ACP (Agent Client Protocol) JSON-RPC 2.0, reusing the
hermesClient transport layer.

Changes:
- New backend: pkg/agent/qwenpaw.go (ACP stdio via hermesClient)
- Register in agent.SupportedTypes, New(), launchHeaders
- Daemon config: probe qwenpaw binary, QWENPAW_ARGS env var
- Display name 'QwenPaw', inline system prompt support
- Runtime config: AGENTS.md injection, .qwenpaw/skills/ discovery
- User-level skills via QWENPAW_HOME env var
- DB migration 224: add qwenpaw to protocol_family whitelist

* fix(ci): add qwenpaw to CLI guard names and rename migration 231

- Add 'qwenpaw' to scripts/agent-cli-command-names.txt (TestAgentCLIGuardCoversDefaultCommands)
- Rename migration to 231 to avoid collision with existing migrations (TestMigrationNumericPrefixesStayUniqueAfterLegacySet)

* fix: address Bohan-J review comments on qwenpaw backend (PR #5986)

Blocking fixes:
1. config.go: Add qwenpaw CLI probe so daemon discovers the binary
2. models.go: Add qwenpaw case returning empty list (same as qwen)
3. qwenpaw.go: Fix resume path — use session/load (QwenPaw implements
   load_session, not session/resume), use resolveResumedSessionID,
   only set resumeRejected on session-not-found errors, fail the run
   on set_model failure instead of silently continuing
4. daemon.go: Add qwenpaw→"QwenPaw" to runtimeDisplayNameOverrides
5. agent.go: Update error message to list qwenpaw

Non-blocking fixes:
6. qwenpaw_test.go: Add 10 tests covering session/new, session/load,
   session/load not-found, set_model failure, ListModels, blocked args,
   protocol verification (session/load not session/resume), timeout,
   usage tracking, and backend construction
7. sidecar_manifest_test.go: Add qwenpaw to allFileBasedProviders

* fix: address Bohan-J second-round review on qwenpaw backend (PR #5986)

Blocking fixes:
1. qwenpaw.go: Send qwenpaw.coding_project_dir inside _meta, not as
   top-level parameter — ACP Pydantic model ignores extra root-level
   keys (model_config has no extra='allow'), only merges field_meta
   into kwargs
2. qwenpaw.go: Handle set_model returning result:null (not RPC error)
   — real QwenPaw server swallows exceptions and returns None, which
   serialises as JSON null; json.RawMessage('null') is non-nil so
   bytes.Equal check is needed
3. context.go: Workspace skills are at <workDir>/skills, not
   <workDir>/.qwenpaw/skills (confirmed against QwenPaw v2.0.1)
4. local_skills.go: QwenPaw resolves global root from
   QWENPAW_WORKING_DIR -> COPAW_WORKING_DIR -> ~/.copaw -> ~/.qwenpaw,
   not from QWENPAW_HOME

Test additions:
5. TestQwenpawSetModelReturnsNull — regression test for result:null
6. TestQwenpawSessionLoadTransientError — transient error does not
   set ResumeRejected=true
7. TestQwenpawSessionNewSendsCodingProjectDir — verifies _meta format
8. TestQwenpawSessionLoadSendsCodingProjectDir — same for session/load
9. TestQwenpawTimeout — deterministic sync via signal file

All 15 Qwenpaw tests pass. Verified against real qwenpaw v2.0.1 binary
(end-to-end: prompt, session ID, resume, system prompt).

* fix: rebase against upstream/main and add per-task qwenpaw workspace/agent isolation

- Rebase add-qwenpaw-backend branch on upstream/main (88 commits ahead)
- Resolve conflict in config.go: keep refactored probe() with shell
  resolution fallback, which already includes qwenpaw probe
- Add --workspace and --agent CLI args to qwenpaw acp for per-task
  skill isolation and agent identity isolation
- Mark --workspace and --agent as blocked in qwenpawBlockedArgs so
  user custom_args cannot override them
- Add deriveQwenpawAgentID() to produce deterministic agent IDs
  from task's issue ID and agent ID
- Wire QwenpawWorkspace and QwenpawAgentID through ExecOptions
- Add prepareQwenpawWorkspace() in execenv to materialize bound
  skills into a per-task workspace directory

* fix: address Bohan-J third-round review on qwenpaw backend (PR #5986)

Three blocking issues resolved:

1. Agent ID registration — remove session/set_model and --agent entirely.
   The simpler route: QwenPaw model override is declared unsupported,
   eliminating the need for agent profile registration in QwenPaw config.

2. Skill revocation — prepareQwenpawWorkspace now does os.RemoveAll
   on the skill_pool dir and manifest before rebuilding, making it
   idempotent. A->empty (revoke all), A->B (replace), and A->A
   (repeated reuse) all work correctly. Added 3 new unit tests.

3. Skill root path — changed 'skills' to 'skill_pool' in both
   prepareQwenpawWorkspace and skillsDirPath to match QwenPaw's
   store.py get_workspace_skills_dir.

Also cleaned up: removed deriveQwenpawAgentID function and its test,
removed QwenpawAgentID from ExecOptions, removed --agent from
qwenpawBlockedArgs, removed session/set_model test cases.

* fix: add missing qwenpaw probe() call in agents_probe.go

The TestDefaultAgentCommandNamesCoversAllProbes test found only 17
probe() calls in agents_probe.go but defaultAgentCommandNames has 18
entries. The probe() call for qwenpaw was missing, causing the
backend CI test failure.

Adding the probe ensures GUI-launched daemons can resolve qwenpaw
via the login shell fallback, matching all other providers.

* fix: sync models.go with upstream/main (remove qwenpaw from ListModels)

* fix: CI failures — migrate prefix 235→236 + update test

- migration prefix 235 was reused by upstream 235_chat_message_quick_actions;
  renamed our qwenpaw migration from 235 to 236
- TestQwenpawListModels called len() on Catalog struct (compile error);
  fixed to expect error for unknown provider type

* fix: bump qwenpaw migration prefix 236->241 (upstream took 236)

* fix: qwenpaw local skill root — 'skills' → 'skill_pool' (MUL-5355)

Bohan-J third-round review blocker #3: local_skills.go still scanned
<QWENPAW_HOME>/skills but the QwenPaw shared skill pool is
<QWENPAW_HOME>/skill_pool (store.py get_workspace_skills_dir).

Verified no other qwenpaw paths in the codebase assume the wrong layout.

* fix: bump qwenpaw migration prefix 241->242 (upstream took 241)

lint test fails with: migration prefix 241 is reused by
[241_comment_parent_lookup_index 241_runtime_profile_add_qwenpaw].
Upstream added 241_comment_parent_lookup_index; bump our migration.

* feat: add QwenPaw integration test and version declaration (MUL-5355)

- New: server/pkg/agent/qwenpaw_integration_test.go with three
  agentintegration build-tagged tests:
    TestQwenpawRealACPSmoke — full end-to-end ACP smoke test with
    session/new → session/prompt and session/load resume validation.
    TestQwenpawRealWorkspaceSmoke — validates skill_pool workspace
    flag handling and per-task skill isolation.
    validateQwenpawVersion — attempt version detection via qwenpaw
    --version, pip show qwenpaw, and python import.

- Document QwenPaw v2.0.1 as the supported baseline version in
  qwenpaw.go package comment, noting the contract details:
    _meta qwenpaw.coding_project_dir for Coding Mode,
    session/set_model NOT supported,
    skill_pool workspace layout.

- This test suite is gated by MULTICA_RUN_REAL_AGENT_SMOKE=1 and
  requires qwenpaw on PATH, matching the pattern used by grok,
  cursor, and traeecli integration tests.

* fix: bump qwenpaw migration prefix 242->243 (upstream took 242 for qoderclicn)

Upstream added 242_runtime_profile_add_qoderclicn in the same rebase
window, colliding with our 242_runtime_profile_add_qwenpaw. The
migration lint test TestMigrationNumericPrefixesStayUniqueAfterLegacySet
catches duplicate prefixes after the legacy range.

Also add 'qoderclicn' to our migration's CHECK constraint so it doesn't
regress the whitelist added by upstream's 242.

* temp: stub out integration test to isolate CI failure

* fix: restore upstream probeAgentCLIs() call in config.go (rebase regression)

Rebase conflict resolution accidentally reverted upstream's MUL-5439
refactor (extracting probe logic to agents_probe.go) back to the old
inline probe block. This also dropped qoderclicn detection and
duplicated the probe logic already in agents_probe.go.

Restore the single 'agents := probeAgentCLIs()' call — qwenpaw is
already probed in agents_probe.go.

* fix: bump qwenpaw migration prefix 243->251 (upstream took 243-250)

Upstream added migrations 243-250 since our last rebase. Bump to 251,
the next available prefix.

* feat: add QwenPaw integration test (agentintegration build tag)

TestQwenpawRealACPSmoke drives the real qwenpaw acp binary end-to-end:
  - session/new + session/prompt produces 'pong'
  - session/load resume with ResumeSessionID works
  - --workspace flag is forwarded correctly

Gated by MULTICA_RUN_REAL_AGENT_SMOKE=1, matching grok/cursor pattern.
Validated against QwenPaw v2.0.1.

* fix: workspace skills dir 'skills' not 'skill_pool' + add skill loading integration test

Two fixes:

1. qwenpaw_workspace.go: write skills to <workspace>/skills/ instead of
   <workspace>/skill_pool/. QwenPaw's get_workspace_skills_dir() looks
   for workspace skills at <workspace>/skills/ (store.py:65-67), not
   skill_pool (which is the shared pool at WORKING_DIR/skill_pool).
   Verified against real qwenpaw acp — skills in skill_pool/ were never
   discovered.

2. Add TestQwenpawRealWorkspaceSkill integration test that proves a
   bound skill is actually loaded and effective: writes a skill that
   overrides the agent's response, sends an unrelated prompt, and
   asserts the skill's marker text appears in the output. This
   addresses R3 review feedback: 'please also add a test that exercises
   an actually-bound skill'.

All three integration tests pass against QwenPaw v2.0.1:
  - TestQwenpawRealACPSmoke (session/new + prompt + session/load resume)
  - TestQwenpawRealWorkspaceSkill (skill discovery + effectiveness)

* feat: add ACP model discovery for qwenpaw via session/new models field

QwenPaw v2.0.1+ now includes a 'models' field (SessionModelState) in
the session/new response, added by agentscope-ai/QwenPaw#6531. This
lets ACP clients discover available models without session/set_model.

- ListModels for qwenpaw now uses discoverACPModels (same pattern as
  traecli/grok/kiro) to spin up 'qwenpaw acp', call session/new, and
  parse the models catalog from the response.
- discoverQwenpawModels mirrors discoverTraecliModels — ACP-native,
  no auth selection needed.
- Model override via session/set_model remains unsupported: it
  persists to agent.json at the agent scope (not session-scoped),
  so calling it would mutate the user's shared agent config. The
  model picker shows available models for display/selection, but
  the daemon does not send set_model.
- Updated TestQwenpawListModels to verify qwenpaw is a recognized
  type (not 'unknown agent type' error).

* fix: address Bohan-J Review 5 — ModelSelectionSupported=false, version bump to v2.1.0-beta.1, remove debug files

- ModelSelectionSupported('qwenpaw') now returns false with rationale
  (session/set_model persists to agent scope, not session scope)
- Add TestQwenpawModelSelectionUnsupported regression test
- Update version references from v2.0.1 to v2.1.0-beta.1 (includes
  agentscope-ai/QwenPaw#6531 — models field in session/new response)
- Remove check_ci.py, jobs.json, runs.json debug artifacts

* fix: bump qwenpaw migration prefix 251->253 (upstream took 251)

Upstream added 251_agent_runtime_unbind. Bump to 253, the next
available prefix after 252_agent_builder_draft.

* fix: address Bohan-J Review 6 — drop unused discovery, always attribute to unknown

- ListModels for qwenpaw returns empty catalog without spawning ACP
  subprocess (model selection is unsupported, so no consumer exists)
- Usage attribution always uses 'unknown' instead of opts.Model
  (the backend never sends opts.Model to QwenPaw)
- Add TestQwenpawUsageModelIgnored regression test
- Fix stale v2.0.1 comment in TestQwenpawListModels

* chore(agent): clean up qwenpaw model-discovery leftovers

Follow-up nits from review 7 on PR #5986:

- Drop discoverQwenpawModels: it lost its only caller when ListModels
  started returning an empty catalog for qwenpaw.
- Correct the version contract in qwenpaw.go. The execution path needs
  only the ACP surface present in v2.0.1 (current stable); the models
  field on session/new landed in v2.1.0-beta.1 but has no consumer now
  that model selection is unsupported.
- Make TestQwenpawListModels actually guard the no-subprocess promise.
  It pointed at a nonexistent path, which the old discovery helper also
  answered with an empty catalog, so it passed either way. It now uses
  an executable fake that records invocation; verified it fails if a
  discovery path is reintroduced.
- gofmt agent.go (ExecOptions alignment broke when QwenpawWorkspace was
  added) and restore the trailing newline in qwenpaw_test.go.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: niudakok <niudakok@users.noreply.github.com>
Co-authored-by: Bohan-J <bohan.optimism@gmail.com>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 18:47:10 +08:00
e4b6f7a31b MUL-5581: add Qoder CN CLI runtime (#6232)
* feat(agent): add Qoder CN CLI runtime

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): address Qoder CN review nits

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): defer Qoder CN version gate

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 16:01:53 +08:00
cac453fe01 MUL-5505: fix(selfhost): one source of truth for the backend host port (#6168)
* fix(selfhost): read the published host port back from Compose (MUL-5505)

`make selfhost` polled the wrong port when the backend host port was passed
through the environment. Make and Compose disagree on variable precedence: an
assignment in an included makefile (the Makefile `include`s .env) overrides a
real environment variable, while for Compose the environment overrides .env.
Because .env.example ships BACKEND_PORT uncommented, .env always pinned the
recipe's view of the port, so `BACKEND_PORT=9000 make selfhost` published the
stack on 9000 while the health check hammered 8080 for 60s and then reported
"Services are still starting" on a perfectly healthy install. The success
message printed the same wrong backend URL. FRONTEND_PORT had the same
inversion, affecting the printed frontend URL.

Stop re-deriving the host port in the recipe and ask Compose, which is the only
authority on what it published. The wait-and-report block (three port
references per target, duplicated across selfhost and selfhost-build) moves into
scripts/selfhost-wait.sh, which resolves both host ports via `docker compose
port` and falls back to the compose default chain when the stack is not up.

Also fix the input contract: `PORT` was the only port variable documented in
SELF_HOSTING_ADVANCED.md, yet compose read only BACKEND_PORT, so setting the
documented variable was silently ignored. The backend host port mapping now
falls back to `${PORT:-8080}`, matching Makefile and scripts/local-env.sh, and
the docs describe both variables as host ports.

Container-internal ports (8080/3000), the global PORT derivation used by local
dev, and the 127.0.0.1 bindings are unchanged; the default self-host experience
is byte-identical.

Fixes #6145

Co-authored-by: multica-agent <github@multica.ai>

* fix(selfhost): make PORT the real backend port and test the whole chain

Addresses review feedback on #6168.

The first round's root-cause claim was wrong. It said `BACKEND_PORT=9100 make
selfhost` published on 9100 while the health check probed 8080. It does not:
when make drives Compose, Compose inherits make's exported values, so both
sides agree on 8080 and the override is silently dropped instead. Re-measured
through the real recipe, the genuine disagreements were:

  - `.env` setting PORT with no BACKEND_PORT: probe followed PORT, Compose
    published ${BACKEND_PORT:-8080} — 9100 vs 8080.
  - `make selfhost PORT=8080` over a .env with BACKEND_PORT=9100: probe 8080,
    Compose published 9100.

Both are the "wrong port variable" defect the GitHub issue names, and reading
the port back from Compose fixes both. The comments and PR text that blamed
make/Compose precedence are corrected.

The PORT fallback added to docker-compose.selfhost.yml was also unreachable:
.env.example shipped BACKEND_PORT=8080 uncommented, so the fallback never
engaged and `SELF_HOSTING_AI.md`'s "edit PORT and FRONTEND_PORT" instruction did
nothing. Pick one contract and apply it everywhere: PORT is the backend port to
edit, BACKEND_PORT/API_PORT/SERVER_PORT are optional aliases that override it.
That is already what Makefile, scripts/local-env.sh, scripts/install.sh and
install.ps1 do; compose and the web dev fallback were the outliers.

  - .env.example ships BACKEND_PORT commented out, like its sibling aliases, so
    editing PORT works out of the box.
  - apps/web/config/runtime-urls.ts walks the same alias chain instead of
    reading BACKEND_PORT alone, so `pnpm dev` follows an edited PORT.
  - SELF_HOSTING_AI.md and SELF_HOSTING_ADVANCED.md state the contract.

Configuration that cannot take effect is now reported instead of ignored.
scripts/selfhost-preflight.sh warns when an alias in the env file overrides an
edited PORT, and when a port from the shell environment is overridden by the env
file. The Makefile captures the pristine origins before its include, because
afterwards there is no way to tell that `PORT=9000 make selfhost` was asked for.

Tests now cross environment -> make -> compose -> report for real. The docker
stub is no longer told what to publish: on `up` it asks the real
`docker compose config` to interpolate the compose file with the environment the
recipe actually handed it, records that, and answers `port` from the recording.
Verified failing on the old code — with the old recipe and compose file the
suite reports published 8080 against probe 9100, the exact defect.

Co-authored-by: multica-agent <github@multica.ai>

* ci: gate the self-host test on scripts/selfhost-preflight.sh too

The new preflight script was missing from the frontend path filter, so a
change to it alone would skip the test that covers it.

Co-authored-by: multica-agent <github@multica.ai>

* fix(selfhost): honour the full backend port alias chain in Compose

Addresses the second review round on #6168.

The contract this PR documents is

  BACKEND_PORT -> API_PORT -> SERVER_PORT -> PORT -> 8080

and Makefile, scripts/local-env.sh, scripts/install.sh, scripts/install.ps1 and
apps/web/config/runtime-urls.ts all implement it. docker-compose.selfhost.yml
implemented only BACKEND_PORT -> PORT -> 8080, so `API_PORT=9100` or
`SERVER_PORT=9200` in .env still published 8080.

That is not only a direct-Compose concern. scripts/install.sh:407 derives its
health-check port from the full chain and then runs this compose file, so an
alias it honours but Compose ignores puts the probe on 9100 while the stack
listens on 8080 — the original #6145 defect, reached through the installer.

  - docker-compose.selfhost.yml publishes the full nested chain, verified to
    keep the same precedence and the same 8080 default.
  - scripts/selfhost-wait.sh's fallback matches, so the degraded path agrees
    with what Compose would have published.
  - The Makefile captures the pristine origins of API_PORT and SERVER_PORT too,
    and the preflight reports them, so no alias can be silently overridden by
    the env file while the others are reported.

Tests pin the chain on the boundaries a recipe-level test cannot see, because
Makefile normalises all four aliases into PORT before any recipe runs: a matrix
over the direct Compose path asserts each alias moves the published port with
the right precedence, and every case additionally asserts that
scripts/install.sh's resolver returns the same value Compose published. That
invariant is the one that actually protects the installer.

Verified failing without the fix: with the two-level chain restored the suite
reports `expected 9200 / compose published 8080 / installer resolved 9200`, and
with the alias warnings narrowed it reports the missing API_PORT notice.

Co-authored-by: multica-agent <github@multica.ai>

* fix(selfhost): resolve one winner before reporting ignored port config

Addresses the third review round on #6168.

The preflight walked each alias independently, so it made a separate "wins"
claim per alias. With

  PORT=9000
  BACKEND_PORT=8000
  API_PORT=7000
  SERVER_PORT=6000

in .env it announced that BACKEND_PORT, API_PORT and SERVER_PORT each won, when
only BACKEND_PORT=8000 can. Two of the three notices were simply false.

It also compared only same-named shell and file variables, so shadowing across
aliases stayed silent: .env with PORT=8000 and BACKEND_PORT=8000 plus a shell
API_PORT=7000 produced no output at all, even though API_PORT was discarded.

Resolve the winner along the documented chain first — within a variable the env
file beats the environment, then BACKEND_PORT beats API_PORT beats SERVER_PORT
beats PORT — and report only the inputs that winner actually shadows. Output now
names one winning input and lists the rest, whichever variable or source they
came from. An input carrying the winning value is not reported: it is redundant
but the user still gets the port they asked for, so there is nothing to fix.

Tests cover both cases the old logic got wrong: every alias set at once must
yield exactly one winner claim and list the others as unused, and an env-file
alias beating a lower-priority shell alias must be reported. Also asserted that
a redundant same-value alias and a default configuration both stay silent.

Measured against the previous implementation, phrasing aside: with all four
values set it made 3 winner claims where there is now 1, and on the cross-alias
case it printed nothing where API_PORT is now reported.

Co-authored-by: multica-agent <github@multica.ai>

* fix(selfhost): track env-file existence so empty port assignments resolve right

Addresses the fourth review round on #6168.

The Makefile passed only values to the preflight, so the script inferred "the
env file does not set this" from an empty string. An explicit empty assignment is
a different input from an absent one: make lets `BACKEND_PORT=` in the env file
override `BACKEND_PORT=9000` from the environment, and the variable then drops out
of the alias chain instead of setting a port.

With .env holding PORT=8080 and BACKEND_PORT=, invoked as
`BACKEND_PORT=9000 make selfhost`, the stack published 8080 and the health check
probed 8080 — correct — while the preflight announced

  the backend host port resolves to 9000 from BACKEND_PORT (environment)
  Set but unused: PORT=8080 (.env)

naming the ignored value as the winner and the winning value as unused. A report
that contradicts the startup is worse than no report. FRONTEND_PORT= had the
matching hole: it shadowed the environment, Compose fell back to 3000, and the
preflight said nothing.

The Makefile now captures ENV_FILE_<VAR>_IS_SET from $(origin) alongside the
value, for all five port variables via one $(foreach) instead of hand-written
lines. The preflight treats a variable the file defines as taking the file's
value even when empty, so it shadows the environment, and an empty value never
wins — resolution continues down the chain and falls back to 8080. Inputs the
winner shadows are still listed, including ones shadowed by an empty assignment.
The frontend check mirrors this and now names the port that actually resolved.

Recipe-level tests cover an empty alias over a shell value, every link of the
chain emptied at once, and the frontend equivalent — each asserting the reported
winner, the Compose published port and the probed port all agree. Against the
previous implementation the first case fails with `reported: 9000 / actual: 8080`.

Co-authored-by: multica-agent <github@multica.ai>

* fix(selfhost): installers read the published port; drop the second parser

Addresses all four blockers from the consolidated review on #6168.

1. Both installers probed and printed the wrong port. scripts/install.sh:402 and
   scripts/install.ps1 start Compose inheriting the current environment, and
   Compose lets that environment outrank .env — but the installers then derived
   the probe port and the summary URL from .env alone. Measured on this branch
   before the fix, every port variable diverged:

     ambient PORT=9100        -> compose 9100, installer 8080
     ambient BACKEND_PORT=9200 -> compose 9200, installer 8080
     ambient API_PORT=9300     -> compose 9300, installer 8080
     ambient SERVER_PORT=9400  -> compose 9400, installer 8080
     ambient FRONTEND_PORT=3100 -> compose 3100, installer 3000

   Both now read the ports back once with `docker compose port` after `up -d`
   and reuse that single result for the health check and the summary. If the
   query fails they fail loudly instead of claiming success. The .env-derived
   helpers are deleted rather than left beside the new path.

2. The Make preflight is removed entirely, as the review recommended. It
   modelled `origin=environment` and `origin=file` but not `command line`, so
   `make selfhost BACKEND_PORT=9000` over a .env with API_PORT=7000 announced
   7000 while Compose published 9000. That was its fourth wrong report in four
   rounds, because it was a second parser of precedence rules — the very thing
   this PR removes elsewhere. `docker compose port` is the runtime truth, so the
   script, the Makefile captures and the macro are gone.

3. Tests now cover the installer paths that were unguarded. scripts/install.test.sh
   gains a `--with-server` matrix (defaults, .env PORT, all four backend aliases,
   FRONTEND_PORT, ambient overriding .env for all five, and explicit-empty
   fallback) plus a case proving an unresolvable port fails loudly. A new
   scripts/install.ps1.test.ps1 drives the same matrix through Start-LocalInstall.
   Every case asserts Compose's published port == the probed URL == the printed
   URL. Ambient variables are explicitly cleared per case so a runner's own PORT
   cannot leak. The Makefile matrix gains the command-line origin cases. CI runs
   the PowerShell suite on windows-latest in the always-on installer job, and the
   frontend filter now includes both installers.

4. Docs state the real contract: the alias order is shared, but which *source*
   wins is per entry point — Compose prefers the environment, make prefers the
   included env file, and a make command-line assignment outranks both. The
   removed preflight's promise is gone from SELF_HOSTING_AI.md, and the compose
   header no longer claims the installer derives its own health-check port.

Verified failing without the fixes: the Bash matrix reports
`[ambient PORT beats .env] compose published 9500 / probed 9100 / printed 9100`
against the old installer, and the PowerShell matrix reports
`printed backend=8080` against an injected summary divergence.

Co-authored-by: multica-agent <github@multica.ai>

* fix(selfhost): keep web dev proxy off frontend port

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 14:46:13 +08:00
YYClaw ee7ba83f53 fix(self-host): apply setup config to daemon (MUL-5269) (#5880)
Fixes two connected self-host setup failures in local worktree development.

Generated worktree environments now expose the backend HTTP origin through
MULTICA_PUBLIC_URL, and existing generated worktrees derive the missing value
at startup through both scripts/local-env.sh and the Makefile. Explicit
values, including an intentionally empty same-origin setting, are preserved.

setup and setup self-host now reconcile an existing same-profile daemon after
authentication so it loads the newly written server URL and token. An idle
daemon is restarted behind the existing restart preflight; when active tasks
exist, setup leaves the daemon running to avoid cancelling work and prints an
actionable profile-aware restart command instead.

Closes #5879
2026-07-30 15:10:40 +08:00
73b0015475 feat(vcs): make self-hosted Git providers self-host-only (MUL-3772, MUL-5138) (#5888)
* feat(vcs): gate self-hosted Git providers to self-host deployments only (MUL-3772)

The Forgejo/Gitea/GitLab integration is intended for self-hosted Multica, where
Multica can reach a Git instance on the operator's own network. On the managed
multi-tenant cloud it adds an SSRF surface (connect validates a user-supplied
instance URL from the server) and would store third-party Git tokens for all
tenants under one key, while only serving the small subset of users whose
instance is publicly reachable. Product decision: offer it on self-host only.

- Add an explicit deployment switch MULTICA_VCS_INTEGRATION_ENABLED (default
  off). Connect, rotate, and webhook now require BOTH the switch on AND a valid
  MULTICA_VCS_SECRET_KEY — the switch is the product boundary, not key presence
  alone. When off, connect/rotate return 404 and the webhook returns a bare 404
  (no config leak), independent of the frontend.
- /api/config exposes vcs_integration_available (mirrors the switch, omitted
  when false) so the Settings UI hides the whole "Git providers" section on
  cloud instead of surfacing an operator-only "missing key" hint.
- docker-compose.selfhost.yml defaults the switch on; .env.example documents it.
- Docs (en/zh) lead with a callout: available on self-hosted Multica only, not
  Multica Cloud, and clarify "self-hosted" means Multica itself, not just Git.

#5006 / #5883 stay in place — the schema and backend capability are retained;
this only gates availability. No cloud VCS connection can exist (connect always
required the key, which the cloud never set), so nothing needs migrating.

Verified: go build/vet + VCS/config handler tests on a fresh migrated DB
(incl. a new disabled-deployment 404 test); pnpm typecheck (core + views) and
the integrations-tab + core schema/config vitest suites pass.

Co-authored-by: multica-agent <github@multica.ai>

* fix(vcs): complete self-host integration gating (MUL-5138)

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-24 16:39:22 +08:00
ba0b711787 chore(ci): run test-go.sh self-test in Go test entry points (MUL-5169) (#5802)
Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-23 02:27:11 +08:00
YYClaw 36533bbc2b fix(test): prevent agent CLI execution in default tests (#5789) 2026-07-23 01:59:06 +08:00
220fa58264 fix: guide SSH installs to token login (#5318)
Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-13 16:42:46 +08:00
AdamQQQandClaude Opus 4.6 3030c803bf fix(scripts): fix version comparison to prevent unnecessary CLI upgrades (#4227)
`multica version` now outputs two lines (version + go info). The old
`awk '{print $2}'` captured both lines, causing the version comparison
to always mismatch and triggering unnecessary upgrades.

Fixes https://github.com/multica-ai/multica/issues/4226

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-17 13:27:24 +08:00
4df6c1468d fix: validate selfhost compose env defaults (#4138)
Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-06-15 15:43:10 +08:00