5 Commits
Author SHA1 Message Date
姚劲andClaude Opus 5 f643dd7da0 feat(safe-mode): add floating repair agent widget in recovery and safe mode (#431)
* feat(safe-mode): add floating repair agent widget in recovery and safe mode

- Add floating repair agent FAB widget on plugin recovery and safe mode overlay
- Provide pre-configured quick action prompts with error diagnostics and logs
- Support switching models and uploading/pasting screenshots for vision models
- Connect with isolated safe mode Harness session and stream responses via IPC

* feat(safe-mode): temporarily disable floating repair agent chat widget

* feat(safe-mode): enable repair agent with 0.9.0 incident knowledge and offline diagnostics

* feat(release): expand release notes generator with performance section and full changelog

* feat(window): persist window bounds, maximized state, and zoom level

* fix(repair-agent): handle unconfigured models with api key drawer and prevent silent stream failures

* fix(dev): strip inherited ELECTRON_RUN_AS_NODE to prevent electron launch crash

* feat(dev): forward runtime logs to console in development and log single-instance conflicts

* fix(repair-agent): inherit working default model route and add dual-track history polling

* refactor(repair-agent): rely purely on websocket stream without redundant history polling

* feat(repair-agent): replace floating widget with seamless native Harness UI diagnostic cards

* fix(windows): stop GPU fallback ladder from degrading on TDR device-loss recoveries, recover the menu view from a lost renderer

gcp3 crash data for v0.9.0 showed 23/30 Windows gpu-crash reports at exitCode=34
— Chromium's own exit code when it detects a lost D3D11 device (typically a
driver TDR reset) and exits the GPU process on purpose so it can restart it.
The fallback ladder was counting this self-recovery as evidence the sandbox is
broken and would degrade hardware acceleration after 3 of these on an
otherwise healthy machine. isGpuLossFatal now takes the exit code and treats
exitCode 34 as non-fatal, only logging a breadcrumb.

Separately, ~6 renderer-crash reports were on windows-menu.html: that view runs
in its own WebContentsView outside installMainWindowRendererRecovery, which
only hooks the main window's webContents, so a lost renderer there left a dead
menu until the whole app restarted. It now reloads itself on render-process-gone.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(session): skip unreadable session logs instead of failing the whole listing

One corrupt JSONL/zstd session log made sessionPersistence.list() throw, which
failed dsh-workspace and stopped the Harness entirely (0.9.0 startup-failure
reports: "corrupt Zstandard session log", corrupt header id). The log is now
skipped with a warning naming its path and left on disk untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(startup): stop reporting attributed plugin failures, clear stale writer locks, fix repair diagnosis

- Harness startup failures that plugin recovery attributes to a user plugin
  are handed to the user and their pending crash report is discarded; only
  the frontend recovery path did this before, so they were still uploaded.
- Remove dsh-atomic-write locks in DSH_HOME and DSH_HOME/profiles whose owner
  pid is gone before launching; a killed Harness left .credentials.yaml.lock
  behind and the next launch timed out on it.
- Repair Agent diagnosis uses loader provenance and the recovery-resolved
  plugins, and only reports a startup timeout when the runtime says so, since
  "waiting for Harness" + SIGTERM matched ordinary failure cleanup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(safe-mode): ensure safe mode uses hoisted node-linker and hide diagnostics when plugins are identified

* feat(runtime): add preset YAML validation step before launch

* refactor(runtime): remove redundant preset check in favor of native Cordis loader logging

* feat(repair-agent): unify diagnostic cards into a single entry with verification and reporting instructions

* fix(repair-agent): restore the 3 fine-grained diagnostic cards in safe mode

* fix(safe-mode): unify safe mode diagnostic cards into a single AI agent entry

* fix(web-import): never replace a desktop home that already holds plugins or credentials

A 0.9.0-rc1 home with plugins or .credentials.yaml but no settings.yaml or
sessions was treated as unused, so the web import deleted it wholesale. It now
counts as used. An unused home is also moved aside before deletion so
delete-pending entries on Windows cannot block links created at the same path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(recovery): don't remove plugins a deferred migration has not installed yet

When the generation migration failed (e.g. pnpm EPERM right after install),
imported plugins stayed manifest-only, failed to prepare, and plugin recovery
removed them as incompatible. Deferred migrations now report the pending
plugins and recovery excludes them. Transient EPERM/EBUSY/EACCES failures no
longer freeze the same profile for six hours, so the next launch retries.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(windows): stop the titlebar drag region from swallowing clicks and churning

The drag region now ignores pointer events, and only semantic dialog markers
hide it: class-name guesses like [class*="modal"] matched permanent elements
and hid it for good. Visibility checks run at most once a frame, since
streaming output mutates the DOM continuously.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(diagnostics): send crashes without an event id right away

Only a plugin-attributed startup failure should wait for recovery to decide
whether to discard it. A crash with no event id compared equal to an unset
pending id and was held back too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(market): pin the exact version a declared dshmarket range names

The installer pins and verifies an exact version, so a declared range such as
^1.47.0 is reduced to the version it names instead of being passed through.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: ignore the local pnpm content store

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(safe-mode): resolve Safe Mode plugins from the installation, not the module fallback

Safe Mode loads only installation-owned bundles, yet it still went through
the shared $DSH_HOME/profiles/node_modules fallback, whose junctions Windows
can refuse to recreate (EPERM) for minutes — the top startup failure of
v0.9.0. The desktop now sets DSH_DESKTOP_HOST_RESOLVED for the Safe Mode
profile only; the patched Harness then skips healing the fallback and
resolves bare plugin names, including entries created at runtime through
ctx.loader.create, from its own installation. A flag leaked from the parent
environment never reaches a normal profile.

This replaces the hoisted node-linker workaround for Safe Mode.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(runtime): mirror Harness logs to the console only in development

A packaged app's stdout may be a closed pipe; writing every Harness log line
to it there buys nothing and can fail. harness.log still receives every line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(recovery): still blame a migration-pending plugin whose legacy copy is installed

f475ce59 exempted every plugin with a pending deferred migration from
recovery. Only a plugin the migration left manifest-only is not broken; one
whose legacy copy is installed still loads and can be the real culprit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(logging): record Harness runtime warnings and errors in harness.log

Cordis's ctx.logger keeps messages only in an in-memory ring, and its default
level filters out warnings; nothing in the shipped composition exported them.
Session activation failures (a preset that cannot be found or mounted) never
reach the logger at all — they are only pushed to the client. So errors after
startup left no trace for a person or the Repair Agent once the process was
gone.

A new desktop plugin, dsh-desktop-log-bridge, is composed first into both the
normal and the Safe Mode profile. It writes warn and error messages, the
errors already in the ring, and api-session/error events to stderr, which
the desktop records into harness.log. It rate-limits floods and writes one
ready line per launch, so a log without it is known to predate the bridge.

Every bridged line carries a [harness-log] prefix, and latestHarnessAttemptLogs
drops such lines: runtime warnings never become recovery's failure cause or
the plugin it blames.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(safe-mode): disable plugins instead of removing them

Safe Mode now switches a plugin off the way the plugin market's own toggle
does, so it can come back without a reinstall: `disabled: true` rows for the
loader entries the package inserts in the user patch layer, which the loader
re-applies on every boot, plus the market's .dsh-market/state.json list,
which is also the only switch for client-only packages. Nothing is deleted.

Disabled plugins show as disabled with a Re-enable action instead of a
selection. Compatibility issues that call for disabling a plugin use the same
path; the old approach of dropping it from dsh.profile.bundles was silently
undone, because the generation projection rewrites that list on every launch.
Compatibility inspection skips disabled plugins, since they never load.

A disable-carrier (a bundle whose patch disables another plugin) cannot be
switched off on its own without stranding the plugin it replaces, so it
keeps the backed-up removal. Startup recovery keeps removal too: a package
that is itself broken fails before the patch layer's disable applies. When
recovery removes a plugin it also clears its market disable entry, as the
market's own uninstall does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(repair-agent): start a fresh, well-briefed repair session from one card

Recovery and Safe Mode now offer a single "start the repair agent" card. Each
click opens a new session in the Safe Mode Harness and makes it the session
the Harness UI shows: the UI restores its current session from desktop
storage when the page loads, so the Recovery path opens the session before
its page loads and the Safe Mode path reloads the already open page. The
floating repair widget is gone.

The session's workspace is the Harness home (profiles, plugins, patch layer)
instead of the empty launch root, and the first turn of each session carries
the diagnosis. The system prompt is rebuilt around real incidents:

- a directory map that separates the normal profile it repairs from the Safe
  Mode profile it runs in;
- the path to harness.log with how to find the failed launch in it, plus a
  short excerpt, instead of pasting logs; Safe Mode's own logs are no longer
  used when no failed launch was captured;
- one bilingual playbook source, so the Chinese and English prompts can no
  longer drift: plugin load failures, broken packages (which disabling cannot
  fix), the Windows module-fallback EPERM (no Developer Mode advice), a
  corrupt patch layer, startup timeouts, the code -> ptc preset rename, and
  schema errors in copied user presets;
- actions named after the real UI, and a confirm-and-back-up rule before any
  file change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(market): stop shipping dshmarket so market upgrades take effect

Harness resolves every profile bundle from its own installation before the
profile (`resolveBundleDir` in dsh-app-boot). Since #352 pinned dshmarket
1.45.1 as a runtime dependency, the packaged copy won that lookup: upgrading
the market wrote 1.48.0 into the profile, the market and the baseline check
both read 1.48.0 back, and Harness kept loading 1.45.1 after every restart.

Nothing in the app imports dshmarket; the profile copy is installed from npm
by the market installer and `ensureMarketBaseline`. It stays a
devDependency for the tests that exercise its real profile reader, which
keeps it out of the packaged app.

`ensureMarketBaseline` now also notes when the copy Harness would load is not
the profile's, so a shadowed market shows up in harness.log instead of as a
version the UI reports but never runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(recovery): let startup recovery repair a plugin market that breaks startup

The market is a core bundle, so startup-failure attribution never named it
and the Recovery page, finding no culprit, could only offer Safe Mode — which
does not load the market and has nothing to act on. A market release that
stopped Harness from booting therefore failed every launch the same way.
Until now the packaged dshmarket shadowed the profile copy and hid this;
with it gone the profile copy really loads, so the loop is live.

When the failed launch's log (or loader provenance) names dshmarket, the
Recovery page shows it as its own row, apart from the third-party plugin
list whose removal path refuses core bundles:

- Upgrade to the newest release the market check finds compatible with this
  Harness, for a market too old for it; a release already tried here is not
  offered again.
- Install the verified version (VERIFIED_MARKET_BASELINE). That moves down
  from a broken release, up from a stale one, or reinstalls a damaged copy.
  It is the primary action when the market is the only culprit.
- Remove the market; installed community plugins are kept.

Versions come from the main process's own check, never from the page. Each
change stops Harness first, as the shared-tree installer requires, pins the
exact version so the next baseline check does not reinstall the broken
release, then relaunches and returns to recovery if startup still fails.
The market removal behind the settings page is shared as `removeMarket`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(recovery): render the plugin market row like any other plugin row

The market row added its own touches: a version after the name, a green
button for restoring the verified version (which can be a downgrade), and a
confirmation dialog plus its own wording for removal. Third-party rows have
none of these, so the market read as a different kind of item.

Show the name only, keep green for the upgrade alone with every other action
in the neutral style, and remove it the way the page removes any plugin —
"Remove this plugin", no extra confirmation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(recovery): offer the plugin market only upgrade and removal

Restoring the verified market version was an action no other plugin row
has. Drop it, so the market row is a plugin row like the rest: upgrade to a
compatible release when the market check finds one, and removal.

When the market alone blocks startup, the page now uses the same buttons as
a lone third-party plugin: "Upgrade plugin and restart" with "Uninstall
plugin" beside it when an upgrade exists, otherwise "Remove this plugin and
continue".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: stop tracking the local pnpm content store

0d083d5d committed the whole `.pnpm-store/` (13,357 files, ~235 MB) along
with its four intended changes, before 7a2cb730 added it to .gitignore. An
ignore rule does not untrack files already committed, so they stayed in the
branch and buried the PR's real diff.

Remove them from the index only; the local store is untouched and stays
ignored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: drop tests that only grep our own source code

About 40 tests read src/main/index.ts, the preload, build/*.html, the NSIS
script or our own packages' client.js as text and asserted that particular
lines were present — `toContain("ipcMain.handle('safe-mode:action'")`,
`indexOf(a) < indexOf(b)` and the like. They execute nothing, so they cannot
catch a behavior regression, yet they fail on any rename or reformat, and on
Windows, where the checkout has CRLF line endings, any expected string that
spans a newline never matches. Two of them failed CI that way on PR #431.

Removed: every test block whose assertions were such text matches, and the
source-text lines inside otherwise behavioral tests (windows-titlebar
allowlist, runtime Node mode, release dev channel, branding postinstall,
harness-node-entry). The dev-channel check now reads electron-builder.dev.cjs
as a config object instead of matching its text.

Kept: tests that execute code, parse package.json/lockfile/YAML/JSON
config, or check the installed third-party bundles in node_modules (patch
and upstream-contract checks), since those inspect what actually ships.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(repair-agent): drop tests that only check the prompt's wording

Two repair-prompt tests asserted fixed sentences of the generated prompt:
exact Chinese phrases with hard-coded POSIX paths, and a fixed count of
seven playbooks plus two error strings. Any edit to the copy broke them
without saying anything about how the agent behaves, and the path one
failed on Windows, where path.join correctly yields backslashes.

Keep the test of real prompt logic: only the tail of the log is included,
and none at all when no failed launch was captured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(repair-agent): drop the remaining repair-prompt test

It asserted which log lines and which sentence end up in the generated
prompt text. Like the other prompt tests, that pins the copy rather than
anything the agent does with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(safe-mode): offer the repair agent as one line under the summary

The repair agent sat in its own section below the actions — a heading, a
divider and a card, about 140px. In a 1280x800 window that pushed the plugin
list into a scrollbar, which the native Windows recovery UI check rejects:
every plugin has to be visible at once. Its text was also Chinese-only.

Show it as a single link under the summary instead, in both languages. It
sends the same diagnostic request as the card did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(recovery): offer the repair agent as a button beside Safe Mode

The recovery page showed the repair agent in its own section under the
technical details — a heading and a card, in Chinese only. Make it a
secondary "Repair agent" button next to "Enter Safe Mode", in both
languages. As before, it appears only when no plugin or market was
identified; a named culprit has its own repair above.

The native Windows UI check expected Safe Mode to be the only action when
nothing is identified; it now expects the agent beside it, and nowhere else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(runtime): keep a clean launch off stderr

The packaged Windows smoke fails any launch that writes to stderr once a
workspace and session exist. Two lines did on every launch, and neither
reported a problem:

- The log bridge's "recording from here on" marker. It is a notice, so it
  now goes to stdout; stderr carries only bridged warnings and errors.
- `patch: entry dsh-market not found`. The desktop patch configured the
  market's entry (restart off, wait for desktopProfiles) unconditionally,
  but the market is optional, and a patch row whose entry is absent makes the
  loader warn. It was always there; the bridge only made it visible. The row
  now lives in dsh-desktop-market.patch.yml, passed as a second --patch only
  when the profile boots dshmarket, so a market profile behaves as before.

Checked by booting a fresh web profile (no market overlay, no bridged
warning) and one that boots dshmarket (overlay applied, no warning).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 21:28:25 +08:00
yaojin3616 6fae210b5a feat(ci): 支持预发布飞书群通知(标明预发布,范围为上一个tag到当前tag) 2026-09-02 00:51:02 -07:00
yaojin3616 e45e921161 feat(release): AI-organized Chinese GitHub release notes generator 2026-09-01 06:57:27 -07:00
yaojin3616 b76be70c66 fix(ci): package macOS development channel on non-tag builds and resolve test suite errors 2026-08-21 13:16:03 +08:00
yaojin3616 6a39b9828b feat(ci): add Feishu release notification and bilingual highlight summary workflow 2026-08-20 21:03:09 +08:00