mirror of
https://github.com/dataelement/dsh-desktop.git
synced 2026-09-28 05:23:01 +08:00
* feat(safe-mode): add floating repair agent widget in recovery and safe mode - Add floating repair agent FAB widget on plugin recovery and safe mode overlay - Provide pre-configured quick action prompts with error diagnostics and logs - Support switching models and uploading/pasting screenshots for vision models - Connect with isolated safe mode Harness session and stream responses via IPC * feat(safe-mode): temporarily disable floating repair agent chat widget * feat(safe-mode): enable repair agent with 0.9.0 incident knowledge and offline diagnostics * feat(release): expand release notes generator with performance section and full changelog * feat(window): persist window bounds, maximized state, and zoom level * fix(repair-agent): handle unconfigured models with api key drawer and prevent silent stream failures * fix(dev): strip inherited ELECTRON_RUN_AS_NODE to prevent electron launch crash * feat(dev): forward runtime logs to console in development and log single-instance conflicts * fix(repair-agent): inherit working default model route and add dual-track history polling * refactor(repair-agent): rely purely on websocket stream without redundant history polling * feat(repair-agent): replace floating widget with seamless native Harness UI diagnostic cards * fix(windows): stop GPU fallback ladder from degrading on TDR device-loss recoveries, recover the menu view from a lost renderer gcp3 crash data for v0.9.0 showed 23/30 Windows gpu-crash reports at exitCode=34 — Chromium's own exit code when it detects a lost D3D11 device (typically a driver TDR reset) and exits the GPU process on purpose so it can restart it. The fallback ladder was counting this self-recovery as evidence the sandbox is broken and would degrade hardware acceleration after 3 of these on an otherwise healthy machine. isGpuLossFatal now takes the exit code and treats exitCode 34 as non-fatal, only logging a breadcrumb. Separately, ~6 renderer-crash reports were on windows-menu.html: that view runs in its own WebContentsView outside installMainWindowRendererRecovery, which only hooks the main window's webContents, so a lost renderer there left a dead menu until the whole app restarted. It now reloads itself on render-process-gone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(session): skip unreadable session logs instead of failing the whole listing One corrupt JSONL/zstd session log made sessionPersistence.list() throw, which failed dsh-workspace and stopped the Harness entirely (0.9.0 startup-failure reports: "corrupt Zstandard session log", corrupt header id). The log is now skipped with a warning naming its path and left on disk untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(startup): stop reporting attributed plugin failures, clear stale writer locks, fix repair diagnosis - Harness startup failures that plugin recovery attributes to a user plugin are handed to the user and their pending crash report is discarded; only the frontend recovery path did this before, so they were still uploaded. - Remove dsh-atomic-write locks in DSH_HOME and DSH_HOME/profiles whose owner pid is gone before launching; a killed Harness left .credentials.yaml.lock behind and the next launch timed out on it. - Repair Agent diagnosis uses loader provenance and the recovery-resolved plugins, and only reports a startup timeout when the runtime says so, since "waiting for Harness" + SIGTERM matched ordinary failure cleanup. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(safe-mode): ensure safe mode uses hoisted node-linker and hide diagnostics when plugins are identified * feat(runtime): add preset YAML validation step before launch * refactor(runtime): remove redundant preset check in favor of native Cordis loader logging * feat(repair-agent): unify diagnostic cards into a single entry with verification and reporting instructions * fix(repair-agent): restore the 3 fine-grained diagnostic cards in safe mode * fix(safe-mode): unify safe mode diagnostic cards into a single AI agent entry * fix(web-import): never replace a desktop home that already holds plugins or credentials A 0.9.0-rc1 home with plugins or .credentials.yaml but no settings.yaml or sessions was treated as unused, so the web import deleted it wholesale. It now counts as used. An unused home is also moved aside before deletion so delete-pending entries on Windows cannot block links created at the same path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(recovery): don't remove plugins a deferred migration has not installed yet When the generation migration failed (e.g. pnpm EPERM right after install), imported plugins stayed manifest-only, failed to prepare, and plugin recovery removed them as incompatible. Deferred migrations now report the pending plugins and recovery excludes them. Transient EPERM/EBUSY/EACCES failures no longer freeze the same profile for six hours, so the next launch retries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(windows): stop the titlebar drag region from swallowing clicks and churning The drag region now ignores pointer events, and only semantic dialog markers hide it: class-name guesses like [class*="modal"] matched permanent elements and hid it for good. Visibility checks run at most once a frame, since streaming output mutates the DOM continuously. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(diagnostics): send crashes without an event id right away Only a plugin-attributed startup failure should wait for recovery to decide whether to discard it. A crash with no event id compared equal to an unset pending id and was held back too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(market): pin the exact version a declared dshmarket range names The installer pins and verifies an exact version, so a declared range such as ^1.47.0 is reduced to the version it names instead of being passed through. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: ignore the local pnpm content store Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(safe-mode): resolve Safe Mode plugins from the installation, not the module fallback Safe Mode loads only installation-owned bundles, yet it still went through the shared $DSH_HOME/profiles/node_modules fallback, whose junctions Windows can refuse to recreate (EPERM) for minutes — the top startup failure of v0.9.0. The desktop now sets DSH_DESKTOP_HOST_RESOLVED for the Safe Mode profile only; the patched Harness then skips healing the fallback and resolves bare plugin names, including entries created at runtime through ctx.loader.create, from its own installation. A flag leaked from the parent environment never reaches a normal profile. This replaces the hoisted node-linker workaround for Safe Mode. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(runtime): mirror Harness logs to the console only in development A packaged app's stdout may be a closed pipe; writing every Harness log line to it there buys nothing and can fail. harness.log still receives every line. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(recovery): still blame a migration-pending plugin whose legacy copy is installedf475ce59exempted every plugin with a pending deferred migration from recovery. Only a plugin the migration left manifest-only is not broken; one whose legacy copy is installed still loads and can be the real culprit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(logging): record Harness runtime warnings and errors in harness.log Cordis's ctx.logger keeps messages only in an in-memory ring, and its default level filters out warnings; nothing in the shipped composition exported them. Session activation failures (a preset that cannot be found or mounted) never reach the logger at all — they are only pushed to the client. So errors after startup left no trace for a person or the Repair Agent once the process was gone. A new desktop plugin, dsh-desktop-log-bridge, is composed first into both the normal and the Safe Mode profile. It writes warn and error messages, the errors already in the ring, and api-session/error events to stderr, which the desktop records into harness.log. It rate-limits floods and writes one ready line per launch, so a log without it is known to predate the bridge. Every bridged line carries a [harness-log] prefix, and latestHarnessAttemptLogs drops such lines: runtime warnings never become recovery's failure cause or the plugin it blames. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(safe-mode): disable plugins instead of removing them Safe Mode now switches a plugin off the way the plugin market's own toggle does, so it can come back without a reinstall: `disabled: true` rows for the loader entries the package inserts in the user patch layer, which the loader re-applies on every boot, plus the market's .dsh-market/state.json list, which is also the only switch for client-only packages. Nothing is deleted. Disabled plugins show as disabled with a Re-enable action instead of a selection. Compatibility issues that call for disabling a plugin use the same path; the old approach of dropping it from dsh.profile.bundles was silently undone, because the generation projection rewrites that list on every launch. Compatibility inspection skips disabled plugins, since they never load. A disable-carrier (a bundle whose patch disables another plugin) cannot be switched off on its own without stranding the plugin it replaces, so it keeps the backed-up removal. Startup recovery keeps removal too: a package that is itself broken fails before the patch layer's disable applies. When recovery removes a plugin it also clears its market disable entry, as the market's own uninstall does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(repair-agent): start a fresh, well-briefed repair session from one card Recovery and Safe Mode now offer a single "start the repair agent" card. Each click opens a new session in the Safe Mode Harness and makes it the session the Harness UI shows: the UI restores its current session from desktop storage when the page loads, so the Recovery path opens the session before its page loads and the Safe Mode path reloads the already open page. The floating repair widget is gone. The session's workspace is the Harness home (profiles, plugins, patch layer) instead of the empty launch root, and the first turn of each session carries the diagnosis. The system prompt is rebuilt around real incidents: - a directory map that separates the normal profile it repairs from the Safe Mode profile it runs in; - the path to harness.log with how to find the failed launch in it, plus a short excerpt, instead of pasting logs; Safe Mode's own logs are no longer used when no failed launch was captured; - one bilingual playbook source, so the Chinese and English prompts can no longer drift: plugin load failures, broken packages (which disabling cannot fix), the Windows module-fallback EPERM (no Developer Mode advice), a corrupt patch layer, startup timeouts, the code -> ptc preset rename, and schema errors in copied user presets; - actions named after the real UI, and a confirm-and-back-up rule before any file change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(market): stop shipping dshmarket so market upgrades take effect Harness resolves every profile bundle from its own installation before the profile (`resolveBundleDir` in dsh-app-boot). Since #352 pinned dshmarket 1.45.1 as a runtime dependency, the packaged copy won that lookup: upgrading the market wrote 1.48.0 into the profile, the market and the baseline check both read 1.48.0 back, and Harness kept loading 1.45.1 after every restart. Nothing in the app imports dshmarket; the profile copy is installed from npm by the market installer and `ensureMarketBaseline`. It stays a devDependency for the tests that exercise its real profile reader, which keeps it out of the packaged app. `ensureMarketBaseline` now also notes when the copy Harness would load is not the profile's, so a shadowed market shows up in harness.log instead of as a version the UI reports but never runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(recovery): let startup recovery repair a plugin market that breaks startup The market is a core bundle, so startup-failure attribution never named it and the Recovery page, finding no culprit, could only offer Safe Mode — which does not load the market and has nothing to act on. A market release that stopped Harness from booting therefore failed every launch the same way. Until now the packaged dshmarket shadowed the profile copy and hid this; with it gone the profile copy really loads, so the loop is live. When the failed launch's log (or loader provenance) names dshmarket, the Recovery page shows it as its own row, apart from the third-party plugin list whose removal path refuses core bundles: - Upgrade to the newest release the market check finds compatible with this Harness, for a market too old for it; a release already tried here is not offered again. - Install the verified version (VERIFIED_MARKET_BASELINE). That moves down from a broken release, up from a stale one, or reinstalls a damaged copy. It is the primary action when the market is the only culprit. - Remove the market; installed community plugins are kept. Versions come from the main process's own check, never from the page. Each change stops Harness first, as the shared-tree installer requires, pins the exact version so the next baseline check does not reinstall the broken release, then relaunches and returns to recovery if startup still fails. The market removal behind the settings page is shared as `removeMarket`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(recovery): render the plugin market row like any other plugin row The market row added its own touches: a version after the name, a green button for restoring the verified version (which can be a downgrade), and a confirmation dialog plus its own wording for removal. Third-party rows have none of these, so the market read as a different kind of item. Show the name only, keep green for the upgrade alone with every other action in the neutral style, and remove it the way the page removes any plugin — "Remove this plugin", no extra confirmation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(recovery): offer the plugin market only upgrade and removal Restoring the verified market version was an action no other plugin row has. Drop it, so the market row is a plugin row like the rest: upgrade to a compatible release when the market check finds one, and removal. When the market alone blocks startup, the page now uses the same buttons as a lone third-party plugin: "Upgrade plugin and restart" with "Uninstall plugin" beside it when an upgrade exists, otherwise "Remove this plugin and continue". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: stop tracking the local pnpm content store0d083d5dcommitted the whole `.pnpm-store/` (13,357 files, ~235 MB) along with its four intended changes, before7a2cb730added it to .gitignore. An ignore rule does not untrack files already committed, so they stayed in the branch and buried the PR's real diff. Remove them from the index only; the local store is untouched and stays ignored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: drop tests that only grep our own source code About 40 tests read src/main/index.ts, the preload, build/*.html, the NSIS script or our own packages' client.js as text and asserted that particular lines were present — `toContain("ipcMain.handle('safe-mode:action'")`, `indexOf(a) < indexOf(b)` and the like. They execute nothing, so they cannot catch a behavior regression, yet they fail on any rename or reformat, and on Windows, where the checkout has CRLF line endings, any expected string that spans a newline never matches. Two of them failed CI that way on PR #431. Removed: every test block whose assertions were such text matches, and the source-text lines inside otherwise behavioral tests (windows-titlebar allowlist, runtime Node mode, release dev channel, branding postinstall, harness-node-entry). The dev-channel check now reads electron-builder.dev.cjs as a config object instead of matching its text. Kept: tests that execute code, parse package.json/lockfile/YAML/JSON config, or check the installed third-party bundles in node_modules (patch and upstream-contract checks), since those inspect what actually ships. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(repair-agent): drop tests that only check the prompt's wording Two repair-prompt tests asserted fixed sentences of the generated prompt: exact Chinese phrases with hard-coded POSIX paths, and a fixed count of seven playbooks plus two error strings. Any edit to the copy broke them without saying anything about how the agent behaves, and the path one failed on Windows, where path.join correctly yields backslashes. Keep the test of real prompt logic: only the tail of the log is included, and none at all when no failed launch was captured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(repair-agent): drop the remaining repair-prompt test It asserted which log lines and which sentence end up in the generated prompt text. Like the other prompt tests, that pins the copy rather than anything the agent does with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(safe-mode): offer the repair agent as one line under the summary The repair agent sat in its own section below the actions — a heading, a divider and a card, about 140px. In a 1280x800 window that pushed the plugin list into a scrollbar, which the native Windows recovery UI check rejects: every plugin has to be visible at once. Its text was also Chinese-only. Show it as a single link under the summary instead, in both languages. It sends the same diagnostic request as the card did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(recovery): offer the repair agent as a button beside Safe Mode The recovery page showed the repair agent in its own section under the technical details — a heading and a card, in Chinese only. Make it a secondary "Repair agent" button next to "Enter Safe Mode", in both languages. As before, it appears only when no plugin or market was identified; a named culprit has its own repair above. The native Windows UI check expected Safe Mode to be the only action when nothing is identified; it now expects the agent beside it, and nowhere else. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(runtime): keep a clean launch off stderr The packaged Windows smoke fails any launch that writes to stderr once a workspace and session exist. Two lines did on every launch, and neither reported a problem: - The log bridge's "recording from here on" marker. It is a notice, so it now goes to stdout; stderr carries only bridged warnings and errors. - `patch: entry dsh-market not found`. The desktop patch configured the market's entry (restart off, wait for desktopProfiles) unconditionally, but the market is optional, and a patch row whose entry is absent makes the loader warn. It was always there; the bridge only made it visible. The row now lives in dsh-desktop-market.patch.yml, passed as a second --patch only when the profile boots dshmarket, so a market profile behaves as before. Checked by booting a fresh web profile (no market overlay, no bridged warning) and one that boots dshmarket (overlay applied, no warning). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>