[rush] Dogfood the Rush daemon in rushstack: allow command-scoped plugins, add snapshot workflow and guide - #6046
Sean Larkin (TheLarkInn) wants to merge 88 commits into
Conversation
…add daemon dogfooding workflow Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…g guide Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…n starts Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…me fingerprint The client only needs to locate rush.json, read daemon settings, parse rushx arguments and build its daemon start command. Move those into small rush-lib and rush-daemon modules so the client no longer loads the whole engine for every invocation, and load version selection and rushx discovery only when needed. The runtime fingerprint now also covers the bundle chunks under dist, which contain the implementation, and recomputes a resolved path only when a file stat changes. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The detached helper waited for readiness only until the requesting client deadline (15 seconds). A slower first start, such as the first start from a freshly deployed snapshot on Windows, left its startup reservation behind. Every later client then refused to use the ready daemon and fell back to native Rush after 15 seconds. The helper now waits for a live launcher for at least 120 seconds; the reservation is still retained if the launcher exits or never becomes ready. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…oader The heft-config-file rig fallback resolves the rig profile synchronously, which serializes the daemon's concurrent uncached project-configuration loads (two per request). Resolve it asynchronously first; RigConfig caches the same result, and a resolution failure still surfaces through a fresh instance exactly as before. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…sh-daemon # Conflicts: # libraries/rush-lib/src/api/test/WorkspaceInputFingerprint.test.ts
…sh-daemon main (#6090) now keeps rush-client from loading @microsoft/rush-lib with on-demand requires and lazy start-command resolution. Take that implementation for the client and drop this branch's overlapping client-only modules (RushJsonLocation, RushXCommandLineArguments, DaemonLaunchCommand and bundledRushVersion) and their change files. The daemon-side changes in this branch are unaffected. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
requested the copilot review on this pr |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The guide’s cleanup command can erase contributors’ pre-existing uncommitted edits.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Enables safe, source-built Rush daemon dogfooding with command-scoped plugin support, improved Windows behavior, startup recovery, and contributor documentation.
Changes:
- Permits plugins that are inert for daemon builds and expands fingerprint coverage.
- Hides daemon child-process consoles on Windows and extends startup handoff time.
- Adds tests, deployment tooling, release notes, and a dogfooding guide.
| File | Description |
|---|---|
README.md |
Links the dogfooding guide. |
libraries/rush-lib/src/pluginFramework/PluginManager.ts |
Detects command-participating plugins. |
libraries/rush-lib/src/pluginFramework/PluginLoader/AutoinstallerPluginLoader.ts |
Exposes plugin shape paths. |
libraries/rush-lib/src/cli/scriptActions/PhasedScriptAction.ts |
Exposes schedulable phases. |
libraries/rush-lib/src/api/WorkspaceInputFingerprint.ts |
Expands and optimizes fingerprints. |
libraries/rush-lib/src/api/test/WorkspaceInputFingerprint.test.ts |
Tests plugin-shape reloads. |
libraries/rush-lib/src/api/test/RushProjectConfiguration.test.ts |
Tests asynchronous rig resolution. |
libraries/rush-lib/src/api/test/PhasedCommandEngine.test.ts |
Tests plugin eligibility rules. |
libraries/rush-lib/src/api/RushProjectConfiguration.ts |
Pre-resolves rig profiles asynchronously. |
libraries/rush-lib/src/api/PhasedCommandEngine.ts |
Applies the narrower plugin gate. |
libraries/rush-daemon/src/WindowsSubprocessConsoles.ts |
Defaults Windows children to hidden. |
libraries/rush-daemon/src/test/WindowsSubprocessConsoles.test.ts |
Tests child-process wrapping. |
libraries/rush-daemon/src/SelectedDaemonBootstrap.ts |
Installs the Windows default. |
libraries/rush-daemon/src/ProductionDaemonRequestResolver.ts |
Updates supported-surface documentation. |
libraries/rush-daemon/README.md |
Documents plugin and Windows behavior. |
libraries/rush-client-core/src/test/connectOrStartDaemon.test.ts |
Tests slow startup handoff. |
libraries/rush-client-core/src/connectOrStartDaemon.ts |
Extends helper readiness timeout. |
libraries/rush-client-core/README.md |
Documents startup handoff behavior. |
docs/rush/environment-variables.md |
Clarifies daemon opt-in scope. |
docs/rush/dogfooding-rush-daemon.md |
Adds the contributor workflow. |
common/config/rush/deploy-rush-daemon-dogfood.json |
Defines the snapshot deployment. |
common/changes/@rushstack/rush-daemon/thelarkinn-windows-hide-subprocess-consoles_2026-09-23-21-30-00.json |
Records the Windows fix. |
common/changes/@rushstack/rush-daemon/thelarkinn-dogfood-rush-daemon_2026-09-23-17-36-00.json |
Records documentation changes. |
common/changes/@rushstack/rush-client-core/thelarkinn-client-startup-perf_2026-09-23-22-20-00.json |
Records startup recovery. |
common/changes/@rushstack/rush-cli-client/thelarkinn-dogfood-rush-daemon_2026-09-23-17-36-00.json |
Records client documentation. |
common/changes/@microsoft/rush/thelarkinn-dogfood-rush-daemon_2026-09-23-17-36-00.json |
Records plugin support. |
common/changes/@microsoft/rush/thelarkinn-client-startup-perf_2026-09-23-22-20-00.json |
Records fingerprint improvements. |
apps/rush-cli-client/README.md |
Links and documents dogfooding. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…ed on #6046) (#6102) * [rush-lib] Workspace fingerprint: ignore build state of projects nested under common/config (and 3 more) Swarm integration step 1; original commit 433c87f24f (merge of swarm/r05 at 190d77ea17). Commits folded into this step (4): - 40b5b24d72 [rush-lib] Workspace fingerprint: ignore build state of projects nested under common/config - 1880e7b2b7 [rush-daemon] Warm set: keep resource-free retained results under memory pressure and idle expiry - 07e908d32d [rush-daemon] Admission: don't spend a default wait budget behind a graph load or across an admitted request's execution - 190d77ea17 [rush-lib] Move the nested common/config fingerprint test to the end of its suite Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Warm set: read status once per maintenance pass, not once per project (and 1 more) Swarm integration step 2; original commit 8553b2be55 (merge of swarm/r07 at d94054e04d). Commits folded into this step (2): - 57dd6a4b45 [rush-daemon] Warm set: read status once per maintenance pass, not once per project - d94054e04d [rush-lib] Share rig profile loads across projects in the uncached rush-project.json load Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Keep RUSH_DAEMON* variables out of lifecycle script environments (and 3 more) Swarm integration step 3; original commit 3b2d24277d (merge of swarm/r06 at 6c10e241a4). Commits folded into this step (4): - 803345396b [rush-lib] Keep RUSH_DAEMON* variables out of lifecycle script environments - b49d8b10cb [rush-daemon] Keep one daemon across per-session environment differences - 36892176cb [rush-daemon-transport] Reap detached operation groups of a daemon that died uncleanly - 6c10e241a4 [rush-daemon] Run each operation with its requester's per-session environment Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Let daemon engines serve builds with declared daemon-compatible plugins (and 1 more) Swarm integration step 4; original commit 1e786a4252 (merge of swarm/r01 at 162ff05eaf). Commits folded into this step (2): - 5fe505bf1a [rush-lib] Let daemon engines serve builds with declared daemon-compatible plugins - 162ff05eaf [rush-lib] Walk each plugin package once in the daemon runtime fingerprint Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Abort a shared iteration's leftover work once no remaining client needs it Swarm integration step 5; original commit 02f83a99ae (merge of swarm/r05 at bd7ac88caa). Refs: board 207. Commits folded into this step (1): - bd7ac88caa [rush-daemon] Abort a shared iteration's leftover work once no remaining client needs it (board 207) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Run the initial script for every command outside watch mode Swarm integration step 6; original commit 519ae33c8d. Refs: task 50, board 115. ShellOperationRunner runs a phase's `:incremental` script whenever the operation has a last state. Native `rush build` executes one iteration, so only watch mode ever had one, and watch mode turns cache writes off. rushd keeps one operation graph across requests, so from its second request on a daemon `build` ran `_phase:build:incremental` (in odsp-web, `heft run --only build --` without --clean) with cache writes on. Outputs of a deleted or renamed source file stayed in lib/ and were cached under the key native Rush computes for the clean build, so later native builds restored them (t04 board 115, board 450). ShellOperationRunnerPlugin now sets the incremental command only in watch mode. Every other command runs the initial script, exactly as a single `rush build` does, which is option (a) of task 50. Native `rush build` is unchanged (it never has a last state) and watch mode keeps `:incremental` with writes off. Tests: a rush-lib unit test runs a repeated operation with and without watch mode, and a daemon integration test runs three warm builds of a project with an `:incremental` script. Before the fix the second daemon build logged "Invoking (incremental): node build.cjs --incremental" followed by "Successfully set cache entry". (cherry picked from commit e382d2ac80f881451b7d527ea9a97e3037880923) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Share input captures between concurrent requests Swarm integration step 7; original commit f7ac88dc9d (merge of swarm/r07 at 55a16f14ef). Refs: F1 coalescer, board 555. GATE OK ch01 board 930 (tree 41b2227d59). CONFIRMED by ch03 board 681, t06 board 701/board 749, t01 board 861. Only 55a16f14ef; b6692d09a5 (task 74) is not gated. Commits folded into this step (1): - 55a16f14ef [rush-daemon] Share input captures between concurrent requests Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Skip cache writes when input files change while the inputs snapshot is taken Swarm integration step 8; original commit bc03023091 (merge of swarm/r06 at 29f86b181c). Refs: task 77. ch01 GATE OK board 1019 (tree 80c9659bac); claim board 698, CONFIRMED by o04 board 738, o03 board 742, o05 board 748. Commits folded into this step (2): - a644819daf [rush-lib] Don't write a build cache entry for inputs saved while the snapshot was taken - 29f86b181c [rush-lib] Make the parameter identity test independent of the host's RUSH_PARALLELISM Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] don't let a request received during a batch's reconcile join that batch Swarm integration step 9; original commit d0ed461613 (merge of swarm/r05 at 199c69347c). Refs: board 371, task 72. ch01 GATE OK board 1152 (gated tree 5c2b6b6cbf). Commits folded into this step (1): - 199c69347c [rush-daemon] Don't let a request received during a batch's reconcile join that batch (board 371) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Another request's graph load doesn't spend any finite wait timeout Swarm integration step 10; original commit 908dcc1f27 (merge of swarm/r05 at 868e4b9299). Refs: task 7 part 3, board 891. ch01 GATE OK board 1196 (gated tree 55698be825). Commits folded into this step (1): - 868e4b9299 [rush-daemon] Admission: another request's graph load doesn't spend any finite wait timeout Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Keep cache trust for unselected retained results and re-verify untrusted ones Swarm integration step 11; original commit 4456f17460 (merge of swarm/r07 at 53654af37c). Refs: board 401, task 74, board 982, board 1024. ch01 GATE OK board 1197 (gated tree 3014a03bb8). Commits folded into this step (3): - b6692d09a5 [rush-lib] Keep cache trust for unselected retained results and re-verify untrusted ones - 492126a051 [rush-lib] Re-verify retained results built against unverified dependency outputs without the build cache - 53654af37c [rush-daemon] Test that input captures are shared only between equal fingerprint environments Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-cli-client] Show engine and plugin warnings in agent output; keep the daemon alive when git hash-object fails Swarm integration step 12; original commit cc5c6aab7e (merge of swarm/r01 at 575cbf5698). Refs: board 601, task 80, board 899. ch01 GATE OK board 1272 and board 1273 (gated tree b691997c35). Commits folded into this step (3): - c11ee10320 [rush-cli-client] Show engine and plugin warnings in agent output - c77aa1dbef [package-deps-hash] Handle a git hash-object failure while ls-files and status still run - 575cbf5698 [package-deps-hash] Name the git command that failed in the error message Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [node-core-library] LockFile ties, TZ/locale rule, /proc start times and own-start memo Swarm integration step 13; original commit 0347cd0dbb. Refs: tasks 67, D and 62. ch01 GATE OK board 1389; o05 board 1360 (task 62 on ch01-t62: 32/32 served in 3 of 3 bursts); t01 board 1106/board 1186. Commits folded into this step (6): - 3dfb75ec86 [node-core-library] Don't grant a LockFile twice when lockfile birthtimes tie - 074f207217 [node-core-library] Keep the LockFile of a running process with another time zone or locale - 4e43a1ef8f [node-core-library] Don't set dirtyWhenAcquired when the holder releases the lock during a check - ad9c1211b4 [node-core-library] Read the start times of other LockFile processes from /proc on Linux - 698af4ef92 [node-core-library] Run ps for the start time of the current LockFile process only once on Linux - 9249e85d63 [rush-client-core] Test clients that start the daemon at the same time Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A request can share an input capture that started after it was received Swarm integration step 14; original commit f2ec9ffc26. ch01 GATE OK board 1430 (tree ae49a66e47 with o07's stack); t06 CONFIRMED board 1365; o05 board 1412 (cold split gone in auto-start mode). Commits folded into this step (1): - a4de3fe5fe [rush-daemon] Share running input captures that started after a request was received Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Hand requests to in-process Rush when a project's configuration cannot load Swarm integration step 15; original commit 69f30e84ab (merge of swarm/r01-t70 at c1b5c95adb). Scope: task 70. ch01 GATE OK board 1522 (tree 3a564b4533; WorkspaceRequestLifecycle.ts resolved with r07's board 1231 recipe: the Reuse capture stays tolerant and keeps task 94's receipt); o01 board 1291; r01 board 1250. Commits folded into this step (5): - 08aed7efd2 [package-deps-hash] Handle a git hash-object failure while ls-files and status still run - 3d786c1f78 [package-deps-hash] Name the git command that failed in the error message - a11aa6b774 [rush-lib] Name the project whose configuration an engine cannot load - 24cf595d6a [rush-lib] Report a missing project dependency file to an engine's host - c1b5c95adb [rush-daemon] Hand requests to in-process Rush when a project's configuration cannot load Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] With cache writes off, a result whose inputs changed during the run is not trusted as up to date Swarm integration step 16; original commit d6ac7c43f3 (merge of swarm/r06-t92 at 72c2f0e4f8). Scope: task 92. ch01 GATE OK board 1563 (tree 5c0b0ac25d); r06 board 1270; CONFIRMED by t03, t04, ch05 and o03. Commits folded into this step (1): - 72c2f0e4f8 [rush-lib] Re-run operations whose inputs changed during the iteration, also with cache writes off Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] A failed pnpm install releases the package-manager lock Swarm integration step 17; original commit 9a58fc2f85 (merge of swarm/r01-t103 at 93eacefb71). Scope: task 103. ch01 GATE OK board 1565 (tree e234c8d86d); o04 CONFIRMED board 1477; r01 board 1440. This is upstream row 16's commit. Commits folded into this step (1): - 93eacefb71 [rush-lib] Release the package manager lock when installing the package manager fails Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon-transport] rush-lib path handoff, startup reservation, lost-connection message and a private runtime folder Swarm integration step 18; original commit 261e4320dd (merge of swarm/r03-t87 at d9366621c8). Scope: tasks 4+5, 68, 86 and 87. ch01 GATE OK board 1675 (tree f85ce62f88). Commits: bb1d272738 (tasks 4+5), 959cbb84ee (task 68), a68e4d96a5, 696dba1d3e (task 86) and d9366621c8 (task 87). Commits folded into this step (5): - bb1d272738 [rush-lib] Hand off a by-name resolvable rush-lib and load built-in cache plugins from deploy outputs - 959cbb84ee [rush-client-core] Resolve a retained startup reservation next to a ready daemon - a68e4d96a5 [rush-client-core] Resolve a remaining startup reservation before a plain daemon stop - 696dba1d3e [rush-client-core] Explain a daemon that exited before delivering a result - d9366621c8 [rush-daemon] Meet one daemon per checkout whatever TMPDIR or XDG_RUNTIME_DIR says (task 87) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-client-core] Wait for a daemon a live process is still starting instead of running Rush in-process Swarm integration step 19; original commit a81a00520d (merge of o07/startup-fallback at 668753ad4d). Scope: task 95. Brings 647a8bf902 [rush-client-core] and 668753ad4d [rush-cli-client]. Gate: ch01 GATE OK board 1731 (tree e4a00622d4); burst: t01 board 1716. Commits folded into this step (2): - 647a8bf902 [rush-client-core] Wait for a daemon that a live process is still starting before permitting in-process Rush - 668753ad4d [rush-cli-client] Don't run Rush in-process while daemon startup is still pending Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Reuse workspace input digests and project configurations while their files are unchanged Swarm integration step 20; original commit 0f987a5b62 (merge of swarm/r07-t93-on-s10 at 18900fe380). Scope: task 93. Brings 143f21a891 and 18900fe380 [rush-lib]. Gate: ch01 GATE OK board 1869 (tree 8892a71708); m01 CONFIRMED board 1738 and board 1818; t04 board 1794. Commits folded into this step (2): - 143f21a891 [rush-lib] Reuse workspace input digests and project configurations while their files are unchanged - 18900fe380 [rush-lib] Test that a project configuration that failed to load is loaded afresh once it is fixed Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-cli-client] The daemon's admission and restart-drain reasons in full in agent output, and rushd's startup wait as a pipe line Swarm integration step 21; original commit 478f4d3af0 (merge of swarm/r04-t83-a81 at 4d1a045fca). Scope: task 83. Brings task 83's chain as resolved onto a81a00520d: 244ed8b1f5 (merge of swarm/r04-t83-clip @9d71dbe9bb, the clip fix for board 1539) and 4d1a045fca, with board 711 (1886d80936), tasks 79, 71 and 50 and the restart-drain timeout. Gate: ch01 GATE OK board 1872 (lifts board 1778); combined with task 93, tree d0e8df0156, full suite board 1888. Commits folded into this step (21): - 2e7d278d8f [rush-daemon] Report the full log path of failed and warning operations - 3b91eec856 [rush-cli-client] Keep spawned test clients on the default output when tests run in an agent shell - 1ae03fc555 [rush-cli-client] Agent output that holds at repository scale - 862ed5e676 [rush-cli-client] Do not report a daemon shutdown as a user cancellation - 39cae5f22b [rush-cli-client] Send no terminal width for a pseudo-terminal without a size - 5f3c36fe17 [rush-daemon] Describe a rejected request without the debug messages of the graph load (task 71) - 563e541fdf [rush-cli-client] Cap a multi-line error message in agent output (task 71) - 2377740ab9 [rush-daemon] Keep the loaded graph when a request only selects an unknown project (task 71) - e382d2ac80 [rush-lib] Run the initial script for every command outside watch mode (task 50, board 115) - b7bba00fb3 [rush-cli-client] Never leave agent output silent on a pipe (task 79) - feacc54bf7 [rush-daemon] Report a restart drain as a queue position (task 79) - 1886d80936 [rush-cli-client] Report Ctrl+C during graph preparation as cancelled (board 711) - ae31023aad [rush-cli-client] Report an empty selection in agent output (board 634) - bacfa3d5af [rush-daemon] Do not apply the client-default queue timeout to a restart drain (task 83) - 5f1c7b4f1b [rush-cli-client] Keep "daemon admission failed (<code>)" in the agent summary; word an empty selection like native Rush (board 785, board 634) - 3b5dc578d6 [rush-daemon] Document the restart drain and its client-default timeout (task 83) - cbf8a3a45e [rush-daemon] Apply the client-default timeout to a restart drain while a rushx script runs or for later requests (task 83) - 6cd768c6b3 [rush-daemon] Suggest stopping the rushx script when a restart drain times out behind one (task 83) - b2f9d78f1f [rush-daemon] Name the waived time when a restart drain times out (task 83) - 9d71dbe9bb Write the daemon's reason for an admission failure in full in agent output (task 83, board 1539) - 4d1a045fca Write rushd's startup wait as a pipe line in agent output (task 83 on task 95) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-cli-client] Report failed operations as they fail in agent output Swarm integration step 22; original commit 6bc6dd8f83 (merge of swarm/r04-t108-a81 at 14faa4e5be). Scope: task 108. Brings task 108 items 1 and 3 onto task 83: agent output prints a failed operation's first diagnostics as soon as that operation finishes, no `running` line before 10 s for quick requests, and each TypeScript error once (t07 board 1930: L4 17 of 41 on 4d1a045fca against 24 of 41 on s11; 38 of 41 with items 1 and 3 on their old base, board 1692). Gate: ch01 GATE OK board 1980 (tree 9079647277) Commits folded into this step (1): - 14faa4e5be Report failed operations as they fail in agent output (task 108 items 1 and 3) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Keep results that a plugin found up to date in long-lived operation graphs Swarm integration step 23; original commit e16ed01d3a (merge of swarm/r01-t85-a81 at 4066e8f5d7). Scope: task 85. Brings task 85 onto task 108 items 1+3: the daemon keeps a result that a plugin such as fstrace found up to date across requests, and checks a Skipped result again when it was marked unverifiable. m01 CONFIRMED the gated tip (board 2024: 72 of 72 hot requests "1066 no op") and measured 24 of 24 hot B6 requests as "no operations needed" (board 2073); ch04 CONFIRMED 6e39ea6527 (board 2003) and posted the product-default A/B (board 2132: -0.24 s [-0.57, +0.09]). ch01's combined suite on this tree passes (board 2150: rush-lib 1169/1, a timing flake in InstallRunScripts.test that task 85 doesn't touch, 5 of 5 on rerun; rush-daemon 515/0, rush-cli-client 367/0). Gate: ch01 GATE OK board 1955, combined suite board 2150 (tree 1c6a3bfa10) Commits folded into this step (2): - 95c8ece420 [rush-lib] Retain results that a plugin found up to date in long-lived operation graphs - 4066e8f5d7 [rush-lib] Check a Skipped result again if it was marked unverifiable Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Pass getOperationEnvironment to the before/after iteration hooks (and 7 more) Swarm integration step 24; original commit 31a8db3e65 (merge of swarm/r01-t85-x106b at 4835a8700e). Scope: tasks 85 and 106. Brings task 106 (guarded :incremental builds in the daemon) and task 84 at de9a503a50 through r01's 85x106 resolution. This is step 1 of the path in board 2046 and board 2161; step 2 brings task 106's test-only tip b767c64dac. Merging b767c64dac straight onto task 85 conflicts in CacheableOperationPluginRetainedResults.test.ts. Gate: ch01 GATE OK board 2161 (tree 8250469c70) Commits folded into this step (8): - c7400b3f9c [rush-lib] Pass getOperationEnvironment to the before/after iteration hooks - e07159a38f [rush-lib] Let the daemon run :incremental scripts when a guard allows it - 90a76234e5 [rush-lib] Register RUSH_DAEMON_INCREMENTAL_BUILDS as a known environment variable - b1647a952a [rush-lib] Test that --only consumers of a retained incremental result skip cache writes - aed31544c2 [rush-lib] Test that a terminated command drops its incremental base - 69e33c03a8 [rush-lib] Format the incremental build changes with Prettier - c3afda3587 [rush-lib] Test that a run whose inputs changed while it ran leaves no incremental base - de9a503a50 [rush-lib] Check dependencies reached through same-project phases in the incremental guard Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Let the daemon run :incremental scripts when a guard allows it Swarm integration step 25; original commit 7d0ffb955b (merge of swarm/r06-t106 at b767c64dac). Scope: task 106. Brings task 106's test-only tip onto step 1 (31a8db3e65). With task 106 the daemon runs an operation's <phase>:incremental script, on top of the outputs of its last successful run in the daemon, when only files that the operation builds were edited since then and its output folders are unchanged; anything else runs the initial command. A result built on an incremental base is never written to the build cache. Kill switch: daemon.incrementalBuilds false in rush.json, or RUSH_DAEMON_INCREMENTAL_BUILDS=0. Task 84 passes getOperationEnvironment to the before/after iteration hooks. Second agents: t03 CONFIRMED (board 2067, 37 of 39; TAMPERI is the declared limitation) and t04 CONFIRMED (board 2100). ch01 on this tree (board 2161): rush-lib 1228/0, rush-daemon 516/0, rush-cli-client 367/0; 50 mutants, 42 killed; 33 of 33 end-to-end rows against a native cache-off oracle. Known costs, per ch01's rulings in board 2161: revisiting a state restores less from the cache (Rule 2), and an in-place rewrite of an output file by another writer is not detected (TAMPERI; task 139, r06's 641c39f120 in board 2171). Gate: ch01 GATE OK board 2161 (tree b1c0c2e690) Commits folded into this step (1): - b767c64dac [rush-lib] Test that an incremental result taints cache writes through an operation without a script Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Test that a configuration is read again while only its rig.json may still be changing Swarm integration step 26; original commit 8a1d1ee8dd (merge of swarm/r07-t93-nit at f5826544d1). Scope: task 93 NIT. Test-only: 7 lines in RushProjectConfiguration.test.ts, no source change. It kills ch01's task 93 mutant M13 (board 1869): with the mutant the new assertion fails, 25/1, and 18900fe380's test file lets it survive, 26/0. Full rush-lib on the cherry-pick onto b1c0c2e690: 1228/0 (board 2178). The deployed daemon files don't change, so snapshot s14 (7d0ffb955b) is unaffected. Gate: ch01 GATE OK board 2178 (tree 09d90522de) Commits folded into this step (1): - f5826544d1 [rush-lib] Test that a configuration is read again while only its rig.json may still be changing Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] The incremental guard detects an in-place rewrite of an output file Swarm integration step 27; original commit 6eae8fadad (merge of swarm/r06-t106 at 641c39f120). Scope: task 139. Brings r06's fold-in on task 106 (board 2171), in place of b767c64dac. The output signature of an operation now records each output file's inode, size and mtime as well as the folder listing, so when another writer appends to or rewrites an output file in place, the next build runs the initial command ("its output folders changed since its last successful run") instead of an incremental run on the stale file. This closes t03's TAMPERI. A touch that changes only mtime now also costs one initial run. r06 measured about 4 µs per output file per read. ch01 on tree 1eb87abdd2 (board 2211): rush-lib 1231/0, rush-daemon 516/0, rush-cli-client 367/0; b767c64dac's manifest fails 4 of the new tests; 10 mutants, 6 killed (N1, N3 and N6 survive; task 138 adds their tests); tamper end-to-end rows 9/9 against 2/9 before, full rows 33/33, cache-off rows 7/7. On this branch the tree is 1eb87abdd2 plus r07's test-only 8a1d1ee8dd change, byte for byte. Gate: ch01 GATE OK board 2211 (tree 1eb87abdd2) Commits folded into this step (1): - 641c39f120 [rush-lib] Detect in-place rewrites of output files in the incremental guard Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] Tests for the incremental guard's configuration inputs, content-hashed outputs, invalidation and skip record Swarm integration step 28; original commit 0d30a5164c (merge of swarm/r06-t138 at dbd5d99147). Scope: task 138 and t106d. Brings r06's task 138 tests (board 2302) and t106d (4c9f23e650, board 2233). Only 3 test files change (+204/-18): ProductionDaemonRequestResolver.test.ts, IncrementalExecutionGuardPlugin.test.ts and OperationOutputManifest.test.ts. No product code changes, so no snapshot is needed. ch01 on tree 45c2c5edd5 (board 2364): rush-lib 1241/0 (1 skipped); rush-daemon 516/0 with t106d. With the new tests, each of the 10 guard and manifest mutants (G9-G12, G19, G24, L1, N1, N3, N6) fails its intended test; all 10 pass the old tests. t106d's kill-switch test fails against a product mutant that drops the daemon.incrementalBuilds check at PhasedScriptAction.ts:666, and 6eae8fadad's version of it passed that mutant. Gate: ch01 GATE OK board 2364 (tree 45c2c5edd5) Commits folded into this step (2): - 4c9f23e650 [rush-daemon] Make the test of daemon.incrementalBuilds: false edit a source file - dbd5d99147 [rush-lib] Test the incremental guard's configuration inputs, content-hashed outputs, invalidation and skip record Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A served rushx script releases the lifecycle gate once it starts, so reloads don't wait for it (board 113) (and 5 more) Swarm integration step 29; original commit 3857bc9926 (merge of swarm/r05-stack-int at c35a77cc5c). Brings r05's stack (board 1931): - Task 55: a served rushx script releases the lifecycle gate once it starts, so reloads don't wait for it. - Task 7 part 3 text: admission timeouts name the client's wait timeout and the time that isn't counted. - Task 96: a rushx script that arrives while a restart is pending waits for the restart instead of postponing it, and admission failures in legacy output and rushx-client print the daemon's reason. - Task 98: a request spends its wait timeout only while it waits for other requests. - The race fix (c35a77cc5c, restacked from 9370662d5f): a client that follows a daemon restart connects to the successor that the restarting daemon launches instead of racing it. ch01 on tree b500e451d4 (board 2429): build rc 0 with 0 warnings; rush-daemon 544/0, rush-lib 1231/0, rush-cli-client 374/0, rush-client-core 129/0 on its own (2 known flakes in the full run), protocol 166/0, transport 73/0, reporter 519/0, wire-e2e 4/0. The red checks fail as intended without each change. 36 of 43 mutants are killed; the 7 survivors are test gaps and change no claim. The end-to-end rows R55, R96 and R96A pass on the stack and fail on s14b's files. Gate: ch01 GATE OK board 2429 (tree b500e451d4 on 6eae8fadad; 87e0e3a10f together with task 138's merge, in either order) Commits folded into this step (6): - c8cd36f574 [rush-daemon] A served rushx script releases the lifecycle gate once it starts, so reloads don't wait for it (board 113) - 4a5a13bddf [rush-daemon] Name the client's wait timeout and uncounted time in admission timeouts - cafa954eb4 [rush-daemon] A rushx script that arrives while a restart is pending waits for it instead of postponing it (task 96) - 7ba7021ed6 [rush-cli-client] Admission failures in legacy output and rushx-client print the daemon's reason (task 96) - 4ddb1fdbfe [rush-daemon] Spend a request's wait timeout only while it waits for other requests (task 98) - c35a77cc5c [rush-client-core] A client that follows a daemon restart connects to the successor the restarting daemon launches instead of racing it (task 96) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A warm-set pass takes the repository lock only when it has work Swarm integration step 30; original commit a4de651852 (merge of swarm/r07-t76 at 6b32aa340e). Scope: task 76. Brings r07's task 76 (board 1730; o05 CONFIRMED board 1904). A warm-set maintenance pass takes the native repository lease only when it has something to evict, a watcher change to apply, or a failed observation to retry. An idle daemon no longer takes common/temp/rush#<pid>.lock every 30 s, and a native command that holds the lock no longer puts the warm set in 'native-busy'. Two commits: 40a9f8f5eb (product) and 6b32aa340e (tests for the two lease exemptions a pass relies on). ch01 on tree fe0b6a452a (board 2511): rush-daemon 549/0 (544 + 5 new tests). The old WorkspaceWarmSet.ts fails exactly the 5 new tests. 15 of 16 WorkspaceWarmSet.ts mutants are killed; B6 (no protected check in #shouldEvict) survives and is a test NIT. E2E on ch01's own daemons: 0 lock holds in 150 s idle (5 before), 0 daemon lock attempts while native holds the lock 40 s (12 before, with 'native-busy'). Gate: ch01 GATE OK board 2511 (tree fe0b6a452a) Commits folded into this step (2): - 40a9f8f5eb [rush-daemon] Warm set: take the repository lock only for a maintenance pass that has work - 6b32aa340e [rush-daemon] Warm set: test the two lease exemptions a pass relies on Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [node-core-library] A leftover LockFile whose pid now belongs to an older process is deleted Swarm integration step 31; original commit 02087b3f11 (merge of o07/lockfile-tzrule at d5e74b1813). Scope: task 121. Brings o07's task 121 (board 1815, re-tipped in board 1949; t05 CONFIRMED board 1923 on faa7675fa6). A rush#<pid>.lock whose pid has been reused by a live process that started before the file was written no longer blocks every command until that process exits. The lock file keeps its holder only if the recorded start time matches the live process's start time in some real time zone, and on Linux that start time is read from /proc instead of spawning ps. Four commits: 595f8f63c7 (the time zone rule), faa7675fa6 (/proc start time for another time zone or locale), 315e1eae81 (ch04's test for a process that started after the file) and d5e74b1813 (t05's note that a clock change can make a start time fail the rule). ch01 on tree 73696a280a (board 2566): build rc 0 with 0 warnings, rush-daemon 549/0; on b10e8c53e0 node-core-library 329/0 (318 + 11 new tests) and rush-client-core 129/0. The new tests fail 7 times on the old LockFile.ts. 21 LockFile.ts mutants, 16 killed; the 5 survivors are 4 test NITs and one equivalent. E2E with 12 live holders in other time zones and locales: all kept, with 0 ps calls after the first attempt (1.2-7 ms per attempt, against 410-520 ms before); the stale files (pid 1 at +7 min, pid 1 with a 2024 start time, a predating process at +3 min) are now deleted. Gate: ch01 GATE OK board 2566 (tree 73696a280a) Commits folded into this step (4): - 595f8f63c7 [node-core-library] Keep a mismatched LockFile only if its start time is in some time zone - faa7675fa6 [node-core-library] Read the start time from /proc for a LockFile with another time zone or locale - 315e1eae81 [node-core-library] Test a LockFile whose PID started after the file, with a start time in another time zone - d5e74b1813 [node-core-library] Note that a clock change can make a LockFile's start time fail the time zone check Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [credential-cache] A credentials.json that isn't valid JSON no longer prints the stored SAS Swarm integration step 32; original commit 0bef326909 (merge of swarm/r03-t143 at 92cbe2336e). Scope: task 143. Brings r03's task 143 (board 2316; t07 CONFIRMED board 2443 and board 2445). CredentialCache no longer rethrows JsonFile's parse error, which quoted jju's excerpt of the failing line, and in credentials.json that line holds the credential. The error now gives only the line and column and says the contents aren't shown because the file stores credentials. A missing file still gives undefined, other read errors are still wrapped as before, and schema errors are unchanged. The branch carries 08e26894e7 (task 101 part 1: compile the credentials file schema once) through r03's merge eae2fc9ad3, so tasks 101 and 143 merge in either order. ch01 on 2c425fc4f5 (board 2576): build rc 0 with 0 warnings, credential-cache 34/0, azure plugin 10/0, http plugin 14/0, amazon-s3 plugin 52/0. The new tests fail exactly the three torn-file tests on eae2fc9ad3. 12 CredentialCache.ts mutants, 9 killed; the survivors are one equivalent and one test NIT (a read error other than ENOENT). E2E with a marker in the SAS: on the old engine a torn file printed the marker in the daemon, --no-daemon and update-cloud-credentials paths; now it is printed 0 times, and 11 of 20 outputs are byte-identical. Task 142 must land with or after this merge. Gate: ch01 GATE OK board 2576 (tree 692db7d548 after task 121) Commits folded into this step (2): - 08e26894e7 [credential-cache] Compile the credentials file schema once per process - 92cbe2336e [credential-cache] Don't quote the credentials file in its parse error (task 143) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A failed agent-output build returns as soon as nothing unfinished can change its result Swarm integration step 33; original commit 1b51e7af08 (merge of swarm/r04-t108-int at c76de10302). Scope: task 108 a. Brings r04's task 108 item 2, "108(a)": 129e04f8bb on the nits 4bad455570 and item 2 2e066a7758, re-merged onto 3857bc9926 as c76de10302 (board 2133, board 2533). With agent output, a failed `build` returns as soon as nothing unfinished can change its result. Its summary line ends `· <n> independent operations continue in rushd`, and those operations finish in the daemon. The next build waits for them (`queued behind another request`) and doesn't run them again; the daemon's own install and rebuild stop the leftover, and a request that falls back in-process stops it first. Until the leftover ends, a native `rush install` or `rush-client --no-daemon` fails with "Another Rush command is already running in this repository." (o01 board 2598, board 2612), and `daemon status` doesn't count it yet (r04's output follow-up, board 2133, board 2652). Legacy output keeps the old behavior. o02 CONFIRMED item 2 with one row refuted (board 2055), which 129e04f8bb addresses; CONFIRMED by t07 (board 2183) and t05 (board 2193). ch01 on 140b42ad50 (board 2678): build rc 0 with 0 warnings, rush-daemon 573/0, rush-cli-client 383/0, rush-client-core 129/0, rush-daemon-protocol 172/0. 31 new and changed tests fail on the old product. 27 mutants, 22 killed; F4 and Y4 are equivalent, and three are test NITs (the stop ignores the client leaving, the offer ignores a record that is already aborted, and silent operations count as continuing). E2E: a failed build returns in 0.3 s where the base takes 15.3 s, and its independent operation finishes 15.0 s later. Gate: ch01 GATE OK board 2678 (tree 140b42ad50) Commits folded into this step (3): - 2e066a7758 Return a failed build's result early in agent output (task 108 item 2) - 4bad455570 Name a failure without output sooner, and pluralize failures, in agent output (task 108 nits) - 129e04f8bb Stop continuing work before an in-process fallback (task 108, coord board 1981 a) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A request that meets a removed or changed installation waits for the old daemon, then restarts it and runs Swarm integration step 34; original commit c0e17d82ab (merge of swarm/r03-t91f at 503c9e6589). Scope: tasks 125, 91 and 116. Brings r03's task 125 (E1 and E2 of t05 board 1726) with task 91 (a note restarts the 25 s status timer) and task 116 (long socket paths, board 1476), re-tipped onto integration 1b51e7af08 with ch01's resolution of the 108(a) conflicts in 2 files and one test mock line (board 2697, board 2732). When the daemon's installation is removed or changes, a request waits for the running requests to drain, the daemon restarts and the request runs, and the wait line names the installation as the reason; `--wait-timeout` and `--no-wait` fail with that reason. CONFIRMED by t05 on r03-t91d (board 2154, board 2197: the daemon side of cold E2 is 19 of 19 rc 0 where s13 fails 3 of 3). r05 reviewed the drain part of the re-tip (board 2680). ch01 on d9f459d151 (board 2756): build rc 0 with 0 warnings, rush-daemon 590/0, rush-cli-client 398/0, rush-client-core 136/0, rush-daemon-protocol 200/0, rush-daemon-transport 81/0. 33 new tests fail on the old product and 4 suites can't load. 14 mutants, 8 killed; the 6 survivors are test NITs. E2E: with the installation renamed away mid-build, the next build is rc 0 in 21.4 s where the base fails at once with `Cannot find module '@microsoft/rush-lib'`. r03's own run (board 2760): rush-cli-client 398/0, rush-daemon 590/0. Gate: ch01 GATE OK board 2756 (tree d9f459d151) Commits folded into this step (4): - 29b5678225 [rush-daemon] Refuse a socket path too long to reach, and keep older daemons in their folder (task 116, board 1476) - 08ed06de32 [rush-daemon] Restart a daemon whose installation was removed or replaced (task 91) - 1393c18030 [rush-daemon] Answer requests with the installation restart once the running requests finish (task 125) - c2218d0792 [rush-daemon] Wait for an installation restart in the restart drain, like a restart for an environment (task 125, aligned with task 83) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A daemon-served build logs one native telemetry entry per request, and the daemon keeps serving when its own PATH repeats an entry Swarm integration step 35; original commit d87702c813 (merge of swarm/r02-int at c9d60eb5a4). Scope: task 49 and the PATH fix. Brings r02's task 49 (daemon-served builds emitted no Rush command telemetry, board 61): a long-lived engine logs one native telemetry entry per served build request, with the request id, and an iteration's requests are reported before the next iteration starts. Engine disposal waits at most 2 seconds for telemetry uploads (8eccf3a6f1, CONFIRMED by ch05 board 1987). The daemon keeps serving after a plugin adds names to process.env (80704de4e3), and keeps serving when its own startup PATH repeats an entry, where it used to turn away every build (202412e57b, board 2657, CONFIRMED by o05 board 2671; the blobs carry over byte for byte, board 2776). Re-tipped onto integration c0e17d82ab with ch01's order for the 125 conflict in WorkspaceRequestLifecycle.ts (board 2756, board 2774); 108(a)'s 'running' early result counts as early in the telemetry entry. ch01 on b8b1dc6e2f (board 2795): build rc 0 with 0 warnings, rush-daemon 607/0, rush-lib 1249/0, rush-cli-client 398/0, rush-client-core 136/0, rush-daemon-protocol 200/0, rush-daemon-transport 81/0. Reverting the PATH fix or the 2 s flush wait each fails its test; 3 of 5 mutants killed, and the 2 survivors are test NITs for s16. E2E: the early row logs earlyResult true, and a PATH that repeats an entry keeps the same daemon where the old product falls back. r02's own run on c9d60eb5a4 (board 2777) gives the same counts. ch01's finding that ODSP_TELEMETRY_TAG restarts the daemon is pre-existing and is task 168. Gate: ch01 GATE OK board 2795 (tree b8b1dc6e2f) Commits folded into this step (10): - 490b3d94f9 [rush-lib] Extract the phased telemetry entry builder from OperationGraph - 251d987b06 [rush-lib] Let a long-lived engine log one native telemetry entry per request - 6f62ff03f2 [rush-daemon] Log one native telemetry entry per served build request - eb0fd9090e Add change files for per-request daemon telemetry - 80704de4e3 [rush-daemon] Keep serving after a plugin adds names to process.env - 6a17f4bbdc [rush-daemon] Test that an iteration's requests are reported before the next iteration starts - e7b848d6a7 [rush-daemon] Add the request id to each served request's telemetry entry - 8eccf3a6f1 [rush-lib] Wait at most 2 seconds for telemetry uploads when an engine is disposed - ddff7a24ff [rush-daemon] Move the native engine test fixture into its own module - 202412e57b [rush-daemon] Keep serving when the daemon's own PATH repeats an entry Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A stuck request no longer keeps a stopped daemon alive Swarm integration step 36; original commit 372742afe7 (merge of swarm/r01-t147-int at a40a901822). Scope: tasks 115 and 147. Brings r01's task 115 (`daemon stop` and the first SIGTERM never ended a daemon whose request was stuck in an await that ignores cancellation) and task 147 (after a clean stop, a ref'd timer or handle kept the process alive after it released its socket and pid.json; t05 board 2199, o04 board 2269), re-tipped onto integration (board 2779, board 2793). The daemon now shuts down within a 10 s deadline, returns a typed result to the running request and releases its socket, lockfile and repository lock. CONFIRMED: 115 by ch05 board 2247 and o04 board 2266; 147 by t05 board 2608 and o04 board 2632. s16 batch A, item 1 of 5. ch01 on the batch (board 2867): build rc 0 with 0 warnings; rush-daemon 628/0, rush-lib 1252/0, rush-cli-client 398/0 on items 1-4; 45 of 48 mutants killed; in the e2e the daemon is gone 10.46 s after `daemon stop`, where s15's tree lingers 39.15 s. Gate: ch01 GATE OK board 2867 (tree 24fc1e6446) Commits folded into this step (2): - fe430fe236 rushd: a shutdown stuck on an await that ignores cancellation no longer hangs the daemon process (task 115) - 444107370d rushd: a daemon process that something else keeps running exits 2 s after the daemon stops (task 147) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A hot request reuses the latest input capture that started after it was received Swarm integration step 37; original commit 7ab68d689d (merge of r05 at a5e20e7c20). Scope: task 137. Brings r05's task 137: in a hot burst each request ran its own full-workspace project-config capture (about 0.09 s each), so n=32 ran 32 of them back to back and held the event loop for 3.1-3.3 s (board 2101). A request now reuses the latest capture that started after the request was received. CONFIRMED by t01 board 2303. s16 batch A, item 2 of 5 (ch01 board 2867; the suites and mutants are in item 1's message). Gate: ch01 GATE OK board 2867 (tree 3b55869396) Commits folded into this step (1): - a5e20e7c20 [rush-daemon] Reuse the latest input capture that started after a request was received Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-azure-storage-build-cache-plugin] The Azure build cache picks up a credential written after the daemon's first cloud read Swarm integration step 38; original commit 0f5e259316 (merge of swarm/r03-t101 at f4a425ada3). Scope: task 101. Brings r03's task 101: the Azure storage build cache provider kept its first container client for the daemon's lifetime, so a credential written later was never used (ch05 board 1202, ch02 board 1229). It now uses the credential cached after the first cloud request. CONFIRMED by ch05 board 2290. s16 batch A, item 3 of 5 (ch01 board 2867). Gate: ch01 GATE OK board 2867 (tree 38dd18ec39) Commits folded into this step (1): - f4a425ada3 [rush-azure-storage-build-cache-plugin] Use a credential that is cached after the first cloud request (task 101) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] build-cache.json accepts azureBlobStorageConfiguration.storageEndpoint Swarm integration step 39; original commit ac2e51d2d2 (merge of swarm/r03-t144 at 04c7af3fbb). Scope: task 144. Brings r03's task 144: rush-lib's build-cache.schema.json rejected azureBlobStorageConfiguration.storageEndpoint, which the Azure plugin has read since microsoft/rushstack PR 5664. CONFIRMED by ch05 board 2396. s16 batch A, item 4 of 5 (ch01 board 2867). Gate: ch01 GATE OK board 2867 (tree 0fbdb27c44) Commits folded into this step (1): - 04c7af3fbb [rush-lib] Allow storageEndpoint in build-cache.json's azureBlobStorageConfiguration (task 144) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A lingering daemon logs its PID and stop time, without a stack Swarm integration step 40; original commit 86fc1ff37e (merge of swarm/r01-t147-nit at 7a76edd6ef). Scope: task 147 NITs. Brings r01's fix for the NITs in t05 board 2608 and o04 board 2632 (board 2814): the report is a timestamped log line through onLog, not an error with frames. t05 read the delta (board 2832), and r01's run through the real entry point is board 2851. ch01 on the final tree: build rc 0 with 0 warnings, rush-daemon 629/0, 7 of 7 mutants killed. s16 batch A, item 5 of 5 (ch01 board 2867). Gate: ch01 GATE OK board 2867 (tree d3a1eb9dbc) Commits folded into this step (1): - 7a76edd6ef rushd: a lingering daemon logs its PID and stop time, without a stack (task 147 NITs) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-cli-client] The client prints a cancelling line at once, agent output shows an operation's error after its own output, and a cancelled request reports never-started operations as not run Swarm integration step 41; original commit cba3c1f17b (merge of swarm/r04-t132-153-int at 4b495d5f81). Scope: tasks 132, 142 and 153. Brings r04's tasks 132 (after Ctrl-C, SIGTERM or SIGHUP during a daemon request the client prints one line at once, plus a notice when the daemon doesn't confirm), 142 (agent output prints a failed operation's error even when the operation wrote output first) and 153 (a cancelled request no longer reports never-started operations with an earlier request's result), re-tipped on d87702c813. Second agents: t05 board 2323 (132), t07 board 2623 (142), t05 board 2524 (153). ch01's e2e on ch01-sB: the cancelling line and the unconfirmed notice appear, and c is reported aborted, not SUCCESS. s16 batch B, item 1 of 5 (ch01 board 2984). Gate: ch01 GATE OK board 2984 (tree 6453dd938e) Commits folded into this step (3): - 85ed03e693 Say at once that a cancelled request waits for rushd to stop it (task 132) - 892c1592eb Print a failed operation's error even after its output in agent output (task 142) - 52b9d3fe3b Report never-started operations of a cancelled request as aborted (task 153) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Tests for a cancel while continuing work, and for the early failure offer Swarm integration step 42; original commit c4693204c1 (merge of swarm/r04-t108-nits-int at 97e473c08e). Scope: task 108(a) NITs. Brings r04's tests for the three gate NITs of task 108(a) (ch01 board 2678). Tests only; they pass on the old product, and ch01's mutants show they're live. s16 batch B, item 2 of 5 (ch01 board 2984). Gate: ch01 GATE OK board 2984 (tree b21b6b1f0b) Commits folded into this step (1): - 97e473c08e Test the three gate NITs of task 108 (a): a cancel while continuing work stops, and the early failure offer's silent and aborted operations (ch01 board 2678) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-client-core] The client relaunches the daemon after a startup helper exited, once its relaunch time passed Swarm integration step 43; original commit e073bd0f4b (merge of swarm/r03-t126 at da64013877). Scope: task 126. Brings r03's task 126: a client takes over a startup reservation whose helper is provably gone and whose daemon never became ready, instead of running in-process until `daemon stop --force` (t05 board 1574 and board 1652, row 13a). Second agent: t07 board 2751. ch01's e2e: the reservation is taken over after the relaunch time. s16 batch B, item 3 of 5 (ch01 board 2984). Gate: ch01 GATE OK board 2984 (tree a7707a430e) Commits folded into this step (1): - da64013877 rush-client-core: relaunch the daemon after a startup helper exited, once its relaunch time passed (task 126) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-client-core] After a daemon crash the client reaps its orphaned operations before it returns Swarm integration step 44; original commit 80219e0892 (merge of swarm/r06-t154-int2 at 25c7f0135c). Scope: task 154. Brings r06's task 154 (board 2382), re-tipped on c0e17d82ab (board 2747): after a daemon crash, the client reaps the orphaned operations before returning, and before `--no-daemon` runs. Second agent: t05 board 2689 on 14baba4338. ch01's e2e: the orphan is reaped after a crash. s16 batch B, item 4 of 5 (ch01 board 2984). Gate: ch01 GATE OK board 2984 (tree 2453314813) Commits folded into this step (2): - e79486fa37 [rush-client-core] Stop a crashed daemon's operations before reporting the crash (task 154) - 14baba4338 [rush-cli-client] Reclaim a crashed daemon before Rush runs in-process (task 154) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Test that an expired protected project takes no warm-set lease Swarm integration step 45; original commit 1b18fcc3eb (merge of swarm/r07-t76-nit at 6c6f638a2e). Scope: task 76 NIT. Brings r07's test for the task 76 NIT. Test only. s16 batch B, item 5 of 5 (ch01 board 2984). Gate: ch01 GATE OK board 2984 (tree 104ca5f27d) Commits folded into this step (1): - 6c6f638a2e [rush-daemon] Warm set: test that an expired protected project takes no lease Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [package-deps-hash] A hot request with no file change since the previous one skips the repo-state work Swarm integration step 46; original commit a9e930338f (merge of swarm/r07-t114-int at e883c78766). Scope: task 114. Brings r07's task 114 (daemon L3b, D20.3): RepoStateCache keeps the last repo-state capture, and a hot request whose files haven't changed since reuses it instead of running git status, ls-tree, the inputs snapshot and each operation's own-state hash again. It also moves assertCompatibleInputs and isGraphDefinitionPath into WorkspaceInputsComparison.ts with the same path list. Second agents: m01 board 2842 and t04 board 2897 on d0f571a420 (t114verify: 0 STALE rows, board 2829); addenda t04 board 3042 and m01 board 3031. ch01's e2e on ch01-sC: 30 of 30 steps equal to --no-daemon, 0 STALE. s16 batch C, item 1 of 7 (ch01 board 3099). Gate: ch01 GATE OK board 3099 (tree 8299cd3e10) Commits folded into this step (4): - 5fbb206eb3 [package-deps-hash] Add RepoStateCache for processes that read the state of a repository repeatedly - b7904dd6d6 [rush-lib] Keep the state of the repository and reuse unchanged inputs snapshots in long-lived hosts - 7cb4c46d89 [rush-daemon] Skip comparing unchanged inputs snapshots on each request - d0f571a420 [package-deps-hash] Copy the index again when the sizes, attributes, configuration or filter change Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] An invalid build command line fails with native Rush's error and exit code instead of running in-process Swarm integration step 47; original commit f84ce3a3cb (merge of swarm/r04-t78-int at cf5921e40f). Scope: task 78. Brings r04's task 78: when a warm daemon's parser rejects a build command line, the client prints the daemon's parse error once and exits 2, where it used to fall back in-process and print about 29 lines. Second agent: o04 board 2801. ch01's e2e on ch01-sC: a warm `build --nope` with agent output returns rc 2, 2 lines, in 0.16 s, answered by the daemon. s16 batch C, item 2 of 7 (ch01 board 3099). Gate: ch01 GATE OK board 3099 (tree d80d2a2ccb) Commits folded into this step (1): - cf5921e40f Fail an invalid build command line with native Rush's error and exit code instead of running it in-process (task 78) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Coalesced requests run with each requester's environment, and each gets its own telemetry key Swarm integration step 48; original commit b572eb50a0 (merge of swarm/r06-t90-int2 at 4b9a258212). Scope: tasks 81, 89 and 90. Brings r06's tasks 81 (coalesced requests no longer run once with the first requester's environment when an operation hashes an ignored variable they differ on), 89 (operations see process.env writes made by the same iteration's beforeExecuteIterationAsync, through a lazy per-participant environment) and 90 (a per-request telemetry key for operations in a coalesced batch), re-tipped on 1b51e7af08 (board 2747). Second agents: o04 board 783 and board 906, o05 board 1020, m03 board 1091 on adcb6f6e0f. It lands after r02's e7b848d6a7 (board 1166), which integration has. ch01's e2e on ch01-sC: two queued requests that differ in WT_SESSION run [ax, ay, c], as --no-daemon does. s16 batch C, item 3 of 7 (ch01 board 3099). Gate: ch01 GATE OK board 3099 (tree 8a64bfc100) Commits folded into this step (6): - 5565ff57ce [rush-daemon] Don't coalesce requests that disagree on a hashed per-session variable - 56ef144e1b [rush-lib] Pass getOperationEnvironment to the before/after iteration hooks - e4d4e90711 [rush-daemon] Let requests share a run when a hashed variable is unset in one and empty in the other - aec3ac8433 [rush-daemon] Start each operation from process.env as it is when the operation starts - 36d9a3a16f [rush-lib] Add getOperationRequestId to the operation graph's iteration options - 3099d5945f [rush-daemon] Give each operation the request id of the request whose environment it gets Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] A build-cache hit no longer restores an earlier failing run's stderr Swarm integration step 49; original commit c3f9fa2ac7 (merge of swarm/o03-t157 at 5b2a400f29). Scope: task 157. Brings o03's task 157: a cache hit no longer writes an older failing run's stderr into rush-logs/<name>.error.log, so a cache entry written after a failure isn't poisoned (ch02 board 2284, t05 board 2456). Second agent: t05 board 2704. ch01's e2e on ch01-sC: no error.log after either hit in the OK, ERR, OK sequence, and none in any cache entry. s16 batch C, item 4 of 7 (ch01 board 3099). Gate: ch01 GATE OK board 3099 (tree 09d5eabe82) Commits folded into this step (1): - 5b2a400f29 Keep cache hits from restoring an earlier run's stderr (task 157) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] An iteration checks an operation's build-cache inputs only when it needs them Swarm integration step 50; original commit 633640c81c (merge of swarm/r07-t152 at 7e98034976). Scope: task 152. Brings r07's task 152: an executed iteration no longer computes the build-cache disabled reason for all of the all-project graph's operations (about 0.7 s at odsp-web) or splices finished records out of the queue one at a time (about 0.1 s). Second agent: m01 board 2727. ch01's e2e on ch01-sC: the cache arm's hits and misses equal --no-daemon. s16 batch C, item 5 of 7 (ch01 board 3099). Gate: ch01 GATE OK board 3099 (tree 86cb10cdf4) Commits folded into this step (1): - 7e98034976 [rush-lib] Check the build cache inputs of an operation only when it's needed Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] A public beta API lets a plugin runner run :incremental under task 106's guard Swarm integration step 51; original commit e39062a035 (merge of swarm/r06-t131-int2 at 52c86b3174). Scope: task 131. Brings r06's task 131: rush-lib exposes a @beta API so that a plugin runner such as rush-fstrace-plugin can run an `:incremental` phase under task 106's L1a guard, re-tipped on 1b51e7af08 (board 2747). Second agent: o03 board 2360 on cb0891d79d. The test-file conflict in 7b8e8c665b keeps both sides. s16 batch C, item 6 of 7 (ch01 board 3099). Gate: ch01 GATE OK board 3099 (tree 07ac5f0139) Commits folded into this step (1): - cb0891d79d [rush-lib] Let operation runners from plugins use the incremental execution guard Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Answer a usage error only from a session that a build bound and checked Swarm integration step 52; original commit e12c8f3632 (merge of swarm/r04-t78-nit1 at 86a46e1d75). Scope: task 78 NIT 1. Brings r04's fix for o04's NIT 1 on task 78 (board 2801): the daemon answers a usage error only from a session that a build bound and checked. Second agent: o04 board 2948. s16 batch C, item 7 of 7 (ch01 board 3099). Gate: ch01 GATE OK board 3099 (tree 1af262606d) Commits folded into this step (1): - 86a46e1d75 Answer a usage error only from a session that a build bound and checked (task 78 follow-up, o04 board 2801 NIT 1) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [node-core-library] Test the edges of LockFile's time zone check Swarm integration step 53; original commit b122243826 (merge of o07/lockfile-tzrule-nits at 48fbc7ee88). Scope: task 121 NITs. Brings o07's fixes for ch01's 4 test NITs on task 121 (board 2972): tests, one doc sentence and a type-none change file. The branch is in <swarm path>, fetched read-only for this merge. s17 batch D1, item 1 of 10 (ch01 board 3310). Gate: ch01 GATE OK board 3104 and board 3310 (tree 20d42b41b6) Commits folded into this step (1): - 48fbc7ee88 [node-core-library] Test the edges of LockFile's time zone check Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon-protocol] Name the variables when a command's environment restarted the daemon Swarm integration step 54; original commit 898ff4dba1 (merge of swarm/r06-t88 at 04c6fb9d5a). Scope: task 88. Brings r06's fix for task 88 (board 2886): when a request restarts the daemon, or follows a restarted one, the client prints one line naming the differing variables (names only). Second agent: t07 board 2961. s17 batch D1, item 2 of 10. Gate: ch01 GATE OK board 3310 (tree ad9aa75f8c) Commits folded into this step (3): - ffbe43c152 [rush-daemon-protocol] Add the environmentChanged restart reason (task 88) - 8afd905246 [rush-daemon] Name the variables when the daemon restarts for an environment (task 88) - 04c6fb9d5a [rush-cli-client] Name the variables when a command's environment restarted the daemon (task 88) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] A restart drain's waiter rechecks its tier Swarm integration step 55; original commit 075d2130b2 (merge of swarm/r05-t127-int at 236c184d94). Scope: task 127. Brings r05's task 127 follow-up (board 2436, re-tipped onto d87702c813 in board 2858): a Restart-tier change that is reverted under load no longer fails its waiter at 30 s. Second agent: o05 board 2552. s17 batch D1, item 3 of 10. Gate: ch01 GATE OK board 3310 (tree bdfad6af20) Commits folded into this step (2): - 961f007cae [rush-daemon] A request that waits for a restart drain rechecks whether it still needs the restart (task 127) - 0b1585667a [rush-daemon] Requests that wait for a restart drain share one recheck capture a second (task 127) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] ODSP_TELEMETRY_TAG no longer restarts the daemon Swarm integration step 56; original commit 442af4ed14 (merge of swarm/r02-t168-int at b9b7f9dedd). Scope: task 168. Brings r02's task 168 (board 2939, re-tipped onto batch B in board 3030): setting, changing or unsetting ODSP_TELEMETRY_TAG keeps the warm daemon. Second agents: t07 board 3004 and m03 board 2975. s17 batch D1, item 4 of 10. Gate: ch01 GATE OK board 3310 (tree eadadd0318) Commits folded into this step (3): - 4a2af75d2b [rush-lib] [rush-daemon] Keep the warm daemon when a request's telemetry tag changes (task 168) - efb14d246a [rush-daemon] Log an early failed result's telemetry once its continuing work settles (task 168) - c888d079b3 [rush-lib] Test that the telemetry flush wait clears its timer (task 168) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] The daemon's failure summary names the error, not the first buffered warning Swarm integration step 57; original commit edbe3f9065 (merge of swarm/r04-t170-int at 6219ed8e65). Scope: task 170. Brings r04's task 170 (board 2953) with its NIT fix 9630e0bd88 (board 3036), re-tipped onto batch B. Second agent: t01 board 2966 and board 3045. s17 batch D1, item 5 of 10. Gate: ch01 GATE OK board 3310 (tree cbdb7a5bb7) Commits folded into this step (3): - 781d08b278 Give the error first in a rejected request's message, not a warning written before it (task 170) - f1d8a34c90 Put the in-process fallback on the first line of the reason (task 170) - 9630e0bd88 Drop a colon that ends the fallback reason's first line (task 170, t01 NIT) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-daemon] Check the installation again after a restart's waits Swarm integration step 58; original commit 13c4debd1f (merge of swarm/r03-t91g at b512743133). Scope: task 167. Brings r03's task 167 (board 2919): tests for ch01's surviving mutants on tasks 91 and 125, and rush-daemon checks the installation again after a restart's waits, before planning a successor. Second agent: t05 board 3064. s17 batch D1, item 6 of 10. Gate: ch01 GATE OK board 3310 (tree 6caa39900a) Commits folded into this step (1): - b512743133 rush-daemon: check the installation again after a restart's waits, before planning a successor (task 167) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-lib] A parameter whose short name the tool also defines can be used by its long name Swarm integration step 59; original commit 49695c5d00 (merge of swarm/r03-t172 at 4526357153). Scope: task 172. Brings r03's task 172 (board 3008): rush update-cloud-credentials --delete no longer fails with an ambiguous -d. Second agent: o03 board 3054. s17 batch D1, item 7 of 10. Gate: ch01 GATE OK board 3310 (tree 29b72dd480) Commits folded into this step (1): - 4526357153 [ts-command-line] A parameter whose short name the tool also defines can be used by its long name (task 172) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c * [rush-azure-storage-build-cache-plugin] Cache Azure credentials per storage endpoint Swarm integration step 60; original commit d950fba8d3 (merge of swarm/r03-t144b-x101 at 5e91eed724). Scope: task 164. Br…
Replace the 489 change files that this branch added with one change file for each of the 18 projects it changes, with shorter comments. Also delete 27 change files that came from main commits 1f51ac8..982b33b. Main has already published and deleted them, and the squash merge of #6102 dropped the merge of main that contained them, so merging this branch would add them back. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
Windows runners set core.autocrlf in the system Git config, which changed hashes in fixture repositories before tests changed their local config. Pin autocrlf and safecrlf in each fixture repository so the tests exercise their intended line-ending transitions. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
On both Windows CI jobs, RepoStateCache missed an untracked file in a new folder after it reused its copy of the index, while "git status" on the real index listed it. The untracked cache decides from the times of a folder whether to read it again, and Git for Windows can take those times from the listing of the parent folder, which is the likely cause. Without the untracked cache, "git status" reads every folder, as it does for the real index. The tests of how a new copy keeps the untracked cache of the previous copy don't apply on Windows, so they skip there. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
Restoring in place deleted every output folder at once after staged restore was skipped. On Windows, nested output folders and names that differ only by case can refer to the same tree, so parallel deletes race with each other and fail with EPERM. The staged restore test also uses the long spelling of the temp folder now. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The fixture copied every file that reports no data blocks as a sparse file. On Windows, NTFS reports no blocks for small files too, so their copies were all NUL bytes. The fixture now copies files as they are on Windows, uses the long spelling of the temp folder, and allows tar more time to exit after it is killed there. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The watch tests now create their repositories under the long spelling of the temp folder, since watching a folder through a short 8.3 path can crash Node.js on Windows. Cleanup retries while a process that the session started still keeps the folder busy, and the watch test ignores an iteration that a later change aborted, which Windows can report. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
On Windows runners, the temp folder has a short 8.3 spelling, while the product reports the long one. The native engine fixture now uses the long path, so expected paths match reported ones. The heft watch adapter test does too, since watching a short path can crash Node.js on Windows. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The test read an overlong name to make reading fail for a reason other than absence, but Windows reports such a name as missing. A name with a NUL character fails the same way on every platform. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The test made stopping slow with a detached process that kept the operation's output open, but on Windows the operation still ended at once. The test now holds the daemon's call that stops the operation until the test releases it, which works the same on every platform, and the fixture option that started the detached process is removed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
Windows lock files don't say which process holds them, so the daemon can't find the holder before it tries the lock. There, each background check that finds the lock held starts a preparation that stops at once, and a lock that the test process holds looks like another process's lock. The background preparation test accepts a later preparation number on Windows, and the usage error test gives its request a 1 s wait timeout and expects that timeout there. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The live owner test read the state and the command line of sleep just after the spawn event. On CI runners, sleep was still starting then: once it was running, and once it was running without its command line yet. The test expects a process that waits, so it now waits until sleep does. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The launch tests turn off every CI marker that the client knows, not only CI, because GitHub Actions sets GITHUB_ACTIONS and the client then ran Rush in-process. The fixture rush.json files suppress the Node.js LTS warning that Node.js 26 prints. On Windows, the slow launch tests, their cleanup and the first native build get more time, and the closed-pipe test is skipped, because libuv never closes file descriptors 0 to 2 there. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
…unners Before it exits, the fixture daemon now waits until its queue and start notices are written, as rushd does for the start notice. On Windows, a pipe finishes a write only on a later turn of the event loop, so the daemon lost the start notice and the client sent the request again. The test for a new daemon that did not start now waits until that daemon records itself, so the cleanup stops it before it removes the folder, which Windows keeps busy while the daemon runs there. The restart-always test gets a 5 s deadline, because a loaded CI machine did not start a successor within 1 s. To keep the test file within the lint limit of 2,000 lines, one assertion checks a shorter part of an error message, which then fits on one line. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
Test names and comments named tasks from a task list outside this repository, and two tests named a repository outside it for their scale. They now say only what they test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The numbers referred to a task list outside this repository. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
Test names and comments named tasks from a task list outside this repository, and a comment named a repository outside it for the shape of its graph. They now say only what they test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
Comments compared the test plugin with a plugin of a repository outside this one. They now describe the test plugin itself. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The helper checks that a wait timeout is an integer in the range that the protocol allows. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
CI fixesCI run 36754086279 on Linux and Windows
Windows only
Also in this push
ValidationOn Linux, the 8 changed projects build in production mode without warnings, and their tests pass on Node.js 22. The exception is heft's |
…ne test At its deadline, the registry lookup ends npm's whole process tree but waits only for the process that it started. In the test, that process is the npm shim and npm is its child, so npm could still be exiting when the test checked that it was gone (the Node.js 24 CI job on Ubuntu). In a local run of the same sequence, npm was still exiting when the shim's exit was seen in 34 of 100 rounds on Node.js 24 and 12 of 100 on Node.js 22. The test now waits up to 10 s for npm to exit. The fake npm otherwise hangs for two minutes, so the wait still shows that the lookup stopped it. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
An earlier commit in this branch built the path of a project dependency file with path.join, so that the new tests could expect a native path. On Windows, that also changed the path that the existing PnpmShrinkwrapFile tests record in their snapshots (the Node.js 24 and 26 CI jobs on Windows). Rush builds the path as before again, and the new tests expect it as Rush builds it: the project's temp folder, then a slash and the file name. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
Rush builds the path of a project dependency file from the project's temp folder, then a slash and the file name. The test now builds the path that it expects the same way. On Windows, joining the whole path puts a backslash before the file name, which does not match the path that Rush reports there. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
The Windows runners turn on core.autocrlf in the system Git configuration. Git then warned on stderr about each text file that the GitIndexFile tests and the submodule test of the RepoStateCache tests added, and the test phase fails when an operation warns (the Node.js 24 and 26 CI jobs on Windows). Those repositories now turn it off, as the main RepoStateCache test repository already does. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
CI fixes, round 2CI run 36771945207 on
The CodeQL check shows one annotation, the existing alert 35 ( ValidationOn Linux, the tests of the 8 projects that this PR changes pass in production mode on Node.js 22 with |
…second changes The test that changes the attributes after the previous copy failed in the Node.js 22 CI job on Ubuntu. The second changed between writing crlf.txt and commit(), so Git trusted the file in the index, but examined it again, under the new attributes, in the copy, which is a second older. The two states then reported different hashes for it. A time in the future makes Git examine the file in both. The documentation of RepoStateCache lists this case with the other rare differences from getDetailedRepoStateAsync. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
… test repository Each time the graph goes idle, the watcher takes a snapshot, and the session doesn't wait for it when it ends. The watch test removed its repository while Git was still reading it, so the snapshot provider wrote a warning to stderr, and Rush reported the tests with warnings, which failed the Node.js 24 CI job on Windows. The test now waits for every snapshot. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
CI fixes, round 3CI run 36778300935 on
The CodeQL check again shows only the existing alert 35. ValidationThe fixed package-deps-hash test passed 8 of 8 times, and all 182 package-deps-hash tests pass, also with a system Git config that turns on |
…n of the file system's clock Two tests of OperationOutputFingerprints failed in one run of the Node.js 24 CI job on Windows. Each adds a file to an output folder that the test wrote shortly before, and expects the modification time of the folder to change. Windows dates the change with the system time, which advances in ticks of up to 16 ms, so when both writes fell in the same tick, the folder kept its time and the stat check missed the new file. The tests now date the outputs an hour back, as if an earlier build wrote them. The documentation of OperationOutputFingerprints mentions this limitation of the stat check. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
…ends The Node.js 26 CI job on Windows failed with the warning that the previous commit fixed in the watch test, this time from the watch cancellation test in RushCommandLineParserReporterLifecycle.test.ts. The cause is in the watcher, not in the tests: when the graph goes idle, it starts a check for edits made during the iteration and doesn't wait for it, and it queues iterations the same way. Both take a snapshot, which could outlive the session, so Git still read the repository after the command returned and the test removed it. The watcher now keeps track of the work that it starts on its own, and once the session is aborted, the command waits for that work before it writes "Watch mode exited.". No new work starts after the abort. Instead of waiting for the snapshots itself, the watch test checks that none is running when the command returns. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6d3da49f-95f3-4291-9026-0b0f8e38675c
CI fixes, round 4CI run 36784157555 on
The CodeQL check again shows only the existing alert 35. ValidationTo emulate the Windows clock with a coarser one, a copy of the fingerprint tests rounded the modification times from |
CI, round 5All required checks pass on
The first attempt of Node.js 26, Linux failed in rush-reporter's "keeps a representative reporter workload within the wall-time regression budget". That test comes from main, and this PR doesn't change rush-reporter. It compares the wall time of a short workload with and without a reporter, taking the best of 7 pairs, against a 3% budget. It passed in every earlier CI job of this PR, so I re-ran the job, and the re-run passed. The CodeQL check again shows only the existing alert 35. |

Summary
After this change, contributors can build rushstack with the opt-in Rush daemon today, from source, on Linux and Windows. A warm no-op
rush-client buildis now faster than native Rush on both platforms.rush-client buildin this repo fell back to in-process Rush. The engine rejected any configured plugin, including ourrush-published-versions-json-plugin, which is only associated withrecord-published-versions.Results of the merged dogfooding fixes
#6102 is now merged into this branch. It fixes the daemon bugs found by an overnight dogfooding run in a large monorepo, and that run measured these results on one 192-core Linux host.
RUSHD_OUTPUT=agentRUSH_DAEMON=1cost per command, no compatible daemon releaseBefore is this branch without the relevant fixes, and after is with them. Each row compares the same scenario, measured the same way. Times are medians of repeated runs, except the failed-build row: a check in a small test repo where an independent 15 s operation is still running when the failure is known.
Dogfooding reported 138 daemon bugs. This merge fixes 130, fixes 4 more with a known residual, leaves 3 open and covers 1 with guidance. #6102 lists each one.
Changes
PhasedCommandEngine.parseAsyncnow uses Rush's own plugin loaders (PluginManager.getPluginsParticipatingInCommand). Configured plugins are allowed only if they are inert for the requestedbuild/rebuild. A plugin is still rejected if it is unassociated (initialized for every command), is associated with the command, or its command-line.json defines the command or one of its phases, or associates a parameter with either. The check fails closed: if a manifest or command-line file can't be read, the request is rejected.captureWorkspaceInputFingerprintAsyncnow also hashes each configured plugin's autoinstallerpackage.jsonand its cached manifest andcommand-line.json. Changing any of them reloads the generation.cmd.exeshell running each operation) its own visible window. Before serving, the launched daemon now makeswindowsHide: truethe default for anynode:child_processcall that doesn't set it. Each hidden child gets a windowless console, and its descendants inherit it. This applies only to the standalone Windows launch path..startingreservation. Every later client then refused the ready daemon, waited 15 s and fell back to native Rush. The helper now waits for a live launcher for at least 120 s. The fail-closed behavior is unchanged: the reservation is still kept if the launcher exits or never becomes ready.rush-clientno longer loads the Rush engine to findrush.json, read daemon settings or parserushxarguments. Those now live in small modules:RushJsonLocation,RushXCommandLineArgumentsandDaemonLaunchCommand. Version selection andrushxdiscovery are loaded only when needed.dist(which contain the implementation and weren't covered before), and it no longer runs a JavaScriptrealpathon every file for every request.PhasedCommandEngine.test.ts: plugin gate casesWindowsSubprocessConsoles.test.tsrush deployscenariocommon/config/rush/deploy-rush-daemon-dogfood.json. It extracts a self-contained copy of the built client closure into gitignoredcommon/temp/rush-daemon-dogfood.docs/rush/dogfooding-rush-daemon.md, linked from the root README, therush-cli-clientREADME andenvironment-variables.md. The client, daemon and client-core READMEs are updated.No CI workflow changes, and no
daemonblock inrush.json: opt-in isRUSH_DAEMON=1, set only forrush-clientinvocations.Performance
@rushstack/tree-pattern, one terminal, file caches warm, two identical rounds:rush-client daemon statusrebuildA rebuild is dominated by Heft compile time, which the daemon doesn't change. On Windows, one cold
rebuildpreviously opened 22 visible windows and took 85 s; it now opens 0.Validation
rush build --to @rushstack/rush-cli-clientpasses, including lint, with no API report changes.rush change --verifypasses.C:\workspaces\rdfclone, Node 26.7.0, Git 2.53 for Windows):graphInitialized: true, with the same PID and generation token across cold, warm, edit-and-rebuild and after-native buildsrush rebuild --only--no-daemonruns nativedaemon stopleaves no processestypescript-v4-testalso builds under the daemonRemaining bottlenecks and open risks
git hash-objectfailure. The failure (exit 0xC0000142, seen earlier under Jest) did not occur here, but its cause is unresolved. It may be related to each spawned tool allocating its own console, which this PR removes, but that is not established.