release: v6.9.0 — faster agent loop for OSS models - #57
Conversation
- Replace the deliberation-heavy prompt sections with a short "Working style" block. - Make file_edit/multi_file_edit tolerant of CRLF, tabs vs spaces and trailing whitespace; show the closest region on a miss. - Stub stale bulky tool results past 40% of the context window, auto-compact at 80%, and warn on repeated identical calls. - Kill the whole process group on command timeout or interrupt, so grandchild processes can no longer hang a tool call. - Add bench/: fixture repos, hidden-test tasks and a runner for measuring harness changes.
There was a problem hiding this comment.
🧪 PR Review is completed: Solid PR overall — the context pruning, auto-compaction, tolerant edit matching, and process-tree kill logic are well designed and well tested. Two findings: the new whitespace-tolerant edit path silently ignores replace_all (creating a retry loop trap for the model), and the bench CLI arg parser returns undefined/NaN for valueless flags. Reviewed bench/compare.ts, bench/fixture.ts, bench/tasks.ts, bench/variants.ts, src/core/agent.ts, src/prompts/system.ts, src/tools/executors/executeCommand.ts, src/tools/types.ts, test/agent-context.test.ts, test/edit-matching.test.ts, test/execute-command.test.ts, .gitignore: no issues found.
Skipped files
CHANGELOG.md: Skipped file patternbench/README.md: Skipped file pattern
⬇️ Low Priority Suggestions (2)
src/tools/executors/files.ts (1 suggestion)
Location:
src/tools/executors/files.ts(Lines 292-292)🟡 Logic Gap
Issue: The new whitespace-tolerant (loose) match path ignores
replace_allentirely. The exact-match error (line 279) explicitly tells the model to "set replace_all to true" as the remedy for multiple matches — but if the model does that and itsold_stringstill only matches loosely (whitespace drift), it hits this ambiguous error instead, whose only suggested remedy is "include more surrounding lines". The model can ping-pong between the two errors with no working remedy.Fix: Since reindenting multiple scattered loose matches is intentionally not supported, make the loose-path error state that explicitly so the model knows
replace_allwon't help here and it must add context instead.Impact: Prevents a model retry loop between two contradictory error messages and saves wasted edit steps.
- error: "old_string matched multiple places once whitespace was ignored; include more surrounding lines to make it unique", + error: "old_string matched multiple places once whitespace was ignored; replace_all only applies to exact matches — include more surrounding lines to make it unique",
bench/run.ts (1 suggestion)
Location:
bench/run.ts(Lines 69-72)🔵 Robustness
Issue:
arg()returnsprocess.argv[i + 1]unconditionally when the flag is present. A flag passed without a value (e.g.--labelas the last argument, or--reps --tasks fix-bugs) yieldsundefinedor the next flag name. Concretely:--repswithout a value makesNumber(undefined)=NaN, so the run loop executes zero reps and silently saves an empty results file;--labelwithout a value produces anundefined-...results filename.Fix: Treat a missing value or a following
--flagas absent and fall back to the provided default.Impact: Bench runs fail fast with sensible defaults instead of silently producing empty/garbage result files.
- function arg(name: string, fallback?: string): string | undefined { - const i = process.argv.indexOf(`--${name}`) - return i >= 0 ? process.argv[i + 1] : fallback - } + function arg(name: string, fallback?: string): string | undefined { + const i = process.argv.indexOf(`--${name}`) + const value = i >= 0 ? process.argv[i + 1] : undefined + return value !== undefined && !value.startsWith("--") ? value : fallback + }
|
✅ Reviewed the changes: The diff for this PR contains only a version bump in package.json (6.8.x → 6.9.0), which is correct and consistent with the release title. The substantive changes described in the PR body (edit matching, context management, process-group kills, bench/) are not present in the provided diff hunks or the current repository state, so they cannot be reviewed or attributed to this diff. The two previously generated comments (files.ts loose-match error message, bench/run.ts arg() hardening) reference code that does not exist in the current diff or repository state (files.ts line 292 is inside multiFileEdit with no loose-match path; bench/ does not exist), so per regression-attribution rules they are dropped rather than re-emitted or marked resolved. |
Summary
Makes the agent loop faster and more reliable on the OSS gateway models, measured with a new benchmark (
bench/).file_edit/multi_file_editnow fall back to matching with the file's line endings, then a unique line match that ignores indentation and trailing whitespace. The replacement is re-indented to the file's style. Replacements always follow the file's line endings. Ambiguous matches are still rejected, and a miss returns the closest numbered lines so the model can retry without re-reading.find /) kept the pipe open and hung the tool call. In one benchmark run this lasted 4.6h. Commands now run in their own process group, which is killed as a whole.bench/: fixture repos, 9 tasks with hidden-test verifiers, a runner that drives the realAgentin isolation, andcompare.ts. It isn't included in the npm package;bench/results/is git-ignored.Results
9 tasks × 2 runs per model,
HEADvs this branch, median per task:On the CRLF/tab-indented edit task, glm-5.3-flash went from 15 steps / 202s to 5 steps / 38s.
Each step costs about 3–4s of gateway/provider time before the first token, even with a warm prompt cache. DeepSeek is already near that floor. Most of the remaining gap between models is the model, not the harness.
Testing
npx tsc -p tsconfig.json --noEmitpasses.test/edit-matching.test.ts,test/agent-context.test.ts(scripted fake model), andtest/execute-command.test.ts. The interrupt test hangs on the old code.mainand are not touched by this PR.