Skip to content

release: v6.9.0 — faster agent loop for OSS models - #57

Merged
code-crusher merged 2 commits into
mainfrom
perf/harness-speed
Sep 30, 2026
Merged

code-crusher merged 2 commits into
mainfrom
perf/harness-speed

Conversation

@code-crusher

Copy link
Copy Markdown
Member

Summary

Makes the agent loop faster and more reliable on the OSS gateway models, measured with a new benchmark (bench/).

  • Leaner system prompt. The "Plan before editing", "Investigation efficiency" and "Verifying tool results" sections, which made models deliberate before every call, are replaced with a four-line "Working style" block.
  • Tolerant edits. file_edit / multi_file_edit now fall back to matching with the file's line endings, then a unique line match that ignores indentation and trailing whitespace. The replacement is re-indented to the file's style. Replacements always follow the file's line endings. Ambiguous matches are still rejected, and a miss returns the closest numbered lines so the model can retry without re-reading.
  • Context management. Once context passes 40% of the window, old bulky tool results are stubbed in the outgoing request; stored history is unchanged. The history is auto-compacted mid-turn at 80%. The 3rd identical call with identical output gets a loop warning.
  • Fix: commands couldn't be stopped. A timeout or Esc killed only the shell, so grandchildren (e.g. find /) kept the pipe open and hung the tool call. In one benchmark run this lasted 4.6h. Commands now run in their own process group, which is killed as a whole.
  • bench/: fixture repos, 9 tasks with hidden-test verifiers, a runner that drives the real Agent in isolation, and compare.ts. It isn't included in the npm package; bench/results/ is git-ignored.

Results

9 tasks × 2 runs per model, HEAD vs this branch, median per task:

model wall steps input tokens pass
glm-5.3 122s → 81s (−33%) 6 → 5 −22% 17/18 → 18/18
glm-5.3-flash 82s → 66s (−20%) 5 → 5 −12% 18/18 → 18/18
deepseek-v4.1-flash 25s → 25s 5 → 5 −25% 18/18 → 18/18
gemini-3.8-flash 72s → 73s 11 → 11 −15% 17/18 → 17/18

On the CRLF/tab-indented edit task, glm-5.3-flash went from 15 steps / 202s to 5 steps / 38s.

Each step costs about 3–4s of gateway/provider time before the first token, even with a warm prompt cache. DeepSeek is already near that floor. Most of the remaining gap between models is the model, not the harness.

Testing

  • npx tsc -p tsconfig.json --noEmit passes.
  • 19 new tests pass: test/edit-matching.test.ts, test/agent-context.test.ts (scripted fake model), and test/execute-command.test.ts. The interrupt test hangs on the old code.
  • Full suite: 77/92 pass. The 15 failures (Axon model registry and UI viewport tests) fail identically on main and are not touched by this PR.

- Replace the deliberation-heavy prompt sections with a short
  "Working style" block.
- Make file_edit/multi_file_edit tolerant of CRLF, tabs vs spaces and
  trailing whitespace; show the closest region on a miss.
- Stub stale bulky tool results past 40% of the context window,
  auto-compact at 80%, and warn on repeated identical calls.
- Kill the whole process group on command timeout or interrupt, so
  grandchild processes can no longer hang a tool call.
- Add bench/: fixture repos, hidden-test tasks and a runner for
  measuring harness changes.
@matterai-app

matterai-app Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Code Quality

Summary By MatterAI MatterAI logo

🔄 What Changed

Updates the package version in package.json to v6.9.0 to release the faster agent loop optimizations, tolerant file edits, context management improvements, and robust process execution features.

🔍 Impact of the Change

Official release bump for version 6.9.0, making the new agent enhancements available via the distribution package.

📁 Total Files Changed

Click to Expand
File ChangeLog
Package Config package.json Bumped version to 6.9.0.

🧪 Test Added/Recommended

Added

  • None required for version bump.

🔒Security Vulnerabilities

  • None detected.

⏳ Estimated code review effort

LOW (~5 minutes)

♫ Tanka Poem

Version six-nine-zero,
Faster loops and cleaner code,
OSS models soar,
Process trees now fully safe,
Agents run with steady grace.

Sequence Diagram

sequenceDiagram
    participant Release as Release Process
    participant PKG as package.json
    participant NPM as NPM Registry

    Release->>PKG: update version to 6.9.0
    PKG-->>Release: version updated
    Release->>NPM: publish v6.9.0 package
Loading

@matterai-app matterai-app Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧪 PR Review is completed: Solid PR overall — the context pruning, auto-compaction, tolerant edit matching, and process-tree kill logic are well designed and well tested. Two findings: the new whitespace-tolerant edit path silently ignores replace_all (creating a retry loop trap for the model), and the bench CLI arg parser returns undefined/NaN for valueless flags. Reviewed bench/compare.ts, bench/fixture.ts, bench/tasks.ts, bench/variants.ts, src/core/agent.ts, src/prompts/system.ts, src/tools/executors/executeCommand.ts, src/tools/types.ts, test/agent-context.test.ts, test/edit-matching.test.ts, test/execute-command.test.ts, .gitignore: no issues found.

Skipped files
  • CHANGELOG.md: Skipped file pattern
  • bench/README.md: Skipped file pattern
⬇️ Low Priority Suggestions (2)
src/tools/executors/files.ts (1 suggestion)

Location: src/tools/executors/files.ts (Lines 292-292)

🟡 Logic Gap

Issue: The new whitespace-tolerant (loose) match path ignores replace_all entirely. The exact-match error (line 279) explicitly tells the model to "set replace_all to true" as the remedy for multiple matches — but if the model does that and its old_string still only matches loosely (whitespace drift), it hits this ambiguous error instead, whose only suggested remedy is "include more surrounding lines". The model can ping-pong between the two errors with no working remedy.

Fix: Since reindenting multiple scattered loose matches is intentionally not supported, make the loose-path error state that explicitly so the model knows replace_all won't help here and it must add context instead.

Impact: Prevents a model retry loop between two contradictory error messages and saves wasted edit steps.

-  			error: "old_string matched multiple places once whitespace was ignored; include more surrounding lines to make it unique",
+  			error: "old_string matched multiple places once whitespace was ignored; replace_all only applies to exact matches — include more surrounding lines to make it unique",
bench/run.ts (1 suggestion)

Location: bench/run.ts (Lines 69-72)

🔵 Robustness

Issue: arg() returns process.argv[i + 1] unconditionally when the flag is present. A flag passed without a value (e.g. --label as the last argument, or --reps --tasks fix-bugs) yields undefined or the next flag name. Concretely: --reps without a value makes Number(undefined) = NaN, so the run loop executes zero reps and silently saves an empty results file; --label without a value produces an undefined-... results filename.

Fix: Treat a missing value or a following --flag as absent and fall back to the provided default.

Impact: Bench runs fail fast with sensible defaults instead of silently producing empty/garbage result files.

-  function arg(name: string, fallback?: string): string | undefined {
-  	const i = process.argv.indexOf(`--${name}`)
-  	return i >= 0 ? process.argv[i + 1] : fallback
-  }
+  function arg(name: string, fallback?: string): string | undefined {
+  	const i = process.argv.indexOf(`--${name}`)
+  	const value = i >= 0 ? process.argv[i + 1] : undefined
+  	return value !== undefined && !value.startsWith("--") ? value : fallback
+  }

@code-crusher code-crusher changed the title perf(harness): faster agent loop for OSS models release: v6.9.0 — faster agent loop for OSS models Sep 30, 2026
@code-crusher
code-crusher merged commit d57bfb0 into main Sep 30, 2026
1 check was pending
@code-crusher
code-crusher deleted the perf/harness-speed branch September 30, 2026 03:02
@matterai-app

matterai-app Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

✅ Reviewed the changes: The diff for this PR contains only a version bump in package.json (6.8.x → 6.9.0), which is correct and consistent with the release title. The substantive changes described in the PR body (edit matching, context management, process-group kills, bench/) are not present in the provided diff hunks or the current repository state, so they cannot be reviewed or attributed to this diff. The two previously generated comments (files.ts loose-match error message, bench/run.ts arg() hardening) reference code that does not exist in the current diff or repository state (files.ts line 292 is inside multiFileEdit with no loose-match path; bench/ does not exist), so per regression-attribution rules they are dropped rather than re-emitted or marked resolved.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant