Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -17,3 +17,4 @@ coverage/
.orb/AGENTS.md
.orb/links.json
.orbcode/AGENTS.md
bench/results/
32 changes: 32 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,38 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [6.9.0] - 2026-09-30

### Added

- **Whitespace-tolerant edits.** `file_edit` / `multi_file_edit` previously required `old_string` to match byte for byte, so a model that reconstructed indentation from memory (tabs vs spaces), or sent LF text for a CRLF file, got "old_string not found" and had to re-read the file and rewrite the edit — an extra model round trip plus another expensive edit-composition step. Matching now falls back, in order, to the same text with the file's line endings, then to a unique line-by-line match that ignores indentation and trailing whitespace (the replacement is re-indented to the file's own style). Replacement text always follows the file's line endings, so a CRLF file is never left with mixed endings. Ambiguous loose matches are still rejected, and successful loose matches say so in the tool result.
- **Actionable "not found" errors.** A failed match now returns the closest region of the file (up to 7 numbered lines with their exact whitespace), so the model can retry without another read.
- **Stale tool-result pruning.** Once context passes 40% of the model's window, bulky results of `read_file`, `search_files`, `list_files`, `execute_command`, `web_fetch` and `web_search` older than the four most recent tool results are sent as one-line stubs. The stored history and session files are untouched; only the outgoing request shrinks, and the boundary advances in batches of six so the request prefix (and the gateway's prompt cache) stays stable between prunes.
- **Automatic compaction.** When context passes 80% of the window, the conversation is summarized mid-turn and the turn continues from the summary, instead of degrading until the user runs `/compact`. A failed compaction is reported once and never retried within the session.
- **Loop warning.** The third identical tool call with identical output and no file edit in between gets an `[OrbCode]` note appended to its result telling the model that repeating it will not change anything. Previously this rule existed only as prompt text.
- **`bench/` harness benchmark.** Fixture repos, tasks with hidden-test verifiers, and a runner (`node --import tsx bench/run.ts --model <id> --label <name> [--reps N] [--suite core|extended|all] [--variant lean] [--steps]`) that drives the real agent loop in isolation and records steps, tokens, reasoning time, time-to-first-token, streaming and tool time, repeated calls and pass/fail. `bench/compare.ts` compares two result files. Not shipped in the npm package.

### Changed

- **Leaner system prompt.** The "Plan before editing", "Investigation efficiency" and "Verifying tool results and avoiding loops" sections asked the model to deliberate before every tool call and to write out a full change plan before editing. They are replaced by a four-line "Working style" block (act directly, locate → edit → check once, batch independent calls, never repeat an identical call more than twice).

### Fixed

- **Interrupts and timeouts now actually stop shell commands.** `execute_command` killed only the shell on timeout, so a grandchild process (`find /`, a pipeline stage) kept the output pipe open and the tool call hung until it exited on its own — in one benchmark run for 4.6 hours — and pressing Esc never stopped a running command at all. Commands now run in their own process group; a timeout or user interrupt kills the whole group, and the tool result says why the command stopped.

### Measured impact

Before/after on the 9-task benchmark (`bench/`, 2 runs per task per model, median per run; all runs pass unless noted):

| model | wall | steps | input tokens | pass |
|---|---|---|---|---|
| glm-5.3 | 122s → 81s (−33%) | 6 → 5 | −22% | 17/18 → 18/18 |
| glm-5.3-flash | 82s → 66s (−20%) | 5 → 5 | −12% | 18/18 → 18/18 |
| deepseek-v4.1-flash | 25s → 25s | 5 → 5 | −25% | 18/18 → 18/18 |
| gemini-3.8-flash | 72s → 73s | 11 → 11 | −15% | 17/18 → 17/18 |

The CRLF/tab-indented edit task shows the tolerant-edit change most clearly: glm-5.3-flash went from 15 steps / 202s to 5 steps / 38s, glm-5.3 from 12 steps / 275s to 6 steps / 53s.

## [6.8.7] - 2026-09-21

### Added
Expand Down
49 changes: 49 additions & 0 deletions bench/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
# Harness benchmark

Measures how the agent loop behaves on small, verifiable coding tasks so harness
changes (system prompt, tool schemas, edit tool, context handling) can be judged
on data instead of feel.

```sh
# baseline vs. a change: run each arm at least 3 times, ideally in parallel
node --import tsx bench/run.ts --model zai/glm-5.3 --label before --reps 3
node --import tsx bench/run.ts --model zai/glm-5.3 --label after --reps 3
node --import tsx bench/compare.ts bench/results/before-*.json bench/results/after-*.json
```

Flags: `--suite core|extended|all` (default `core`), `--tasks a,b`, `--steps`
(print each step's tools / input tokens), `--variant <name>` (system-prompt
override from `variants.ts`, for A/B tests before touching `src/`).

## What it does

- Every run gets a fresh throwaway git repo (`fixture.ts`) and the real `Agent`
loop with auto-approve on. `HOME` and the config dir are redirected to a temp
dir, so your sessions, `AGENTS.md` and skills never leak in; the login token is
read first and passed explicitly. Unknown model ids are rejected (they would
otherwise fall back to the default silently).
- Success is decided by tests the agent never saw (`tasks.ts`), plus checks such
as "test files untouched" or "CRLF and tabs preserved".
- Results are saved to `bench/results/` (git-ignored). Real model calls cost
money: the core suite is about $0.01 per run on `glm-5.3-flash` and about
$0.08 on `glm-5.3`.

## Suites

| suite | tasks |
|---|---|
| `core` | fix-bugs, rename, add-method, question, trivial-edit, multi-file-feature |
| `extended` | legacy-edit (CRLF + tabs), big-repo-bug (40 generated modules), long-feature (5-part change across layers) |

## Reading the numbers

- Wall time is dominated by the gateway (time to first token) and is very noisy:
the same task has taken 13s and 41s with identical code, and single steps have
stalled for minutes. Compare **medians over several reps**, and lean on
step count, input tokens and reasoning time, which are steadier.
- Runs that hit a gateway stall (`ECONNRESET`, the 10 minute timeout) fail for
reasons unrelated to the harness; check `detail` / `timedOut` before treating a
FAIL as a regression.
- Context pruning and auto-compaction only trigger at 40% / 80% of the model's
window (about 93k / 186k tokens), which these tasks do not reach; they are
covered by unit tests in `test/agent-context.test.ts` instead.
37 changes: 37 additions & 0 deletions bench/compare.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
/** node --import tsx bench/compare.ts <a.json> <b.json> — per-arm medians/sums and per-task medians. */
import * as fs from "node:fs"
import type { RunMetrics } from "./run.js"

const load = (f: string) => JSON.parse(fs.readFileSync(f, "utf8")) as { label: string; results: RunMetrics[] }
const median = (xs: number[]) => {
const s = [...xs].sort((a, b) => a - b)
const m = s.length >> 1
return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2
}
const [a, b] = process.argv.slice(2).map(load)
const metrics: [string, (r: RunMetrics) => number][] = [
["wall s", (r) => r.wallMs / 1000],
["think s", (r) => r.reasoningMs / 1000],
["ttft s", (r) => r.ttftMs / 1000],
["stream s", (r) => r.streamMs / 1000],
["steps", (r) => r.steps],
["input k", (r) => r.inputTokens / 1000],
]
const pct = (x: number, y: number) => (x === 0 ? "n/a" : `${(((y - x) / x) * 100).toFixed(0)}%`)
console.log(`${"".padEnd(10)} ${a.label.padStart(10)} ${b.label.padStart(10)} change (median per run, all tasks)`)
for (const [name, f] of metrics) {
const x = median(a.results.map(f))
const y = median(b.results.map(f))
console.log(`${name.padEnd(10)} ${x.toFixed(1).padStart(10)} ${y.toFixed(1).padStart(10)} ${pct(x, y)}`)
}
const sum = (rs: RunMetrics[], f: (r: RunMetrics) => number) => rs.reduce((t, r) => t + f(r), 0)
console.log(`\npass: ${a.label} ${a.results.filter((r) => r.pass).length}/${a.results.length} ${b.label} ${b.results.filter((r) => r.pass).length}/${b.results.length}`)
console.log(`total wall s: ${(sum(a.results, (r) => r.wallMs) / 1000).toFixed(0)} vs ${(sum(b.results, (r) => r.wallMs) / 1000).toFixed(0)} total think s: ${(sum(a.results, (r) => r.reasoningMs) / 1000).toFixed(0)} vs ${(sum(b.results, (r) => r.reasoningMs) / 1000).toFixed(0)}`)
console.log("\nper task (median wall s / think s / steps):")
for (const id of [...new Set(a.results.map((r) => r.task))]) {
const row = (x: { results: RunMetrics[] }) => {
const rs = x.results.filter((r) => r.task === id)
return `${median(rs.map((r) => r.wallMs / 1000)).toFixed(0).padStart(4)} / ${median(rs.map((r) => r.reasoningMs / 1000)).toFixed(0).padStart(3)} / ${median(rs.map((r) => r.steps)).toFixed(0)}`
}
console.log(`${id.padEnd(20)} ${row(a).padEnd(16)} ${row(b)}`)
}
Loading