Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
143 changes: 143 additions & 0 deletions docs/ai/design/2026-09-22-feature-session-compact-jev.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,143 @@
---
phase: design
title: Jev Session Compaction Design
description: Architecture for typed session-event classification and deterministic compact artifacts
---

# Jev Session Compaction Design

## Architecture Overview

The feature extends the existing `agent session` command group. The command owns session selection and output. A small session-compaction service owns availability, redaction, Jev classification, deterministic assembly, and rendering. Provider-specific transcript parsing stays in `agent-manager`.

```mermaid
flowchart TD
CLI[agent session compact] --> Gate{TYPESAFE_API_KEY present?}
Gate -->|no| Unavailable[Jev unavailable result]
Gate -->|yes| Sessions[AgentManager.findSessionsById]
Sessions --> Resolve[Resolve exact ID and optional type]
Resolve --> Adapter[Adapter.getConversation verbose]
Adapter --> Redact[Local credential redaction]
Redact --> Jev[Bounded Jev worker pool]
Jev --> Filter[Exclude discard, irrelevant, sensitive]
Filter --> Builder[Deterministic compact builder]
Builder --> Markdown[Markdown renderer]
Builder --> JSON[JSON serializer]
```

## Data Models

```ts
type CompactCategory =
| "user_instruction"
| "decision"
| "code_change"
| "command_evidence"
| "validation_evidence"
| "blocker"
| "next_step"
| "memory_candidate"
| "discard";

type CompactImportance = "irrelevant" | "useful" | "important" | "critical";

interface ClassifiedSessionEvent {
role: "user" | "assistant" | "system";
content: string;
timestamp?: string;
keep: boolean;
category: CompactCategory;
importance: CompactImportance;
sensitive: boolean;
}

interface SessionCompact {
intent: string;
currentState: string;
decisions: string[];
changedFiles: string[];
commands: string[];
validation: string[];
openQuestions: string[];
nextStep: string;
memoryCandidates: string[];
resumePrompt: string;
jev: { available: true; model: string; classifiedEvents: number };
}

type JevUnavailable = {
jev: { available: false; reason: "TYPESAFE_API_KEY is not set" };
};
```

The builder preserves source wording rather than claiming generated facts. It uses the first retained user instruction as intent, the latest retained operational events as current state, the latest next-step event as next step, and category groups for array fields. Blockers populate open questions. The resume prompt is a fixed template composed from those fields.

## API Design

### CLI

```text
ai-devkit agent session compact --id <session-id> [--type <type>] [--format markdown|json]
```

- Default format: `markdown`.
- Unavailable Markdown is the single required sentence.
- Unavailable JSON is the `JevUnavailable` object.
- Availability is checked before `AgentManager.findSessionsById()` so no transcript is read without a usable configuration signal.

### Internal boundaries

```ts
interface SessionEventClassifier {
readonly model: string;
classify(message: ConversationMessage): Promise<ClassifiedSessionEvent>;
}

async function compactSession(
messages: ConversationMessage[],
classifier: SessionEventClassifier,
): Promise<SessionCompact>;

function renderSessionCompactMarkdown(result: SessionCompact): string;
```

The default classifier wraps `@typesafe-ai/sdk` and asks one `choice` question for category, one `choice` for importance, and two `noul` questions for retention and sensitivity. It validates returned labels before constructing a classified event.

## Component Breakdown

- `packages/cli/src/services/session-compact/session-compact.types.ts`: stable feature types and enums.
- `packages/cli/src/services/session-compact/jev-classifier.ts`: SDK adapter and response validation.
- `packages/cli/src/services/session-compact/session-compact.service.ts`: redaction, filtering, deterministic assembly, and Markdown rendering.
- `packages/agent-manager/src/AgentManager.ts` and built-in adapters: exact session-ID resolution using provider-native storage, with a compatibility fallback for external adapters.
- `packages/cli/src/commands/agent.ts`: Commander wiring, early availability gate, exact session resolution, output selection.
- `packages/cli/src/__tests__/services/session-compact/*`: classifier/service unit tests with no network.
- `packages/cli/src/__tests__/commands/agent.test.ts`: command-level unavailable and successful wiring tests.
- `skills/session-compact/SKILL.md` and `skills/built-in.json`: agent instructions and distribution manifest.

## Design Decisions

1. **Use `agent session compact`, not a new top-level `session`.** Historical discovery and detail already live under `agent`; extending that namespace minimizes concepts and shares resolution behavior.
2. **Select by session ID, not raw input path.** Adapters already own heterogeneous file/database formats and normalize them to `ConversationMessage[]`.
3. **No silent fallback.** Missing configuration is a typed, successful unavailable result; API/runtime failures are errors.
4. **Deterministic output after classification.** Jev is a decision model, not a prose generator. Keeping source content also makes the artifact auditable.
5. **One injectable classifier boundary.** Tests avoid network access and the SDK can be replaced without changing the builder or CLI.
6. **Bounded classification concurrency.** Eight workers overlap independent Jev calls; each result is written to its source index so deterministic artifact ordering is preserved. A bounded pool avoids the request spike of unbounded `Promise.all`.
7. **Redact before Jev and exclude Jev-sensitive events.** This reduces exposure and prevents sensitive events from entering the artifact, while acknowledging regex redaction is not comprehensive DLP.
8. **Direct lookup is adapter-wide and additive.** All built-in adapters implement exact-ID lookup using their native filesystem or database shape. The interface method is optional so external adapters continue to work through manager-level list-and-filter fallback.

### Rejected alternatives

- Top-level `session compact`: duplicates the existing `agent session` resource namespace.
- `--input <path>` first: leaks provider parsing into the command and cannot naturally represent database-backed sessions.
- Put compaction into `agent-manager`: provider-independent classification and presentation are CLI application concerns.
- Add a general AI-provider layer: no second current caller.
- Ask Jev to generate the artifact: unsupported by its typed-decision role.

## Non-Functional Requirements

- Do not log, serialize, interpolate into errors, or transmit `TYPESAFE_API_KEY` as state.
- Redact common bearer tokens, private-key blocks, and credential assignments before classification.
- Validate all Jev labels at the integration boundary; never coerce unknown labels to a normal category.
- Preserve script-friendly stdout. Normal JSON/Markdown output goes to stdout; runtime errors follow existing stderr/exit-1 handling.
- No filesystem writes by the command.
- Test all new branches without live credentials or network calls.
88 changes: 88 additions & 0 deletions docs/ai/implementation/2026-09-22-feature-session-compact-jev.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
---
phase: implementation
title: Implementation Guide
description: Technical implementation notes, patterns, and code guidelines
---

# Jev Session Compaction Implementation

## Development Setup

**How do we get started?**

- Worktree: `.worktrees/feature-session-compact-jev` on `feature-session-compact-jev`.
- Install with `npm ci`; build with `npm run build`.
- Runtime Jev access requires `TYPESAFE_API_KEY`. Tests never require a live key.
- Official SDK dependency: `@typesafe-ai/sdk@0.6.0` in the CLI workspace.

## Code Structure

**How is the code organized?**

- `services/session-compact/session-compact.types.ts`: compaction domain types and unavailable constants.
- `services/session-compact/session-compact.service.ts`: local redaction, filtering, assembly, and Markdown rendering.
- `services/session-compact/jev-classifier.ts`: official SDK adapter and typed answer validation.
- `__tests__/services/session-compact/`: network-free service and adapter tests.

## Implementation Notes

**Key technical details to remember:**

### Core Features

- Messages are locally redacted, classified by an eight-worker bounded pool, filtered, and grouped without generative rewriting. Indexed result placement preserves source order.
- Jev asks typed category/importance choice questions and retention/sensitivity noul questions.
- Unknown category or importance labels throw explicit integration errors.

### Patterns & Best Practices

- `SessionEventClassifier` is the only injectable boundary needed by the deterministic service.
- `createJevSessionEventClassifier` constructs the production SDK client with logging disabled.
- Threshold `>= 0.5` converts Jev noul probabilities to keep/sensitive booleans.
- Category/importance labels and both noul probabilities are validated at the SDK boundary; blank optional model configuration normalizes to `jev-latest`.

## Integration Points

**How do pieces connect?**

- `TypeSafeClient.systemOne` uses `jev-latest` unless `TYPESAFE_DEFAULT_MODEL` is configured.
- No database or persistent state is added.
- CLI wiring will pass adapter-normalized `ConversationMessage[]` to `compactSession`.
- `agent session compact --id <id> [--type <provider>] [--format markdown|json]` is registered beside historical session detail.
- The missing-key branch returns before `createAgentManager`, so it cannot list sessions or read a transcript.
- Successful resolution calls `AgentManager.findSessionsById`. Every built-in adapter implements provider-native lookup; optional adapter-method semantics preserve list-and-filter compatibility for external adapters.
- Codex, Claude, Grok, Copilot, Pi, and OpenCode exploit ID-addressable paths or SQL. Gemini metadata does not encode IDs in filenames, so it scans until the exact embedded ID is found without building summaries for later files.

## Error Handling

**How do we handle failures?**

- Missing key will be handled before constructing the SDK client or reading a transcript.
- SDK/API failures propagate through existing CLI error handling; there is no fallback.
- SDK logging is disabled so request bodies cannot be emitted by this feature.

## Performance Considerations

**How do we keep it fast?**

- Classification uses eight concurrent requests; the fixed internal limit keeps the public service contract small.
- Each source message produces exactly one Jev request.

## Security Notes

**What security measures are in place?**

- The API key is passed only to `TypeSafeClient` and never put in Jev state.
- Common bearer tokens, private-key blocks, and secret assignments are redacted locally.
- Jev-sensitive events are excluded from the compact artifact.
- Redaction is defense in depth, not complete DLP.

## Delivered Skill

- `skills/session-compact/SKILL.md` routes long-context, handoff, stale-resume, complex-close, memory-candidate, and task-progress use cases to the CLI.
- It requires agents to report missing `TYPESAFE_API_KEY` clearly and forbids claiming a manual fallback was Jev-backed.
- `skills/built-in.json` and the offline fallback list both include the skill.

## Design Alignment

The implementation now includes the measured performance follow-up: adapter-wide exact-ID lookup and bounded concurrent classification. It adds no persistent index/database state, automatic memory/task mutation, fallback summarizer, or output-file behavior.
124 changes: 124 additions & 0 deletions docs/ai/planning/2026-09-22-feature-session-compact-jev.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
---
phase: planning
title: Jev Session Compaction Plan
description: Ordered implementation and validation tasks for session compaction
---

# Jev Session Compaction Plan

## Milestones

- [x] Milestone 1: Typed compaction domain and Jev adapter are covered by unit tests.
- [x] Milestone 2: `agent session compact` implements explicit unavailable and successful output flows.
- [x] Milestone 3: Built-in skill, lifecycle docs, and repository verification are complete.
- [x] Milestone 4: Measured performance bottlenecks are addressed with adapter-wide direct lookup and bounded Jev concurrency.

## Task Breakdown

### Phase 1: Test-first service foundation

- [x] Task 1.1: Add the official `@typesafe-ai/sdk` CLI dependency.
- Outcome: reproducible Jev SDK integration on Node 20.
- Dependencies: none.
- Evidence: lockfile diff and CLI package build.
- [x] Task 1.2: Write failing tests for compaction filtering, grouping, resume-prompt assembly, Markdown headings, and secret redaction.
- Outcome: deterministic artifact contract is executable before production code.
- Dependencies: requirements/design schemas.
- Evidence: focused Vitest failure for missing modules/behavior.
- Covers: service and security scenarios in the testing strategy.
- [x] Task 1.3: Implement compaction types and service until Task 1.2 passes.
- Outcome: provider-independent, network-independent compaction core.
- Dependencies: Task 1.2.
- Evidence: focused service tests pass.
- [x] Task 1.4: Write failing Jev adapter tests, then implement the injectable SDK boundary and strict answer validation.
- Outcome: typed questions map to validated domain events without exposing credentials.
- Dependencies: Tasks 1.1 and 1.3.
- Evidence: mocked-classifier tests pass; no network calls.

### Phase 2: CLI integration

- [x] Task 2.1: Write failing command tests for Markdown/JSON unavailable results, early key gate, format validation, session resolution, and successful rendering.
- Outcome: user-visible semantics are locked before wiring.
- Dependencies: Phase 1.
- Evidence: focused command tests initially fail for absent command, then pass.
- [x] Task 2.2: Register `agent session compact` using existing session ID/type resolution and verbose adapter conversation parsing.
- Outcome: discoverable command with no fallback and no transcript read when the key is absent.
- Dependencies: Task 2.1.
- Evidence: command tests and built `--help` smoke test.
- [x] Task 2.3: Verify built unavailable behavior in both formats.
- Outcome: status 0 and exact output from compiled CLI.
- Dependencies: Task 2.2 and CLI build.
- Evidence: shell smoke commands with `TYPESAFE_API_KEY` unset.

### Phase 3: Skill, docs, and validation

- [x] Task 3.1: Create `skills/session-compact/SKILL.md` and add it to the built-in manifest/fallback list with validation tests.
- Outcome: installed agents know when and how to use the command and report Jev unavailable.
- Dependencies: stable CLI syntax from Phase 2.
- Evidence: skill tests/lint and manifest assertions.
- [x] Task 3.2: Reconcile implementation and testing docs with actual files, decisions, and evidence.
- Outcome: lifecycle artifacts reflect delivered behavior rather than the initial forecast.
- Dependencies: implementation complete.
- Evidence: feature lint.
- [x] Task 3.3: Run focused tests, CLI package tests/build, workspace build/test as proportionate, formatting/lint, and final lifecycle review.
- Outcome: evidence-backed completion or documented blockers.
- Dependencies: all preceding tasks.
- Evidence: fresh command outputs recorded in testing docs and durable task.

### Phase 4: Measured compaction performance

- [x] Task 4.1: Benchmark a representative Codex session and isolate lookup, parsing, classification, and rendering costs.
- Outcome: full historical enumeration and sequential remote classification identified as the dominant avoidable costs.
- [x] Task 4.2: Add exact-ID lookup to `AgentManager` and every built-in adapter, retaining fallback compatibility for external adapters.
- Outcome: the compact command no longer builds every historical session summary before opening one transcript.
- [x] Task 4.3: Replace sequential classification with an order-preserving bounded worker pool using default concurrency eight.
- Outcome: independent Jev round trips overlap without changing artifact order.
- [x] Task 4.4: Add adapter, manager, command, and concurrency regression tests and rerun repository verification.
- Outcome: direct lookup, provider narrowing, ambiguity, fallback compatibility, concurrency bounds, and ordering are executable contracts.

## Dependencies

```mermaid
flowchart LR
T11[SDK dependency] --> T14[Jev adapter]
T12[Service tests] --> T13[Service implementation]
T13 --> T14
T14 --> T21[CLI tests]
T21 --> T22[CLI wiring]
T22 --> T23[Built smoke]
T22 --> T31[Skill]
T23 --> T32[Docs reconciliation]
T31 --> T32
T32 --> T33[Final verification]
```

- The TypeSafe API is external, but automated tests use mocks and do not require credentials.
- Live-key validation is optional and cannot block CI or local verification.
- No migration or persistent state dependency exists.

## Timeline & Estimates

- Service and adapter: small-to-medium, approximately half a day.
- CLI/test integration: small, approximately two hours.
- Skill/docs/final verification: small, approximately two hours.
- Buffer is reserved for SDK response-shape or existing command-test fixture differences.

## Risks & Mitigation

- SDK/API schema changes: isolate in `JevSessionEventClassifier`, validate labels, pin via lockfile.
- Long sessions can still require many API calls: cap concurrency at eight, preserve ordering, and leave adaptive throttling/batching as a measured follow-up.
- Secret leakage: early key gate, local redaction before calls, sensitive-event exclusion, synthetic security tests.
- Command namespace confusion: reuse existing `agent session` and document discovery via `agent sessions`.
- Command tests are already broad: add focused cases and avoid changing shared behavior.
- Live API unavailable: mock integration boundary; distinguish configuration unavailability from API failure.

## Resources Needed

- Existing `@ai-devkit/agent-manager` session discovery/conversation APIs.
- Official `@typesafe-ai/sdk` package.
- Vitest, Commander test harness, lifecycle lint, and existing built-in skill validation.
- No database, new service, or deployment infrastructure.

## Progress Summary

All planned tasks are complete. Final evidence: CLI package 98 files / 1171 tests passed; feature coverage reached 97.1% statements, 96.77% branches, 90.9% functions, and 98.36% lines; all workspace test/build/lint targets passed; changed TypeScript files are formatted; feature lint and compiled unavailable smoke tests passed. A live paid-key call remains intentionally optional.
Loading
Loading