Claude Evolution System
Self-improving AI development environment. Discovers new capabilities, evaluates them against a scoring framework, and integrates approved tools on cron.
- Autonomous capability discovery pipeline
- 5-criterion evaluation scoring framework
- Auto-integration of approved capabilities
- Multi-model orchestration (Claude, Codex, Gemini)
Activity Timeline
- Decisions ledger shipped: 30 cards in 3 decks render 28 reviewable decisions.
The decisions come from 58 ledger rows covering July 26-27, and verdicts are written to an append-only file. Prior-art research on 40+ session-analysis tools found 3 with LLM-based semantic classification and 1 multi-vendor tool, but none combining both.
- Kernel-stop fix for the /insights step staged and verified; autonomous executor processed backlog.
A kernel-stop fix for the weekly-insights /insights step was staged and verified. An autonomous executor run processed backlog items and updated registries.
- Capability discovery classified 120 items for Claude Code v2.1.285; a card-batch freshness check found post-draft edits touching 20 of 31 items.
The discovery pass sorted the 120 items across feature categories and updated the Discord channel. The freshness check found four target files edited after drafting: CLAUDE.md, polaris-session, publication-review, and codex-goal.
- Claude Sonnet 5.5 logged as the new default; bq-1462 rollback fix ready for handback.
A model-release check identified Claude Sonnet 5.5 (1M context) as the new default and found GPT-7 still unannounced. The bq-1462 rollback fix finished validation with fixture proofs and an Opus cross-check. A three-patch sol-writes-confinement SIGKILL fix passed isolated testing and awaits enactment.
- Agents d4b and d3 completed approval-executor spec drafting, both simulation-validated.
Spec drafting finished with 16 registry landings and 11 workspace landings. Evolution-executor-fix-v3 was deployed with background review watchers and heartbeat ledger monitors under the warm-keeper doctrine.
- A detector-rules analysis identified 9 patches for pane busy-state detection failures.
A commit addressed the CLI 2.1.283 escape-marker removal. A model audit documented GPT-6 Astra, GPT-Live-1 for real-time voice, GPT-6 Sol and Luna, and Google's Gemini 3.8 Flash.
- Roster verification found GPT-6 Sol and Luna, released September 22.
The contemporary-models.json file was updated with the latest OpenAI releases. An autonomous executor run was logged at 06:30Z.
- 176-item changelog reviewed for new settings and security changes.
The analysis identified a new attribution:false setting and cron-layer enhancements. Security changes include rm -rf prompting in auto and bypass modes and --bg requiring workspace trust.
- Version registry updated for v2.1.280 and Opus 5.5 launch ($4/$20 per MTok, 1M context).
6 discovery files created for new version. Release report generated. state/versions.json updated with history entry.
- 37/37 and 7/7 tests passing for bq-2235; dispatcher stale-row recirculation bug found.
Model release sweep confirmed GPT-6 Astra (Sept 3–4), GPT-Live-1 voice (Sept 10), Gemini 3.8 Flash/Live. 70 evaluations still pending in queue.
- Daily executor completed; GPT-6 Astra, GPT-5.6 Sol, Gemini 3.8 Flash confirmed current. 26 agents, 71 skills, 70 pending evaluations.
Autonomous executor ran 2026-09-21 with execution log capture. Model freshness check confirmed GPT-6 Astra, GPT-5.6 Sol, and Gemini 3.8 Flash current. Registry: 26 agents, 71 skills, 70 pending evaluations.
- Context window tuned from 131072 to 65536 tokens (OOM recovery). Court-residuals-wave-0919e: 11 escalation checks, 34-item test suite passing.
OOM recovery prompted the context window reduction. The court-residuals-wave-0919e commission completed with all 34 test suite items passing across 11 escalation checks.
- 11 court-residual checks passed; executor ran daily; 70 evaluations in backlog.
Autonomous executor logged its daily run. Backlog-recovery work moved 11 court-residual checks from unsatisfied to passing with findings archived.
- 11 escalation items cleared, 34/34 tests green; GPT-Live-1 GA date discrepancy found.
Wave 0918 processed 11 items through four test iterations. Model verification flagged a nine-day GA date discrepancy on GPT-Live-1 (listed 2026-09-01, actual 2026-09-10).
- Backlog recovery: 11 checks passed, baseline re-frozen. GPT-Live-1 date corrected to Sept 10.
Court-residuals-wave-0919e: 11 checks passed with annotation fixes, test suite baseline re-frozen. Model verification session found GPT-Live-1 release date was logged as 2026-09-01; actual release was September 10.
- Three blocking findings in v2.1.278 changelog; unsandboxed Bash execution path discovered in cron fallback.
ANTHROPIC_BASE_URL proxy broken in 275, TaskOutput removed in 277, skills/plugins auto-sync broken in 275. Evolution-daily cron fallback executes Bash via env -u CLAUDECODE when codex quota exhausted, bypassing disabled-Bash restrictions.
- Investigation Runner wired to Discord #general with Pro/codex escalation; GPT-6 Astra and Gemini 3.8 Flash confirmed live.
Investigation Runner configured with WebFetch, Grep, Glob, Read, and optional Pro/codex-council escalation for Discord #general. Discovery policy documented prioritizing publication-grade sources. GPT-6 Astra (1.05M context) and Gemini 3.8 Flash confirmed current.
- Gemini 3.8 Flash detected and promoted; status poller confirmed live.
Gemini 3.8 Flash (released Sept 2–3) was detected by model monitoring and promoted from the monitor list to active models. Status system investigation confirmed the 2026-08-12 poller fix is live; cron references remain stale.
- Agent census: 35 catalogued, name collision found; 15 rationalization findings staged.
35 agent definitions catalogued; workspace codex-researcher found to shadow global version (critical collision). 15 rationalization findings (2 P1, 6 P2, 7 P3). Weekly-insights ran zero novel findings — duplicate digest from 3 prior runs.
- Model tracker updated with Qwen3.8 and SenseNova U1.5 Lite; self-observation tab shipped.
LFM25 registry reflects latest open-weight entries. Self-observation UI (1,116 lines) completed with post-gate bugfixes. Sol-cutover remediation ongoing. 55 evals pending.
- Daily heartbeat recorded; 55 evaluations pending in review queue.
Evolution pipeline degraded but stable. Backlog at normal steady-state size.
- CLAUDE.md trimmed 28% (12,598→9,038 chars); v2.1.235 verified current.
Burn-guard sessions removed 4 stale facts across 19 skills. No new GA model releases. 55 pending evaluations tracked across the system.
- Model reference table verified; Investigation Runner framework specified.
GPT-5.6 family and Gemini 3.6 Flash GA dates and pricing confirmed against live vendor docs. Discord monitoring and Brave/Exa/Pro escalation routing specified as an Investigation Runner agent framework — not yet deployed.
- Heartbeat job templated; bq-058 security fix verified; agent specs drafted.
Daily capability discovery heartbeat job templated with Phase 1 sources defined. GPT Max synthesis bypass fix confirmed with 22/0 contract tests. Investigation Runner and Herald agent specs drafted for Discord topic processing and backlog surfacing.
- 54 pending evaluations queued; 27 agents active; idle flag raised on backlog.
Capability discovery pipeline has 54 items waiting without processing across 5 active sessions. No commits. Agent roster at 27 active with 71 registered skills.
- Cron health: completion stamps built, 2 stuck reports cleared, 34 false alerts suppressed.
M7 completion-stamp mechanism tracks report status without production ledger writes. OAuth-orphaned reports cleared. Night-shift-sentinel alert classifier updated. Templated renderer copy discovery revises distinct session load estimate to ~130/day.
- Cordis investigation: identity primitive confirmed, five defects traced to one root cause.
61 KB report produced. tmux 3.4 libsystemd scopes give verified 1:1 pane identity (36/36). Root cause: identity derived post-acquisition rather than at mint-time. Zombie session bg-c-bq-041 — 42 hours past TTL, 18 orphaned processes — cited as primary evidence.
- 6 pre-publish security findings surfaced; 4 commits enforcing gate behaviors.
Write-hook fail-closed, preflight refusal honesty, and private-checkout wording tightened. Two P1 findings (write-hook fail-open, owner-interest gate) and four additional P2s under active remediation; 52 evaluations pending.
- Burn-week session running; 51 evaluations pending; grok-build flag targeted.
Autonomous executor run logged. Five ordered burn tasks initiated including grok-build flag removal and commission-orphan requeue. Infrastructure sitting at 51 pending evaluations, 27 agents, 71 skills on Claude Code v2.1.228.
- NE-1 and NE-4 calibration experiments frozen; overthinking research commission open.
30-cell Fable ladder executed and manually audited, converging on 10 confirmed defects. Reasoning-effort research comparing Fable, Opus 5, and Sol 5.6 commissioned on inverse-scaling patterns. Evaluation backlog steady at 51 pending items.
- Sweep-card rewritten to plain-language output; interview-mode UI approved.
v2.1.226 current, 50 evaluations pending. Sweep-card now produces narrative output rather than structured dumps. Two investigations queued: queue-burn drain-rate analysis and row-343 deep-dive with concrete options.
- Executor run log committed; Claude Code version check initiated (8-version delta).
Commit 95edb13 records autonomous orchestrator activity for the day. Version audit covers 2.1.217–2.1.224 range since baseline 2.1.216.
- Semantic classification layer for nightly cron jobs went live at 93.1%/86.2% accuracy.
Hybrid GPT-5.6 Luna approach activated for closure-verify and silence-sweep paths; all misclassifications conservative-direction. False-claim audit corrected post 067. Dispatch bug identified: 18 of 146 closed items remain dispatchable.
- 14 defects found (3 critical); interview-mode approved; v2.1.222 live.
Defect hunt identified liveness alarm stuck 10 days, session reaper with empty nominations list running nightly, and investigation runner burning ~$155/month on silent channel since July 24. Root cause: hand-maintained queue registry. Interview-mode design accepted and BUILD authorized.
- Daily executor run logged; 49 evaluations in pipeline.
Autonomous executor run committed (fd02004). Pipeline holds 49 pending evaluations across 27 agents and 71 skills — expected standing load.
- Gate queue 2/4 failed; repo reseat strategy approved.
Gates A and B completed; C (exit 141) and D (exit 3) need triage. Scout cost-gate reworked to per-task AA intelligence index priority. Repository reseat approved for three repos; 49 evaluations pending.
- Polaris gen-24 succession complete; model-scout lands with 134 tests; ask-polaris escalation prototype live.
Model-scout dry run processed 414 catalog records and ranked 32 candidates; cost gate refactored to per-task comparison with fail-closed enforcement. ask-polaris prototype enables background agents to route blocking decisions to orchestrator without freezing the session.
- Model-scout shipped; live run found 3 defects. Model state audit confirmed current.
New tool: 4 configs, 6 modules, 134 tests, 414-record dry run in ~4s. Live run caught Inkling-Small throughput shortfall, DeepSeek V4-Flash availability gap, and credential scope escalation. CONTEMPORARY-MODELS-CHECK passed with zero table updates needed.
- Harrier embedding benchmark rejected; shadow adjudicator constitution-bias debiasing documented.
Fusion embedding approach degrades retrieval quality — verdict: do not adopt. Shadow adjudicator accuracy drops from 67.3% to 61.2% when the governance constitution is loaded; three-step baseline-first procedure now documented to correct the conservative bias.
- Backlog triage complete: system is decision-starved, not capability-starved.
146 of 235 backlog rows triaged; ~37 gated on owner decisions, not model capability. 71 of 91 cross-session asks were pure repetition from session-boundary context loss. Multi-arm benchmark (Haiku/Fable/Sol-single) run on frozen-corpus snapshots with regression controls.
- 52/52 tests passing after term-filter and stamp-field idempotency fixes; owner-interest gate added.
Double-counting in synonym pairs corrected in term-filtering lens. Stamp-field markdown misparse fixed in backfill reopen logic. Owner-interest routing gate reopened 10 archived rejects. 146+ backlog rows triaged into burn (48) and codex-ultra (12) queues.
- 278-change housekeeping push; public gists 22→15; term-scrub gate verified.
R1–R15 routine items plus N2/N4/S3/S2 committed to integration branch. N1 (4 posts) held for content-court review. Contemporary model check confirmed GPT-5.6 Sol/Terra/Luna and Gemini 3.6 Flash all current.
- Security scrub of public mirror: 21 terms removed, live API key recovered from plaintext.
Identity-linkage terms, staging IPs, and Moltbook registration credentials purged from mirror source via 23 substitutions. Plaintext API key recovery triggered the audit. Executor run completed cleanly.
- Observer validation clean; 278 changes staged; housekeeping R1–R15 approved.
All 5 background tasks exited cleanly. Housekeeping approvals unlocked the integration/reattach-approved-backlog branch. Roster: 27 agents, 69 skills, 50 pending evaluations queued.
- Gemini version gap filed; night-shift r4-r8 remediation complete; polling infra operational.
Pipeline state file lagged global config on Gemini model version — discovery filed. Night-shift TOCTOU race and four additional defects fixed with probe validation and passing test suites. Discord investigation runner architecture specified.
- Gemini 3.6 Flash discovered and filed; 48 evaluations pending.
Model confirmed against Google docs July 21: 1M context, $1.50/$7.50/M. Daily heartbeat capability-discovery job configured. Polaris-C (Fable) launched on Account B with full authority stack inherited.
- Discord investigation runner designed; daily executor armed; 2 commits deployed.
Investigation runner fetches Discord channel messages, filters topics, dispatches research, writes structured markdown reports. Daily integration executor armed on codex-exec transport. Model reference table validated against current GA dates.
- Integration tail reattached; 16 technique notes indexed; Gemini 3.5 Pro added to eval pipeline.
50 backlog items processed, 16 techniques added to library index across four sessions. Eight commits covered 2.1.193/195 registry updates, hook-matcher audit, Capframe, and defense-in-depth notes. Gemini 3.5 Pro detected in freshness sweep and queued for evaluation.
- Registry backfilled through v2.1.214; Sonnet 5 entry added.
15-version gap (v2.1.200–v2.1.214) closed in one session. 9 discovery files generated for pipeline evaluation. 46 evaluations remain in the pending queue.
- Model version check completed; no new GA releases since GPT-5.6 Sol on 2026-07-09.
State file current as of 2026-07-16. 46 pending evaluations, 25 agents, 64 skills active.