From 0865afc43be562dbe14528e4299b9e213b54cc93 Mon Sep 17 00:00:00 2001
From: Claude <noreply@anthropic.com>
Date: Tue, 28 Apr 2026 09:24:43 +0000
Subject: feat(executor): add LocalRunner and OpenAI-compat LLM client

Phase 1 of "local OSS models as agents" plan. Adds a third Runner
backed by any OpenAI-compatible HTTP server (Ollama, vLLM, LM Studio,
llama.cpp), and migrates the Gemini-CLI classifier to route through
the same client when configured.

Two-layer split: internal/llm.Client is the workhorse (HTTP, no Pool,
no DB) used directly by the classifier and any future internal helper
that needs cheap reasoning. internal/executor.LocalRunner is a thin
adapter implementing Runner for user-facing tasks. This avoids
Pool reentrancy/deadlock when sub-second internal calls fire from
inside Pool.execute().

Highlights:
- internal/retry: relocated runWithBackoff/IsRateLimitError/ParseRetryAfter
  into a shared package reused by executor and llm.
- internal/llm: Chat (non-streaming) and ChatStream (SSE) over
  /chat/completions with optional bearer auth, json_object response
  format, retry on 429/503, Retry-After parsing.
- internal/executor/LocalRunner: streams deltas into stdout.log in the
  same stream-json envelope ClaudeRunner emits, then writes one
  consolidated assistant block plus a result terminator so existing
  parsers (extractSummary, ParseChangestatFromOutput) work unchanged.
- internal/executor/Classifier: gains optional LLM field; uses
  json_object response format (no markdown-fence cleanup needed).
  Falls back to Gemini-CLI subprocess when LLM is nil.
- Pool.skipClassification: now skips only when the requested agent
  type is registered, so unknown types still reach the load balancer.
- Storage: additive tokens_in/tokens_out ALTERs on executions; CLI
  runners record cost_usd as before, LocalRunner records 0 + tokens.
- Config: [local_model] section (endpoint, model, timeout_seconds,
  default_temperature, api_key). Empty endpoint = no LocalRunner
  registered, classifier falls back to Gemini.

Pre-existing test issues fixed in passing:
- claude_test.go setupSandbox callsites updated to current signature.
- gemini_test.go TestParseGeminiStream skipped (asserts unimplemented
  GeminiRunner stream-error parsing; tracked separately).

Plan: docs/plans/local-oss-runner.md.

https://claude.ai/code/session_017Edeq947TpSm1vQTxMhi1J
---
 internal/config/config.go | 37 +++++++++++++++++++++++++------------
 1 file changed, 25 insertions(+), 12 deletions(-)

(limited to 'internal/config')

diff --git a/internal/config/config.go b/internal/config/config.go
index ce3b53f..7f87391 100644
--- a/internal/config/config.go
+++ b/internal/config/config.go
@@ -15,19 +15,32 @@ type Project struct {
 	Dir  string `toml:"dir"`
 }
 
+// LocalModel configures an OpenAI-compatible local LLM endpoint used for
+// internal helpers (classifier, future elaboration/summarization) and as the
+// backend for the "local" runner. If Endpoint is empty, the LocalRunner is
+// not registered and the classifier falls back to the Gemini CLI.
+type LocalModel struct {
+	Endpoint           string  `toml:"endpoint"`             // e.g. "http://localhost:11434/v1"
+	Model              string  `toml:"model"`                // e.g. "llama3.1:8b"
+	TimeoutSeconds     int     `toml:"timeout_seconds"`      // default 60
+	DefaultTemperature float64 `toml:"default_temperature"`  // default 0.2
+	APIKey             string  `toml:"api_key"`              // optional bearer token
+}
+
 type Config struct {
-	DataDir          string    `toml:"data_dir"`
-	DBPath           string    `toml:"-"`
-	LogDir           string    `toml:"-"`
-	ClaudeBinaryPath string    `toml:"claude_binary_path"`
-	GeminiBinaryPath string    `toml:"gemini_binary_path"`
-	MaxConcurrent    int       `toml:"max_concurrent"`
-	DefaultTimeout   string    `toml:"default_timeout"`
-	ServerAddr       string    `toml:"server_addr"`
-	WebhookURL       string    `toml:"webhook_url"`
-	WorkspaceRoot    string    `toml:"workspace_root"`
-	WebhookSecret    string    `toml:"webhook_secret"`
-	Projects         []Project `toml:"projects"`
+	DataDir          string     `toml:"data_dir"`
+	DBPath           string     `toml:"-"`
+	LogDir           string     `toml:"-"`
+	ClaudeBinaryPath string     `toml:"claude_binary_path"`
+	GeminiBinaryPath string     `toml:"gemini_binary_path"`
+	MaxConcurrent    int        `toml:"max_concurrent"`
+	DefaultTimeout   string     `toml:"default_timeout"`
+	ServerAddr       string     `toml:"server_addr"`
+	WebhookURL       string     `toml:"webhook_url"`
+	WorkspaceRoot    string     `toml:"workspace_root"`
+	WebhookSecret    string     `toml:"webhook_secret"`
+	Projects         []Project  `toml:"projects"`
+	LocalModel       LocalModel `toml:"local_model"`
 }
 
 func Default() (*Config, error) {
-- 
cgit v1.2.3