No description
  • TypeScript 93.7%
  • JavaScript 6.3%
Find a file
2026-09-23 12:58:11 +02:00
docs feat: recover durable LocalAI job outcomes 2026-09-23 12:58:11 +02:00
examples feat: integrate LocalAI embedded agents with OpenCode 2026-09-18 09:07:12 +02:00
patches feat: use verified OpenCode command interception 2026-09-18 11:02:28 +02:00
scripts feat: use verified OpenCode command interception 2026-09-18 11:02:28 +02:00
src feat: recover durable LocalAI job outcomes 2026-09-23 12:58:11 +02:00
test feat: recover durable LocalAI job outcomes 2026-09-23 12:58:11 +02:00
.gitignore feat: integrate LocalAI embedded agents with OpenCode 2026-09-18 09:07:12 +02:00
package-lock.json feat: integrate LocalAI embedded agents with OpenCode 2026-09-18 09:07:12 +02:00
package.json feat: integrate LocalAI embedded agents with OpenCode 2026-09-18 09:07:12 +02:00
README.md feat: recover durable LocalAI job outcomes 2026-09-23 12:58:11 +02:00
tsconfig.json feat: integrate LocalAI embedded agents with OpenCode 2026-09-18 09:07:12 +02:00

opencode-plugin-localai

Independent integration for OpenCode 1.18.31 and LocalAI's embedded LocalAGI pool, with durable-job recovery targeting LocalAI commit 6d4797d7e95c98fb590ff8476ed67ceaa160c89e. OpenCode owns the conversation and overall plan. Explicit human approval starts sequential worker tasks; worker questions, plans, progress and results return to the originating OpenCode session while normal conversation continues.

Recommended host: FiloSpaTeam/opencode command-interception, tested at 52169044c1c821a477d4386b207da7480f05c68c. On that host, successful /localai commands return HTTP 200 and one host-owned receipt without an error banner. Capability detection selects the hook; both patched and stock hosts use package version 1.18.31. Genuine errors remain errors. Native question dialogs are still outside this change. See adapter contract and validation.

On unpatched hosts, /localai and localai_propose are disabled by default; discovery and status remain available. Remove manually defined localai commands from markdown/config when using this disabled mode. Explicit LOCALAI_LEGACY_COMMANDS=1 restores the old path for idle-session compatibility testing: successful commands still show the sentinel error banner. Read the session receipt and do not retry because of that banner. An idle-status check rejects busy/unknown state, but cannot atomically reserve an idle session against another concurrent caller; prefer the patched host.

Install

npm ci
npm run check

Load the built module through a local OpenCode plugin file, for example ~/.config/opencode/plugins/localai.js:

export { default } from "/absolute/path/to/opencode-plugin-localai/dist/index.js"

Set these in the environment that starts OpenCode. Supply the real token through your secret manager or a protected environment file, not chat:

export LOCALAI_URL="https://localai.example.invalid"
export LOCALAI_AGENTS="coding-worker"
# LOCALAI_TOKEN must be supplied securely.
# Optional LOCALAI_STATE_DIR: private directory (0700), shared across projects.

HTTP is allowed for loopback development. Other plaintext endpoints require LOCALAI_ALLOW_HTTP=1. Credentials and tenant IDs cannot be selected by tool arguments. Only discovery aliases both visible to the authenticated user and locally allowlisted are exposed. The plugin never retrieves agent configuration secrets.

Use one OpenCode pool owner, one state directory and one credential identity for this integration. The private pool-scoped state preserves active reservations across projects and normal shutdown; directory instances within the same OpenCode process share one coordinator; a second process fails closed while the pool lock is held. A crash can leave a .lock file: verify its recorded PID is no longer running before removing only that lock and restarting. Keep the JSON state. Credential rotation uses a different namespace; first externally settle existing work. Cross-machine/global tenant leases require the backend addition below.

Conversation workflow

  1. Ask OpenCode to discover workers and propose a plan with localai_discover and localai_propose. Each task contains id, allowlisted agent, message, and dependsOn (earlier task IDs). Review the exact tasks returned. Proposing does not dispatch.

  2. Type the approval command yourself, substituting the proposed plan ID:

    /localai {"action":"approve","id":"PLAN_ID"}
    

    To edit the overall plan, supply a complete tasks array in that same approval command. The submitted array becomes the approved version. Later tasks start only after their dependencies complete with a captured final result. Failure stops the plan.

  3. Continue chatting with OpenCode. Worker notices are session-scoped. Command receipts on the patched host are excluded from model history and do not start or resume inference. Worker progress/results still use separate noReply messages: those do not directly start a model request, but may steer an already-running loop. This step does not make background notification delivery passive. For a bounded foreground view, ask OpenCode to call localai_status with waitSeconds up to 30. It does not cancel background work when the wait ends.

  4. Respond to worker interactions directly:

    /localai {"action":"answer","id":"QUESTION_ID","selected":["JSON"],"text":"Include a summary"}
    /localai {"action":"plan","id":"WORKER_PLAN_ID","decision":"approve","subtasks":["Implement dry run","Test no writes"]}
    /localai {"action":"plan","id":"WORKER_PLAN_ID","decision":"replan","feedback":"Limit changes to the CLI"}
    /localai {"action":"plan","id":"WORKER_PLAN_ID","decision":"reject"}
    /localai {"action":"status"}
    

    Options are exact backend labels. Text must be allowed by the question. Omitting subtasks preserves the proposed list; [] sends an explicitly empty edited list, which the backend may reject. Final rejection always sends empty feedback. Child questions and plans require human decisions just like root interactions. Commands are registered with a discard placeholder so OpenCode never expands raw JSON as shell or attachment markup; do not replace the command template with $ARGUMENTS.

The model has no answer/approve tool. Commands are intended to be submitted by the human; ordinary chat does not silently approve or answer worker questions. The host command endpoint authenticates server access, not human provenance: an authorized script or privileged model shell can call it. Patched command receipts confirm handling, not durable backend completion or user authorship. This assumes a trusted local OpenCode runtime: arbitrary code running as your OS user can access its credentials and state, so this is not a sandbox against malicious local shell execution.

Reconnect and uncertainty

The client connects authenticated SSE and checks the pending endpoint before submission. Reconnect and pending refresh retry with bounded backoff. Questions are recoverable; pending contains only the oldest plan, so sibling plans remain until explicitly resolved. Duplicate events, wrong sessions and child completion do not advance root tasks.

Approval is stored before dispatch. If preflight fails, the command errors and leaves the task queued; repair connectivity, then use resume. If submission fails ambiguously, the command errors and preserves an uncertain reservation. Do not repeat approve or resubmit the task to work around an error.

The root receipt’s job_id is saved against the originating OpenCode session/task. Reconnect and restart query authenticated GET /api/agents/:name/jobs/:job_id independently of SSE/pending availability, including when the agent has been removed. Terminal SSE triggers that same lookup. Job, alias, root message and conversation must match before accepting a record. Answers and live-loop injection receipts do not create durable jobs.

Lookup outcome Plugin behavior
accepted, running, waiting_user, waiting_agents Keep the reservation and refresh pending interactions. These are last persisted states, not liveness heartbeats.
completed Save the result (including an empty result), queue one job-keyed report, then advance approved dependencies.
failed Report the safe error and stop remaining tasks. Side effects may have occurred.
interrupted Stop the plan, retain the reservation and ask the user to inspect the task/repository. No execution resume.
401/403 Outcome unknown; fix access and explicitly check again.
404/410/501 Unknown/purged record, expired result, or unsupported executor. Retain uncertainty and reservation; pause automatic lookup. A 404 does not prove non-execution.
503/500/network/invalid response Outcome unknown; retry only GET with exponential backoff capped at 30 seconds.

Empty pending does not mean completed. A lost submission receipt without a saved job ID still cannot be safely recovered or retried. Older records without job IDs retain the existing requirement for both root final message and completed event. persistence_unavailable means unknown persistence outcome, never confirmed execution failure. No outcome or lookup error automatically resubmits a task.

Logical outcome delivery is deduplicated by job ID, with terminal state, marker and notice committed together. Notice transport remains at least once: a crash after OpenCode accepts a notice but before local acknowledgement is saved can repeat that same notice ID. Durable lookup has retention limits (default 30 days, then an expired tombstone for another 30 days); archive important results. Questions/plans remain in a volatile registry and may disappear on backend restart; durable job recovery does not restore them.

/localai {"action":"resume"}

Retries a pre-submission connection or forces a fresh lookup/reconnect for existing work, including a lookup blocked by an earlier error; it never retries an uncertain task POST. The status action reads local state without contacting LocalAI. After you externally verify that an uncertain old worker has stopped, you may release the reservation:

/localai {"action":"release","confirmed":true}

Release stops the remaining plan locally. It does not cancel the backend worker. There is no verified per-job cancel endpoint. Local state is sensitive task data, retained until manually archived/deleted after all workers are settled. Delivery to a deleted OpenCode session needs operator reconciliation; the outbox is retained rather than silently discarded.

Model serving is separate

examples/opencode.json configures an OpenAI-compatible LocalAI model provider using placeholders and LOCALAI_INFERENCE_API_KEY. Delegation uses LOCALAI_URL, LOCALAI_TOKEN and /api/agents, never the model provider protocol. Substitute the actual installed DeepSeek V4 Flash model identifier. No GB10 image, quantization, memory budget, context size or throughput is assumed or provisioned. LocalAI scheduling must arbitrate continued OpenCode inference and worker inference; one worker does not guarantee only one global inference request.

Configure repository/MCP access, user questions, planning/plan approval and embedded child agents on the worker. OpenCode's local skills/files do not automatically exist in LocalAGI. Remote /v1/responses workers and the separate LocalAGI container are outside v1.

Verification

npm run check
# After npm run build, with an independently installed pinned host:
# Patched binary:
OPENCODE_BINARY=/path/to/patched-opencode node scripts/opencode-smoke.mjs
# Or run the pinned fork from a prepared source checkout:
OPENCODE_SOURCE=/path/to/opencode BUN_BINARY=/path/to/bun node scripts/opencode-smoke.mjs
# Stock-host safety and opt-in compatibility:
LOCALAI_SMOKE_MODE=disabled OPENCODE_BINARY=/path/to/stock-opencode node scripts/opencode-smoke.mjs
LOCALAI_SMOKE_MODE=legacy OPENCODE_BINARY=/path/to/stock-opencode node scripts/opencode-smoke.mjs
npm pack --dry-run

Tests include a real HTTP/SSE server with pinned-source-derived payloads, early events before receipt, approval edits, isolation, duplicate/out-of-order completion, lost submission, reconnect, stale interactions and private persistence. The host smoke test checks tool registration, receipt persistence, command consumption, real error propagation, multi-directory isolation, literal shell syntax handling and zero model requests. On the patched path it asserts HTTP 200, matching returned/persisted receipt ID and the commandReceipt marker. It does not invoke a real LocalAI worker.

The live DeepSeek/GB10 small-repository coding scenario requires a configured endpoint, secure credential, embedded worker alias and disposable repository accessible to its tools. It has not been run. See acceptance checklist, backend additions, and implementation plan.