Zero-egress PR scoring for regulated engineering teams

Many engineering organizations want automated PR scoring and inline diff walkthroughs, but security policies strictly forbid transmitting proprietary source diffs to external model providers.

CodeOtter was architected from day one so that local models are first-class citizens alongside hosted providers. Every local model is defined in models.json and managed directly from Admin → Local Models.

Three specialized local runtimes orchestrated by server.mjs

Depending on whether a local checkpoint is a System One rubric engine or a language/comment generator, CodeOtter boots the matching sidecar process on demand:

  1. System One Sidecar (kind: "s1"): Served by the ggmlc laya runtime speaking POST /v1/systemone. Requests are automatically chunked to maxQuestions (8) and diffs are sized to contextChars for deterministic latency.
  2. Local GGUF Language Model Sidecar (kind: "llm"): Served by llama.cpp's llama-server for full prose summaries, file-by-file walkthroughs, and structured findings.
  3. Microsoft CodeReviewer Sidecar (api: "codereviewer"): Served by the Python venv runtime (runtimes/codereviewer/serve.py) for specialized per-hunk review comments, paired with a local System One model for scoring.
// models.json catalog entry example
{
  "id": "laya-s1-q4_k_m",
  "kind": "s1",
  "maxQuestions": 8,
  "contextChars": 24000,
  "files": ["laya-s1-q4_k_m.gguf"]
}

Automatic GPU-to-CPU runtime fallback and multi-file verification

When provisioning a local sidecar on Linux, macOS, or Windows, CodeOtter tries platform runtime archives in catalog order—attempting the hardware-accelerated GPU build (CUDA / Metal / Vulkan) first and seamlessly falling back to the CPU build if no GPU is available.

Multi-file model checkpoints are verified atomically via modelReady(m) across all required shards before any review job is dispatched.