Details

A self-hosted agent runtime, built around your own API keys.

What it does

CodePuppet lets a user:

  1. 01Create an account and store their own LLM provider API key (OpenAI, Anthropic, or Google), encrypted per-user on the backend — the backend never uses a shared or global key.
  2. 02Start an agent session against a chosen provider/model, streamed back token-by-token over Server-Sent Events (SSE).
  3. 03Let the model call tools mid-conversation. Some tools (reading a file, running a shell command in the developer's own workspace) are executed by the CLI on the user's machine; backend-category tools are executed by the API server itself. Tool results are fed back into the same conversation to continue the turn.
  4. 04Log the CLI into their account via a device-authorization flow — type a short code, approve it from the browser, no pasting long tokens into a terminal.

The whole system is intentionally “bring your own key”: the server stores and uses only credentials the authenticated user has explicitly saved.

Architecture

CLI — developer's machine

Runs agent sessions locally and talks to the API over HTTPS with a bearer token. Executes file, process, git and user tools in your workspace.

Web app — browser

Sign-in, sign-up, and device approval only. Talks to the API over HTTPS with a cookie session.

API server — Express

Auth (better-auth), controllers (agent-session, credential, catalog, admin), provider registry, tool registry, and the AES-256-GCM credential vault.

Postgres — via Prisma

Stores users, sessions, interactions, turns, messages, and the encrypted provider API keys.

Provider APIs

The provider registry calls out to the real OpenAI, Anthropic, and Google APIs using the decrypted per-user key.

CLI (your machine) ──HTTPS + bearer──┐
                                     ├──▶ API server (Express) ──▶ Postgres (Prisma)
Web app (browser) ──HTTPS + cookie───┘            │
                                                  └──▶ OpenAI / Anthropic / Google

Tech stack

LayerTechnology
LanguageTypeScript everywhere
Runtime / package managerBun, Node.js ≥ 20
Monorepo toolingTurborepo, Bun workspaces
Backend frameworkExpress 4
DatabasePostgreSQL + Prisma ORM
Authbetter-auth — email/password, bearer tokens, admin plugin, OAuth-style device-authorization plugin
ValidationZod v4
LLM providersOpenAI, Anthropic, Google (Gemini) — via a shared streaming adapter interface
FrontendNext.js (App Router), React, Tailwind, shadcn/ui
CLICommander, Inquirer, Axios, Chalk
TestingJest
InfraDocker Compose (Postgres for local dev)

How an agent session works

  1. 01Client sends a message plus the chosen provider, model, and credential to the API.
  2. 02The API validates the model/provider and credential, creates a session and interaction, and streams a request to the LLM provider using the decrypted API key.
  3. 03As the provider streams back text and tool calls, the API forwards every event to the client immediately over SSE.
  4. 04Once the stream finishes, everything — the turn, resulting messages, updated session state — is persisted to Postgres in a single transaction.
  5. 05If the model asked to use a tool, the client (or the server, for backend-only tools) executes it and sends the result back to continue the same interaction. This repeats until the model produces a final answer with no pending tool calls.
  • Session — one conversation, pinned to a provider/model, owned by a user.
  • Interaction — one logical message exchange within a session; can span multiple turns if the model calls tools.
  • Turn — one actual request/response round-trip to the LLM provider (records input/output token counts).
  • Message — every user message, assistant message, and tool result, in one strictly-ordered, replayable log per session.

Authentication

Web — sign-up / sign-in

Standard email and password, with a session cookie issued by better-auth.

CLI — device authorization

  1. CLI requests a device code from the API.
  2. API returns a short user code and a verification URL.
  3. CLI prints: “Go to this URL and enter code XXXX-XXXX.”
  4. Developer opens the URL, signs in if needed, and approves the request.
  5. CLI polls in the background and receives a bearer access token once approved.
  6. Every future CLI request carries that token; the CLI stores it locally.

Credential security

Provider API keys are never stored in plaintext and never read from a shared config value. Each user's key is:

  1. 01Encrypted with AES-256-GCM, using a per-user key derived via HKDF-SHA256 from a single master key.
  2. 02Bound to the specific user, provider, and label, so a ciphertext can't be decrypted under a different user/provider/label than it was written for.
  3. 03Decrypted only in-memory, immediately before a provider call is made.

Tools & providers

  • Provider registry — one adapter per LLM provider (OpenAI, Anthropic, Google), each speaking a shared internal streaming protocol, so the rest of the system never has to know which provider it's talking to.
  • Tool registry — tool definitions with a name, category, and JSON-schema input. Categories (file-read, file-update, process, user, backend) determine where a tool runs: everything except backend runs on the developer's own machine via the CLI (git included — every git operation is just a process shell command, no separate category); backend tools run on the API server itself.
  • Modes — every interaction runs in one of four modes, and the mode is what decides which tool categories the model is even offered: Ask (no tools, answer only), Plan (read files, no writes), Code and Auto (full read/write/process/user access).

Project status

Status

Actively under development — Google and OpenAI provider streaming work end-to-end today; Anthropic support is currently a stub.

Latest status on GitHub

Get started

$npm install -g code-puppet

Requires Node.js 20+. Then run code-puppet login.