A self-hosted agent runtime, built around your own API keys.
What it does
CodePuppet lets a user:
- 01Create an account and store their own LLM provider API key (OpenAI, Anthropic, or Google), encrypted per-user on the backend — the backend never uses a shared or global key.
- 02Start an agent session against a chosen provider/model, streamed back token-by-token over Server-Sent Events (SSE).
- 03Let the model call tools mid-conversation. Some tools (reading a file, running a shell command in the developer's own workspace) are executed by the CLI on the user's machine; backend-category tools are executed by the API server itself. Tool results are fed back into the same conversation to continue the turn.
- 04Log the CLI into their account via a device-authorization flow — type a short code, approve it from the browser, no pasting long tokens into a terminal.
The whole system is intentionally “bring your own key”: the server stores and uses only credentials the authenticated user has explicitly saved.
Architecture
CLI — developer's machine
Runs agent sessions locally and talks to the API over HTTPS with a bearer token. Executes file, process, git and user tools in your workspace.
Web app — browser
Sign-in, sign-up, and device approval only. Talks to the API over HTTPS with a cookie session.
API server — Express
Auth (better-auth), controllers (agent-session, credential, catalog, admin), provider registry, tool registry, and the AES-256-GCM credential vault.
Postgres — via Prisma
Stores users, sessions, interactions, turns, messages, and the encrypted provider API keys.
Provider APIs
The provider registry calls out to the real OpenAI, Anthropic, and Google APIs using the decrypted per-user key.
CLI (your machine) ──HTTPS + bearer──┐
├──▶ API server (Express) ──▶ Postgres (Prisma)
Web app (browser) ──HTTPS + cookie───┘ │
└──▶ OpenAI / Anthropic / GoogleTech stack
| Layer | Technology |
|---|---|
| Language | TypeScript everywhere |
| Runtime / package manager | Bun, Node.js ≥ 20 |
| Monorepo tooling | Turborepo, Bun workspaces |
| Backend framework | Express 4 |
| Database | PostgreSQL + Prisma ORM |
| Auth | better-auth — email/password, bearer tokens, admin plugin, OAuth-style device-authorization plugin |
| Validation | Zod v4 |
| LLM providers | OpenAI, Anthropic, Google (Gemini) — via a shared streaming adapter interface |
| Frontend | Next.js (App Router), React, Tailwind, shadcn/ui |
| CLI | Commander, Inquirer, Axios, Chalk |
| Testing | Jest |
| Infra | Docker Compose (Postgres for local dev) |
How an agent session works
- 01Client sends a message plus the chosen provider, model, and credential to the API.
- 02The API validates the model/provider and credential, creates a session and interaction, and streams a request to the LLM provider using the decrypted API key.
- 03As the provider streams back text and tool calls, the API forwards every event to the client immediately over SSE.
- 04Once the stream finishes, everything — the turn, resulting messages, updated session state — is persisted to Postgres in a single transaction.
- 05If the model asked to use a tool, the client (or the server, for backend-only tools) executes it and sends the result back to continue the same interaction. This repeats until the model produces a final answer with no pending tool calls.
- Session — one conversation, pinned to a provider/model, owned by a user.
- Interaction — one logical message exchange within a session; can span multiple turns if the model calls tools.
- Turn — one actual request/response round-trip to the LLM provider (records input/output token counts).
- Message — every user message, assistant message, and tool result, in one strictly-ordered, replayable log per session.
Authentication
Web — sign-up / sign-in
Standard email and password, with a session cookie issued by better-auth.
CLI — device authorization
- CLI requests a device code from the API.
- API returns a short user code and a verification URL.
- CLI prints: “Go to this URL and enter code XXXX-XXXX.”
- Developer opens the URL, signs in if needed, and approves the request.
- CLI polls in the background and receives a bearer access token once approved.
- Every future CLI request carries that token; the CLI stores it locally.
Credential security
Provider API keys are never stored in plaintext and never read from a shared config value. Each user's key is:
- 01Encrypted with AES-256-GCM, using a per-user key derived via HKDF-SHA256 from a single master key.
- 02Bound to the specific user, provider, and label, so a ciphertext can't be decrypted under a different user/provider/label than it was written for.
- 03Decrypted only in-memory, immediately before a provider call is made.
Tools & providers
- Provider registry — one adapter per LLM provider (OpenAI, Anthropic, Google), each speaking a shared internal streaming protocol, so the rest of the system never has to know which provider it's talking to.
- Tool registry — tool definitions with a name, category, and JSON-schema input. Categories (file-read, file-update, process, user, backend) determine where a tool runs: everything except backend runs on the developer's own machine via the CLI (git included — every git operation is just a process shell command, no separate category); backend tools run on the API server itself.
- Modes — every interaction runs in one of four modes, and the mode is what decides which tool categories the model is even offered: Ask (no tools, answer only), Plan (read files, no writes), Code and Auto (full read/write/process/user access).
Project status
Actively under development — Google and OpenAI provider streaming work end-to-end today; Anthropic support is currently a stub.
Latest status on GitHubGet started
npm install -g code-puppetRequires Node.js 20+. Then run code-puppet login.