ra-yavuz › lillycoder

lillycoder logo: a terminal window with ASCII glyphs and the silhouette of a girl with long hair seen from the back

lillycoder

A local-first coding assistant for projects, analysis, code, and long conversations. It discovers local model endpoints, can start downloaded Hydra models, and gives the model permission-gated project tools. No cloud, no API key, no telemetry, no account.

What sets it apart: a real persona system. Six bundled voices (default kid coder, tsundere, yandere, sweet, calm-adult, analytical), live switching, copy-on-write user shadows, and a model-driven evolve mode where Lilly rewrites her own system prompt over time and the shape persists across sessions.

What it is

Run lillycoder in a project directory and start typing. The model receives a focused set of tools appropriate to the request (read_file, write_file, edit_file, bash, mkdir, mv, rm, grep, find, list_dir, pkg_install) to do what you asked.

Every mutating action prompts:

🦊 lilly wants to: write_file("src/index.js", 142 chars)
   [y]es  [n]o  [a]lways for this tool  [p]ath: always for this exact target
   >

What it is not

lillycoder does not ship or download a model and does not manage Docker itself. It works with any OpenAI-compatible endpoint. When the optional hydra-llm command is installed, LillyCoder can list Hydra's downloaded models and delegate starting the one you select.

Install

One line (Debian / Ubuntu)

Sets up the signed ra-yavuz apt repo if not already added, refreshes the package index, and installs lillycoder. Idempotent, safe to re-run:

sudo bash -c 'set -e; install -m 0755 -d /etc/apt/keyrings && curl -fsSL https://ra-yavuz.github.io/apt/pubkey.gpg -o /etc/apt/keyrings/ra-yavuz.gpg && echo "deb [signed-by=/etc/apt/keyrings/ra-yavuz.gpg] https://ra-yavuz.github.io/apt stable main" > /etc/apt/sources.list.d/ra-yavuz.list && apt update && apt install -y lillycoder'

If you already added the ra-yavuz apt repo earlier, all you need is sudo apt update && sudo apt install lillycoder. The sudo apt update step is required: without it apt will not see new packages or new versions.

One line via the bundled installer script

Equivalent to the above, with extra prerequisite checks and a friendlier output summary:

curl -fsSL https://raw.githubusercontent.com/ra-yavuz/lillycoder/main/scripts/get.sh | sudo bash

If you would rather read the script first (recommended for any curl | bash):

curl -fsSL https://raw.githubusercontent.com/ra-yavuz/lillycoder/main/scripts/get.sh -o get.sh
less get.sh
sudo bash get.sh

Step by step (manual repo setup)

# 1. Trust the signing key
sudo install -d -m 0755 /etc/apt/keyrings
curl -fsSL https://ra-yavuz.github.io/apt/pubkey.gpg \
  | sudo tee /etc/apt/keyrings/ra-yavuz.gpg >/dev/null

# 2. Add the apt source
echo "deb [signed-by=/etc/apt/keyrings/ra-yavuz.gpg] https://ra-yavuz.github.io/apt stable main" \
  | sudo tee /etc/apt/sources.list.d/ra-yavuz.list

# 3. Refresh the package index, then install
sudo apt update
sudo apt install lillycoder

From source (any Linux, also macOS via pip)

git clone https://github.com/ra-yavuz/lillycoder.git
cd lillycoder
pip install --user -e .

Platform support

Tested on Ubuntu (Linux only). Should also work on WSL2 Ubuntu / Debian (it is a Linux distro, the apt path applies). On macOS, the .deb and apt install paths do not apply, but the from-source pip install --user -e . path is expected to work because the dependencies (httpx, prompt_toolkit, rich, pydantic) are all cross-platform and lillycoder shells out to standard POSIX tools that exist on Darwin. macOS support is not regularly tested by the author, so if you hit a portability issue please open an issue.

Quick start

In any project, run:

cd ~/myproject
lillycoder
🦊 scanning localhost for LLM servers...
🦊 found http://localhost:18102/v1 (hydra:my-coder, 1 models)
✓ hydra:my-coder · qwen-coder.gguf
🦊 lilly is awake in /home/you/myproject · /help · /new · /sessions · /exit
› what files are in this folder?

If no server is running and Hydra is installed, LillyCoder lists the models already downloaded by Hydra. Select one by number or alias and it is started for you. You can still use lillycoder --api http://your.host:8080/v1 for any other compatible server.

Long conversations

The complete conversation stays in the project session archive while LillyCoder keeps a bounded working context for the model. It checks the prompt and tool schemas before every completion, keeps recent turns, folds older turns into durable memory, and replaces used tool payloads with short records. Compaction snapshots persist across restarts, and a local fallback works even if the model cannot summarize.

Use /context to inspect usage, /compact to checkpoint manually, /new to start a fresh conversation, and /sessions plus /resume to return to one. The live working window is capped conservatively at 32K tokens even if a server advertises a much larger trained context.

Personalities

lillycoder ships six bundled personas, all written in first person with explicit anti-roleplay rules so a local model still sounds like Lilly typing rather than narrating about her:

namevoice
defaultnine-and-a-half-year-old kid coder, warm and curious
tsunderesnippy, grumpy, still does the work
yanderedoting, focused on the user, mildly possessive about the code
sweetgentle, encouraging, low-key cheerful
adultcalm senior engineer voice, no exclamation marks
analyticalprecise, methodical, distinguishes "checked" from "assumed"

Switch live with /personalities load <name>. Add your own with /personalities add <name> <text> or drop a markdown file into ~/.config/lillycoder/personas/; user files shadow bundled ones of the same name.

When you shadow a bundled persona, lillycoder snapshots the bundled text at the moment of override (a .bundled-base.md sidecar). Later you can run /personalities diff <name> to see your edits AND any upstream drift since you forked. Bundled files are never overwritten by an update if you have a shadow.

Lilly can manage her own personalities through real tool calls (add_persona, clone_persona, set_active_persona, set_evolve). Tell her "make a pirate persona and switch to it" and she does it through the tool registry, not by writing files in your repo.

Flip /persona-evolve on to snapshot the current in-memory persona to disk and switch to it. From then on, every persona rewrite (whether by the model itself via set_persona, or by you inline) gets persisted to that file. Next launch, lillycoder reloads the last active persona automatically.

Token budget

The default /max-tokens auto computes a per-reply cap from your model's reported context window (about 85% of remaining headroom, with a 4096-token ceiling). That matters because:

Set an explicit cap any time: /max-tokens 256 for snappy answers, /max-tokens 4096 for long-form. Or via CLI: lillycoder --max-tokens 4096.

Compatible servers

ServerDefault portNotes
hydra-llm18080+recommended pairing (sibling project)
llama.cpp llama-server8080OpenAI /v1 shape native
ollama11434OpenAI surface at /v1
LM Studio1234built-in local server
any otherany--api http://your.url/v1

The model on the other end matters. Tool-calling reliability needs a model trained for it. lillycoder warns when the chosen model is not in its known-tool-capable allowlist (Qwen 2.5+, Qwen 3, Gemma 3+, Llama 3.1+, Mistral Small 3, Dolphin 3 R1). Pass --force to silence.

Pairs with hydra-llm

hydra-llm is a sibling project that manages local LLM servers: it wraps llama.cpp in Docker, ships a curated GGUF catalog with anonymous downloads, and exposes each running model as an OpenAI-compatible endpoint on a stable local port. LillyCoder asks Hydra for all running ports and can start an already-downloaded model when none is running:

cd ~/myproject
lillycoder

Hydra handles downloads, lifecycle, safe llama.cpp resource settings, and its optional KDE Plasma widget. LillyCoder is the coding assistant on top: project tools, permissions, context, conversations, and personas. It remains compatible with other OpenAI-compatible servers.

Safety

Hard-banned commands cannot be turned off by --bypass-permissions: sudo, rm -rf /, rm -rf ~, mkfs, dd of=/dev/*, recursive chmod / chown of / or ~, fork bombs. They are refused at the safety classifier before exec. Writes outside the working directory are also blocked by default; widen with LILLY_ALLOW_OUTSIDE_CWD=1 if you really mean it.

Disclaimer: lillycoder is provided AS IS, WITHOUT WARRANTY OF ANY KIND. The LLM can read, write, and delete files in the current working directory and run shell commands on your behalf. LLMs hallucinate; they may invent file paths or write incorrect commands. You alone are responsible for any damage to your data, hardware, or system. By installing or running this software you accept full risk. See README for the full text.