A text adventure where the world is generated turn-by-turn as you explore. Nothing is placed; everything is improvised.
There's a chest. It's locked.
The objective says: open the chest. But the game doesn't tell you how, because it doesn't know yet.
There's no designated key waiting in a designated drawer. The lock might be jammed, in which case the key you find won't help. So you wander.
A few rooms later, you find an axe. The axe wasn't placed there to solve the chest. Nothing is placed to solve anything. Nothing was placed at all until you walked into the room.
But you're holding an axe, and there's a locked chest. When you swing at the lid, the world agrees that's what happens.
Rooms, objects, and complications are generated turn by turn as you explore. Some of what appears will turn out to matter. You won't know which things matter until you try something.
Replay the same preset and the axe might not exist. Maybe there's a crowbar. Maybe the chest is rusted shut and now you're looking for water.
Maybe the chest contains something you really wish it hadn't.
The seed scenario is the same; what fills it in isn't.
You can also skip the presets entirely and start in an empty open world. Type what you do, and the world assembles itself around you.
Note: Some of the videos are older and have different interfaces, but the concepts are still the same.
Demo: Promotional Video
Demo: Gameplay video #1 — Full game of Merlin's Daughter (20m)
Demo: Gameplay video #2 — streaming TTS narration and per-turn image generation in action. Slightly old — new UI and logic now, but still a good example.
Demo: Editor video #1 — Create your own scenarios!
- Quick Start
- How to Play
- Optional: Narration
- Optional: Per-Turn Images
- Configuration
- Scenario Editor
- My Personal Configuration
- Technical
- Roadmap
- Known Issues
- Recent Changes
- License
World Engine is a workspace of two pieces:
- Webserver: SvelteKit
- Backend Server: Bun server
Both of these require Bun to run. If you don't have Bun, you'll need to install it.
1. Install dependencies (both sides)
bun install && cd src/web && bun install2. Run both dev servers
Easiest path — if you have Zellij installed, use the bundled layout that opens both panes plus a Socket.IO monitor:
./start.shOtherwise, run them in two terminals:
# terminal A — backend on :3000
bun --hot src/server.ts
# terminal B — frontend on :5173 (proxies /socket.io + /api to :3000)
cd src/web && bun run dev3. Open the app
Visit http://localhost:5173.
On first run you'll land on a config page where you set API keys and pick which LLMs run the narrator / archivist / interpreter stages. There are built-in preset suggestions plus full custom mode.
Settings save automatically to config.local.json.
Don't have a good GPU? You can use a fully hosted configuration. Or, if you have the hardware, you can run it locally using LM Studio and the local TTS module.
The bundled stories — Cellar of Glass, Lunar Rescue, and The Last Train — give you a setting and a goal. Great for a focused session or first-time play.
But for the best experience, start with no preset and type your opening line. You genuinely don't know what you're going to get, and that's the point. Your mind shapes the adventure. The world grows around whatever premise you bring.
Note on quality The game is driven by the narrative. The creativity of the model you use is directly responsible for the quality of the gameplay. The narrator is the MOST IMPORTANT model — it creates the world, the items you find in that world, and the puzzles.
Look around often! Rooms, objects, and details only exist after the narrator establishes them. The more you look around, the more clues and items you'll find.
Empty World mode has no preset, so the narrator improvises a setting from your first input. look around gives it nothing to anchor to, so you'll get a generic room.
A scene-setting line does the work for you:
Don't open the airlock! We're on a deep-space station with over 300,000 people aboard. Open it and you'll suck them all into space.
One sentence gives the narrator a location, stakes, and a constraint to honor. From there, the world fills itself in around you.
World Engine can read each turn aloud. Pick a path.
All three paths are toggleable from the in-app config page (Narration section). Settings persist to config.local.json.
Uses ResembleAI Chatterbox Turbo via a small Python sidecar that Bun spawns automatically on startup.
# One-time install (Python 3.11+)
python -m venv tts_sidecar/.venv
source tts_sidecar/.venv/bin/activate
pip install -r tts_sidecar/requirements.txt
# One-time: generate the bundled voice references
python tts_sidecar/generate_voices.pyRoughly 5 GB disk and 5–10 minutes for the install. Chatterbox uses ~1.5 GB VRAM. CPU mode works but is ~10× slower. See tts_sidecar/README.md for details.
In the config page → Narration: Enable narration + Use local Chatterbox.
Hand off TTS to ElevenLabs. No Python, no GPU, but you'll need a paid plan. A free account gets you ~10k credits — enough for a few full games. Six bucks gets you a lot more.
In the config page → Keys: paste your ElevenLabs API key. Then Narration: Enable narration + Use ElevenLabs (cloud) + pick a model (e.g. eleven_flash_v2_5). Add voices as label:voice_id rows. Labels show up in the narration dropdown — rename them to whatever you want. Browse voice IDs at the ElevenLabs voice library. eleven_flash_v2_5 is the half-price tier with great quality; eleven_multilingual_v2 is the higher-cost flagship.
In the config page → Narration: leave Enable narration unchecked. The game runs fine silently. Skip the Python install too.
Toggle narration with the voice off / voice on button in the action bar, where you'll also find a voice selector. Audio is cached per turn, so replays are instant.
World Engine can generate an image for each turn with Nano Banana.
Toggle images with the images off / images on button in the action bar. Images are cached per turn, so replays are instant. You can also change the art style.
When a turn produces an image you want to keep, click the ★ button next to the ▦. This saves the image to the server's media/ directory.
To browse everything you've saved, type /gallery in the chat window. A wide modal opens with the selected image on top, prev/next arrows on either side, and a horizontally scrolling thumbnail strip below.
Click any thumbnail to jump to it. Click the big image to lightbox it. Use ← / → to navigate and ESC to close.
Everything is configured from the in-app Config page (the Configure button on the title screen). Pick a preset bundle, drop in API keys, toggle narration and images. Settings autosave.
The config page ships with four ready-to-go bundles that set all three model roles (narrator, archivist, interpreter) at once:
- Hybrid premium — premium narrator on OpenRouter, lightweight cloud models for archivist / interpreter.
- All cloud — Anthropic Claude family across all three roles.
- All cloud cheap — Gemini Flash everywhere; the cheapest cloud path.
- All local — everything runs against your LM Studio instance.
Pick one with a single click, then tweak till it's perfect.
The Scenario Creator button on the title screen opens an in-app editor for building your own presets — no files to touch, no markdown to hand-edit. Pick an existing scenario to edit, or save under a new name to create one. Everything you change in the editor saves back automatically.
Four tabs, and your edits stick as you move between them:
Scenarios — pick a scenario to load. You'll see banner thumbnails for every scenario you've got.
Character
Set a name and bio (notes for you, the author), plus a list of appearance details — anything you want the model to know about your character (hair, age, build, piercings, scars). A live portrait preview renders your character as you type, with shot options (T-pose / close-up / mid / custom) so you can dial in the look without leaving the page.
Narrative
Write the story-starter — the briefing your player reads at the very beginning — plus a set of trait cards. Pick a class (Human / Magic / Animal) for sensible defaults, then add traits and limits as tags. The portrait preview here drops your character into the story-starter as a backdrop, so you can see them in their world before saving.
World & Save
Place starting objects and objectives (objectives can optionally be pinned to a specific tile with @ n,e), and generate a title image — a 21:9 ultra-wide banner rendered through OpenRouter's image model.
The banner button disables itself if no OpenRouter key is set; it uses GPT Image 2 to generate title images.
- Keep objects loose. Use suggestive names like "Note" or "broken device" instead of "a note from Dr. Chen explaining the coordinates." The model fills in the details differently each playthrough — one run's "Note" becomes a cryptic warning from the navigator, another becomes a supply manifest with a coffee stain. Vague objects create variety.
- Use culturally iconic items. Models already know what a Tricorder, a lightsaber, or a spellbook does. You don't need to explain them — just name them and the model handles the rest.
- Write the briefing, not the plot. The story-starter sets the stage; the model writes what happens. Give it a world with tension and let it surprise you.
- Narrator — anthropic/claude-opus-latest via OpenRouter
- Archivist / Interpreter — google/gemma-3-12b via LM Studio
- Audio — ElevenLabs (cloud TTS)
- Images — Gemini (Nano Banana, per-turn)
The narrator has the most important job. It drives the story and creates the world. Different narrator models produce radically different results. Experiment till you figure out what works best for you.
eleven_flash_v2_5 is the half-price tier with excellent quality and ~300ms latency — fast enough that narration starts almost instantly after the narrator's first sentence streams. A free account gets ~10k credits (a few full games); six bucks gets you a lot more. See Optional: Narration → Path B for setup.
Nano Banana renders a per-turn scene at low cost and consistent quality. The preset's appearance: field anchors the character so renders stay visually coherent across turns. See Optional: Per-Turn Images.
We extensively tested local models — here are our findings. As all computers and hardware configurations are different, your mileage may vary.
Each turn runs three model passes:
- Interpreter parses the player's input into a structured action: movement, look, interact, or freeform.
- Narrator receives the established world state, active threads, and parsed action, then writes 1–3 sentences of narrative.
- Archivist reads the narrative and extracts new world facts and any objectives that just got achieved.
The world state lives in world-stack.json: an append-mostly list of established facts (e.g. damaged transmitter half-buried in regolith), an active-threads list, an objectives list, and a position. Every turn also appends to a single play-log.jsonl for postmortem.
- Backend: Bun +
node:http+ Socket.IO —src/server.ts(engine, dispatch, presets, services) - Frontend: SvelteKit + Vite under
src/web/(Svelte 5 runes; routes insrc/web/src/routes/) - Transport: Socket.IO for game events; one HTTP route (
GET /api/initial-state) for the title-screen preset list. In dev, Vite (:5173) proxies both to the Bun server (:3000); in prod, the SvelteKit Node adapter serves the frontend alongside Socket.IO. - Providers:
src/socket-io/services/*.ts— drop-in service classes (openrouter, lmstudio, gemini, elevenlabs) discovered at runtime bysrc/providers.ts. Adding a new provider is a single file drop. - State: plain JSON file (
world-stack.json); fine for single-user, but multi-user would need per-user namespacing.
sampling-presets.json lets you override sampling parameters (temperature, top_p, min_p) per model for the narrator. This is a narrator-only tuning knob — the archivist and interpreter need deterministic structured output and should not be touched.
{
"anthropic/claude-opus-latest": { "temperature": 1.2, "top_p": 0.9 },
"google/gemma-3-12b": { "temperature": 1.2, "min_p": 0.1 },
"default": { "temperature": 1.0 }
}Keys are model IDs as they appear in your provider config. The "default" key is the fallback when no model-specific entry matches. The file is re-read on every narrator call, so you can edit values while the game is running and hear the difference on the next turn — no restart required.
This is most useful when running multiple models (e.g. Opus for narration, Gemma for extraction), since each model responds differently to the same temperature. A temperature that produces rich prose from Opus might produce incoherent output from a smaller model.
Bake-off script: To systematically compare sampling parameters across a matrix of settings, use the included CLI tool:
bun src/sampling-bakeoff.ts
bun src/sampling-bakeoff.ts --provider openrouter --model anthropic/claude-opus-latest
bun src/sampling-bakeoff.ts --provider local --model google/gemma-3-12b --output results.mdThis fires the same narrator prompt across a grid of temperature / top_p / min_p values and dumps a readable markdown report. Reads provider config from config.local.json.
- Self-building exploration map
- Local Stable Diffusion image generation (maybe)
-
Cardinal-based movement. The world grid supports four cardinal directions. Compound directions (northeast, SW, etc.) are accepted and resolve to the primary cardinal component. Relative directions like
forward,up, ordownare not yet supported — usenorth,south,east, orwest. -
Threads paraphrase across turns. Active narrative threads can subtly rephrase between turns ("find out what caused the leak" → "what caused the leak?"). They have no stable IDs yet; the archivist re-writes them on each round-trip.
-
No first-class inventory. Items you "take" remain in the world's established entries; there's no inventory data structure yet. Type
inventoryand the narrator synthesizes a list from what it knows, but a specific picked-up item may not make the cut.
Balance is key. We're working on it. 😅
-
Scenario Editor — author presets in-browser. New
/editorroute opens a four-tab UI for creating or editing scenarios end-to-end without touching the filesystem. Scenarios tab picks a preset; Character tab edits name, bio, and a dynamic appearance property list with a live PortraitPanel (T-pose / close-up / mid / custom shots) rendered from the current form state; Narrative tab handles story-starter prose and categorized attribute cards (Human / Magic / Animal presets with chip-add traits) with scene-aware character renders; World & Save tab manages objects, objectives (with optional@ n,eanchors), banner generation (21:9 ultra-wide via OpenRouter, using the latest scene render as multimodal reference), and the save flow. All tabs share a single runed store so edits persist across navigation; save serializes back topresets/<slug>.md+presets/<slug>.png, with rename detection that cleans up old files. Server side,/api/preset/:slugreturns the full Preset shape for load,/api/initial-stateships a capabilities flag so banner controls self-disable without an API key, andimage:generate/banner:generate/preset:saveround-trip through dispatch. -
Shared
AppHeadernav. Title page and config page now use the same sticky header component (New Game / Scenario Creator / Image View / Configure) instead of ad-hoc corner buttons and back links. Active route gets an ember underline. Replaces the previous.title-page-actionsfloating-button row on/and the← back to menurow on/config. -
Structured
appearance:field for visual coherence. Image models render a prose paragraph dramatically more consistently than a bullet list of mixed visual/gameplay facts, so the preset schema gained an optionalappearance:block that's used in place of the attribute dump when building image prompts. Authored either as a single prose line or as a YAML sub-block of key/value pairs (hair: red, eyes: green, …) — both round-trip through the editor. Presets withoutappearance:fall back to the old bullet path, no behavior change. All four bundled presets were rewritten with tight appearance prose (subject → face → hair → outfit → palette → archetype) and new title images rendered through the character probe. -
Player appearance preserved across turns.
processInputrebuilt the stack each turn but only carried over attributes, not appearance — so the field dropped toundefinedon turn 1 and stayed there. First image rendered fine (seeded byapplyPresetToStack); every subsequent turn fell through to the bullet-only fallback path, letting the image model freestyle the character's gender, hair, and outfit. Now copies appearance like every other persistent field;engine-init.generateImagealso backfills from the live presets map when a saved stack predates the field.
MIT — see LICENSE







