Skip to content

About

A text adventure where the world is generated turn by turn by an LLM.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

World Engine — A world that pushes back

World Engine

A text adventure where the world is generated turn-by-turn as you explore. Nothing is placed; everything is improvised.


There's a chest. It's locked.

The objective says: open the chest. But the game doesn't tell you how, because it doesn't know yet.

There's no designated key waiting in a designated drawer. The lock might be jammed, in which case the key you find won't help. So you wander.

A few rooms later, you find an axe. The axe wasn't placed there to solve the chest. Nothing is placed to solve anything. Nothing was placed at all until you walked into the room.

But you're holding an axe, and there's a locked chest. When you swing at the lid, the world agrees that's what happens.

Rooms, objects, and complications are generated turn by turn as you explore. Some of what appears will turn out to matter. You won't know which things matter until you try something.

Replay the same preset and the axe might not exist. Maybe there's a crowbar. Maybe the chest is rusted shut and now you're looking for water.

Maybe the chest contains something you really wish it hadn't.

The seed scenario is the same; what fills it in isn't.

You can also skip the presets entirely and start in an empty open world. Type what you do, and the world assembles itself around you.

Note: Some of the videos are older and have different interfaces, but the concepts are still the same.

Demo: Promotional Video

Demo: Gameplay video #1 — Full game of Merlin's Daughter (20m)

Demo: Gameplay video #2 — streaming TTS narration and per-turn image generation in action. Slightly old — new UI and logic now, but still a good example.

Demo: Editor video #1 — Create your own scenarios!


Contents


Quick Start

World Engine is a workspace of two pieces:

  • Webserver: SvelteKit
  • Backend Server: Bun server

Both of these require Bun to run. If you don't have Bun, you'll need to install it.

1. Install dependencies (both sides)

bun install && cd src/web && bun install

2. Run both dev servers

Easiest path — if you have Zellij installed, use the bundled layout that opens both panes plus a Socket.IO monitor:

./start.sh

Otherwise, run them in two terminals:

# terminal A — backend on :3000
bun --hot src/server.ts

# terminal B — frontend on :5173 (proxies /socket.io + /api to :3000)
cd src/web && bun run dev

3. Open the app

Visit http://localhost:5173.

On first run you'll land on a config page where you set API keys and pick which LLMs run the narrator / archivist / interpreter stages. There are built-in preset suggestions plus full custom mode.

Settings save automatically to config.local.json.

Config Page

Don't have a good GPU? You can use a fully hosted configuration. Or, if you have the hardware, you can run it locally using LM Studio and the local TTS module.


How to Play

Preset scenarios are fun; Empty World is where the real magic happens

The bundled stories — Cellar of Glass, Lunar Rescue, and The Last Train — give you a setting and a goal. Great for a focused session or first-time play.

But for the best experience, start with no preset and type your opening line. You genuinely don't know what you're going to get, and that's the point. Your mind shapes the adventure. The world grows around whatever premise you bring.

Note on quality The game is driven by the narrative. The creativity of the model you use is directly responsible for the quality of the gameplay. The narrator is the MOST IMPORTANT model — it creates the world, the items you find in that world, and the puzzles.

The world builds as you look

Look around often! Rooms, objects, and details only exist after the narrator establishes them. The more you look around, the more clues and items you'll find.

Seed the scene on turn one in freeplay

Empty World mode has no preset, so the narrator improvises a setting from your first input. look around gives it nothing to anchor to, so you'll get a generic room.

A scene-setting line does the work for you:

Don't open the airlock! We're on a deep-space station with over 300,000 people aboard. Open it and you'll suck them all into space.

One sentence gives the narrator a location, stakes, and a constraint to honor. From there, the world fills itself in around you.


Optional: Narration

World Engine can read each turn aloud. Pick a path.

All three paths are toggleable from the in-app config page (Narration section). Settings persist to config.local.json.

Path A — Local TTS (free, requires Python + GPU)

Uses ResembleAI Chatterbox Turbo via a small Python sidecar that Bun spawns automatically on startup.

# One-time install (Python 3.11+)
python -m venv tts_sidecar/.venv
source tts_sidecar/.venv/bin/activate
pip install -r tts_sidecar/requirements.txt

# One-time: generate the bundled voice references
python tts_sidecar/generate_voices.py

Roughly 5 GB disk and 5–10 minutes for the install. Chatterbox uses ~1.5 GB VRAM. CPU mode works but is ~10× slower. See tts_sidecar/README.md for details.

In the config page → Narration: Enable narration + Use local Chatterbox.

Path B — ElevenLabs (cloud, paid, no local install)

Hand off TTS to ElevenLabs. No Python, no GPU, but you'll need a paid plan. A free account gets you ~10k credits — enough for a few full games. Six bucks gets you a lot more.

In the config page → Keys: paste your ElevenLabs API key. Then Narration: Enable narration + Use ElevenLabs (cloud) + pick a model (e.g. eleven_flash_v2_5). Add voices as label:voice_id rows. Labels show up in the narration dropdown — rename them to whatever you want. Browse voice IDs at the ElevenLabs voice library. eleven_flash_v2_5 is the half-price tier with great quality; eleven_multilingual_v2 is the higher-cost flagship.

Path C — Skip narration entirely

In the config page → Narration: leave Enable narration unchecked. The game runs fine silently. Skip the Python install too.

In-game controls

Toggle narration with the voice off / voice on button in the action bar, where you'll also find a voice selector. Audio is cached per turn, so replays are instant.


Optional: Per-Turn Images

World Engine can generate an image for each turn with Nano Banana.

Toggle images with the images off / images on button in the action bar. Images are cached per turn, so replays are instant. You can also change the art style.

Gallery

When a turn produces an image you want to keep, click the ★ button next to the ▦. This saves the image to the server's media/ directory.

To browse everything you've saved, type /gallery in the chat window. A wide modal opens with the selected image on top, prev/next arrows on either side, and a horizontally scrolling thumbnail strip below.

Click any thumbnail to jump to it. Click the big image to lightbox it. Use ← / → to navigate and ESC to close.


Configuration

Config Page — Selection Filter

Everything is configured from the in-app Config page (the Configure button on the title screen). Pick a preset bundle, drop in API keys, toggle narration and images. Settings autosave.

Preset bundles

Config Page — Preset Bundles

The config page ships with four ready-to-go bundles that set all three model roles (narrator, archivist, interpreter) at once:

  • Hybrid premium — premium narrator on OpenRouter, lightweight cloud models for archivist / interpreter.
  • All cloud — Anthropic Claude family across all three roles.
  • All cloud cheap — Gemini Flash everywhere; the cheapest cloud path.
  • All local — everything runs against your LM Studio instance.

Pick one with a single click, then tweak till it's perfect.


Scenario Editor

Scenario Editor overview

The Scenario Creator button on the title screen opens an in-app editor for building your own presets — no files to touch, no markdown to hand-edit. Pick an existing scenario to edit, or save under a new name to create one. Everything you change in the editor saves back automatically.

Four tabs, and your edits stick as you move between them:

Scenarios — pick a scenario to load. You'll see banner thumbnails for every scenario you've got.

Character

Character tab

Set a name and bio (notes for you, the author), plus a list of appearance details — anything you want the model to know about your character (hair, age, build, piercings, scars). A live portrait preview renders your character as you type, with shot options (T-pose / close-up / mid / custom) so you can dial in the look without leaving the page.

Narrative

Narrative tab

Write the story-starter — the briefing your player reads at the very beginning — plus a set of trait cards. Pick a class (Human / Magic / Animal) for sensible defaults, then add traits and limits as tags. The portrait preview here drops your character into the story-starter as a backdrop, so you can see them in their world before saving.

World & Save

World and Save tab

Place starting objects and objectives (objectives can optionally be pinned to a specific tile with @ n,e), and generate a title image — a 21:9 ultra-wide banner rendered through OpenRouter's image model.

The banner button disables itself if no OpenRouter key is set; it uses GPT Image 2 to generate title images.

Tips for Better Scenarios

  • Keep objects loose. Use suggestive names like "Note" or "broken device" instead of "a note from Dr. Chen explaining the coordinates." The model fills in the details differently each playthrough — one run's "Note" becomes a cryptic warning from the navigator, another becomes a supply manifest with a coffee stain. Vague objects create variety.
  • Use culturally iconic items. Models already know what a Tricorder, a lightsaber, or a spellbook does. You don't need to explain them — just name them and the model handles the rest.
  • Write the briefing, not the plot. The story-starter sets the stage; the model writes what happens. Give it a world with tension and let it surprise you.

My Personal Configuration

  • Narrator — anthropic/claude-opus-latest via OpenRouter
  • Archivist / Interpreter — google/gemma-3-12b via LM Studio
  • Audio — ElevenLabs (cloud TTS)
  • Images — Gemini (Nano Banana, per-turn)

The narrator has the most important job. It drives the story and creates the world. Different narrator models produce radically different results. Experiment till you figure out what works best for you.

Audio — ElevenLabs

eleven_flash_v2_5 is the half-price tier with excellent quality and ~300ms latency — fast enough that narration starts almost instantly after the narrator's first sentence streams. A free account gets ~10k credits (a few full games); six bucks gets you a lot more. See Optional: Narration → Path B for setup.

Images — Gemini

Nano Banana renders a per-turn scene at low cost and consistent quality. The preset's appearance: field anchors the character so renders stay visually coherent across turns. See Optional: Per-Turn Images.


Technical

Model research

We extensively tested local models — here are our findings. As all computers and hardware configurations are different, your mileage may vary.

How it works

Each turn runs three model passes:

  1. Interpreter parses the player's input into a structured action: movement, look, interact, or freeform.
  2. Narrator receives the established world state, active threads, and parsed action, then writes 1–3 sentences of narrative.
  3. Archivist reads the narrative and extracts new world facts and any objectives that just got achieved.

The world state lives in world-stack.json: an append-mostly list of established facts (e.g. damaged transmitter half-buried in regolith), an active-threads list, an objectives list, and a position. Every turn also appends to a single play-log.jsonl for postmortem.

Architecture

  • Backend: Bun + node:http + Socket.IO — src/server.ts (engine, dispatch, presets, services)
  • Frontend: SvelteKit + Vite under src/web/ (Svelte 5 runes; routes in src/web/src/routes/)
  • Transport: Socket.IO for game events; one HTTP route (GET /api/initial-state) for the title-screen preset list. In dev, Vite (:5173) proxies both to the Bun server (:3000); in prod, the SvelteKit Node adapter serves the frontend alongside Socket.IO.
  • Providers: src/socket-io/services/*.ts — drop-in service classes (openrouter, lmstudio, gemini, elevenlabs) discovered at runtime by src/providers.ts. Adding a new provider is a single file drop.
  • State: plain JSON file (world-stack.json); fine for single-user, but multi-user would need per-user namespacing.

Sampling presets (experimental)

sampling-presets.json lets you override sampling parameters (temperature, top_p, min_p) per model for the narrator. This is a narrator-only tuning knob — the archivist and interpreter need deterministic structured output and should not be touched.

{
  "anthropic/claude-opus-latest": { "temperature": 1.2, "top_p": 0.9 },
  "google/gemma-3-12b": { "temperature": 1.2, "min_p": 0.1 },
  "default": { "temperature": 1.0 }
}

Keys are model IDs as they appear in your provider config. The "default" key is the fallback when no model-specific entry matches. The file is re-read on every narrator call, so you can edit values while the game is running and hear the difference on the next turn — no restart required.

This is most useful when running multiple models (e.g. Opus for narration, Gemma for extraction), since each model responds differently to the same temperature. A temperature that produces rich prose from Opus might produce incoherent output from a smaller model.

Bake-off script: To systematically compare sampling parameters across a matrix of settings, use the included CLI tool:

bun src/sampling-bakeoff.ts
bun src/sampling-bakeoff.ts --provider openrouter --model anthropic/claude-opus-latest
bun src/sampling-bakeoff.ts --provider local --model google/gemma-3-12b --output results.md

This fires the same narrator prompt across a grid of temperature / top_p / min_p values and dumps a readable markdown report. Reads provider config from config.local.json.


Roadmap

  • Self-building exploration map
  • Local Stable Diffusion image generation (maybe)

Known Issues

  • Cardinal-based movement. The world grid supports four cardinal directions. Compound directions (northeast, SW, etc.) are accepted and resolve to the primary cardinal component. Relative directions like forward, up, or down are not yet supported — use north, south, east, or west.

  • Threads paraphrase across turns. Active narrative threads can subtly rephrase between turns ("find out what caused the leak" → "what caused the leak?"). They have no stable IDs yet; the archivist re-writes them on each round-trip.

  • No first-class inventory. Items you "take" remain in the world's established entries; there's no inventory data structure yet. Type inventory and the narrator synthesizes a list from what it knows, but a specific picked-up item may not make the cut.

Balance is key. We're working on it. 😅


Recent Changes

  • Scenario Editor — author presets in-browser. New /editor route opens a four-tab UI for creating or editing scenarios end-to-end without touching the filesystem. Scenarios tab picks a preset; Character tab edits name, bio, and a dynamic appearance property list with a live PortraitPanel (T-pose / close-up / mid / custom shots) rendered from the current form state; Narrative tab handles story-starter prose and categorized attribute cards (Human / Magic / Animal presets with chip-add traits) with scene-aware character renders; World & Save tab manages objects, objectives (with optional @ n,e anchors), banner generation (21:9 ultra-wide via OpenRouter, using the latest scene render as multimodal reference), and the save flow. All tabs share a single runed store so edits persist across navigation; save serializes back to presets/<slug>.md + presets/<slug>.png, with rename detection that cleans up old files. Server side, /api/preset/:slug returns the full Preset shape for load, /api/initial-state ships a capabilities flag so banner controls self-disable without an API key, and image:generate / banner:generate / preset:save round-trip through dispatch.

  • Shared AppHeader nav. Title page and config page now use the same sticky header component (New Game / Scenario Creator / Image View / Configure) instead of ad-hoc corner buttons and back links. Active route gets an ember underline. Replaces the previous .title-page-actions floating-button row on / and the ← back to menu row on /config.

  • Structured appearance: field for visual coherence. Image models render a prose paragraph dramatically more consistently than a bullet list of mixed visual/gameplay facts, so the preset schema gained an optional appearance: block that's used in place of the attribute dump when building image prompts. Authored either as a single prose line or as a YAML sub-block of key/value pairs (hair: red, eyes: green, …) — both round-trip through the editor. Presets without appearance: fall back to the old bullet path, no behavior change. All four bundled presets were rewritten with tight appearance prose (subject → face → hair → outfit → palette → archetype) and new title images rendered through the character probe.

  • Player appearance preserved across turns. processInput rebuilt the stack each turn but only carried over attributes, not appearance — so the field dropped to undefined on turn 1 and stayed there. First image rendered fine (seeded by applyPresetToStack); every subsequent turn fell through to the bullet-only fallback path, letting the image model freestyle the character's gender, hair, and outfit. Now copies appearance like every other persistent field; engine-init.generateImage also backfills from the live presets map when a saved stack predates the field.


License

MIT — see LICENSE

About

A text adventure where the world is generated turn by turn by an LLM.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages