Skip to content

About

An Ai powered agentic chatbot with capability of tool orchestrating, mcp and personalized responses with the use of long term memory

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

Twilight

An async, memory-having, tool-wielding chatbot that remembers you're allergic to peanuts three conversations later.


Most "I built a chatbot" projects are a while True: input() loop wrapped around an API call. This isn't that. This is what happens when you ask "what would it actually take to build this properly" and then don't stop until the answer includes async subprocess orchestration, a Postgres-backed vector store, and a context-window management system that decides — on its own, every single turn — what's worth remembering and what's safe to forget.

Twilight is a personal AI assistant with a name, a memory, a toolbox, and opinions about your Desktop folder organization (it once tried to help me clean up shortcuts). It's built on LangGraph, runs entirely on free-tier infra, and every architectural decision in it was made — and re-made, and debugged, and re-debugged — one deliberate step at a time rather than scaffolded out by a framework CLI. That distinction matters more than it sounds like it should.

Table of contents

What it actually does

🧠 A real agentic core, not a prompt template

Built on LangGraph, with an explicit state machine — not a chain, not a single prompt. The graph routes between a chat node, a tool-execution node, a summarization node, and a memory-retrieval node, every turn, conditionally, based on what the LLM actually decides to do.

🔧 Tool use across local and remote MCP servers

Three local tools (calculator, web search, weather) plus a live MultiServerMCPClient connection to a calculator MCP server, a filesystem MCP server (scoped, not free-range), and a fetch MCP server — all spawned as subprocesses over stdio, connected to concurrently, not one at a time. Tool provenance is tracked, so the UI can tell you not just that a tool ran, but which server it came from.

📡 Real streaming, not a spinner

Token-by-token streaming over a persistent WebSocket, with live "connecting to X MCP server…" status updates rendered inline as the model works — the same transparency you'd want from any agent you're trusting to touch your filesystem.

🧵 Threads that actually persist

Every conversation is a resumable thread, backed by a Postgres-backed LangGraph checkpointer (AsyncPostgresSaver) — not an in-memory dict that dies on restart. Auto-generated, LLM-written conversation titles. A sidebar you can click back into, mid-conversation, and pick up exactly where you left off — full history intact.

✂️ Context that doesn't just... overflow

Long conversations don't get silently truncated or blow the context window. A dedicated summarizer node watches token count, folds aged-out messages into a compact, structured running summary, and keeps only the freshest exchanges in full — while the complete, untrimmed history stays persisted for the frontend and for resuming threads later. The model sees a compressed version of the past; you never lose any of it.

🧬 Long-term, cross-conversation memory (the fun part)

This is the feature that makes Twilight feel like your assistant instead of an assistant. After every turn, a background task asks: is there anything here worth remembering permanently? If yes, it's embedded and stored in Postgres via pgvector. On every new message — in any thread, any day — a semantic similarity search silently checks whether anything relevant is already known, and if so, injects it into context. No re-explaining your allergies. No re-introducing yourself. It just remembers, the way a person would — including declining to bring anything up when nothing's actually relevant, rather than forcing irrelevant facts into every reply.

Architecture

graph TD
    A[Browser: WebSocket] -->|user message| B[FastAPI /ws endpoint]
    B --> C[LangGraph]
    C --> D[summarizer_node]
    D --> E[memory_retrieval_node]
    E --> F[chat_node]
    F -->|tool call?| G[tools node]
    G --> F
    F -->|response| B
    B -->|streamed tokens| A

    F -.async, non-blocking.-> H[background: extract + store memory]
    H --> I[(Postgres: long_term_memories + pgvector)]
    E -.reads.-> I
    C -.checkpoints every turn.-> J[(Postgres: AsyncPostgresSaver)]

    F -.tool calls.-> K[Local tools]
    F -.tool calls.-> L[MCP servers: calculator / filesystem / fetch]
Loading

Every box marked "Postgres" is the same Dockerized instance — one docker-compose up, no separate vector database, no separate cache layer. Checkpointing, thread metadata, and semantic memory all live in one place, because reaching for a new piece of infrastructure every time you need a new kind of persistence is how projects turn into distributed-systems homework nobody asked for.

Screenshots

The interface — sidebar with resumable threads, streaming responses, live tool-usage indicators:

Interface overview

Persistent, titled conversation threads — every conversation survives a restart, gets an LLM-generated title, and is one click away:

Sidebar with conversation threads

Cross-conversation memory, actually working — this is a brand new thread. Twilight was never told about a peanut allergy in this conversation. It remembered it from a completely separate one, days earlier:

Long-term memory recall across separate conversations

Worth calling out honestly: notice the first phrasing — "am I allergic to anything" — comes back with a generic non-answer, while the more specific "what do you know about me and peanuts" successfully retrieves the stored memory. That gap is a direct, visible consequence of running a small, free, local embedding model (see Known limitations) rather than a fault in the retrieval logic itself — a vaguer query embeds further away from the specifically-worded stored fact than a more targeted one does. The fix isn't "rebuild the system," it's "use a stronger embedding model and tighten the phrasing the model stores facts in" — both already scoped as next steps.

Tech stack

Layer Tools
Orchestration LangGraph (async StateGraph), LangChain core
LLM inference Groq API (openai/gpt-oss-120b)
Tool protocol Model Context Protocol — MultiServerMCPClient, stdio transport
Embeddings sentence-transformers (all-MiniLM-L6-v2), local, free, zero API dependency
Persistence PostgreSQL 16 (pgvector extension), AsyncPostgresSaver, psycopg + connection pooling
Backend FastAPI, native WebSockets, asyncio throughout
Frontend Hand-rolled HTML/CSS/vanilla JS — no framework, no build step
Infra Docker + Docker Compose, named volumes for persistence

The debugging trail

If you're evaluating this project, the commit history (or the conversation log this was built from) tells a more honest story than the feature list does. Getting here meant working through genuine, unglamorous systems problems: Windows subprocess spawning silently failing because npx is a batch script and not a real executable; an MCP reference server breaking because its pinned SDK dependency shipped a breaking rename upstream; a Docker Desktop containerd corruption that took five escalating fixes — image re-pulls, a full system prune, a WSL2 restart — before it was actually resolved; a hosted embeddings API returning a mysterious, undocumented 403 that turned out to be a known bug affecting other developers, not a config mistake. None of these were fixed by pattern-matching to a tutorial. Every one required reading the actual traceback, isolating the variable, and testing the fix standalone before trusting it — the same discipline that separates "it works on my machine" from a system you can actually reason about.

Known limitations & honest gaps

Being upfront about what's not here is part of what makes the rest of it credible.

  • No real authentication or multi-user isolation yet. USER_ID is currently hardcoded. The long-term memory architecture is fully multi-user-capable (every row is already scoped by user_id), but there's no login layer sitting in front of it yet — a deliberate sequencing choice, to get the memory mechanics right before adding identity on top.
  • Running on Groq's free tier, which means occasional tool-calling reliability hiccups — malformed streamed JSON, or the model hallucinating a plausible-sounding tool name that was never actually bound. These are documented, understood, and not yet wrapped in the retry/fallback logic they deserve — that's the next infrastructure investment, not a mystery bug.
  • Embeddings run locally on a small, free model (all-MiniLM-L6-v2, 384 dimensions) rather than a stronger hosted model. It's fast and costs nothing, but it leans more heavily on literal word overlap than a larger model would — a euphemistically-phrased memory can rank below an unrelated one for a very plainly-worded query. This was caught through actual similarity-distance testing (not assumed away), and the fix — tighter extraction prompting plus a recalibrated relevance threshold — is scoped and pending.
  • RAG (retrieval over external documents) isn't built yet. It was the deliberately-deferred next step after long-term memory, on the reasoning that RAG's retrieval pattern is close enough to LTM's that building LTM first makes RAG faster to add later, not slower.

What's next

  • Lightweight authentication → genuine multi-user support (the schema's already ready for it)
  • Retry/fallback handling around Groq's occasional streaming and tool-call failures
  • Recalibrated, better-phrased long-term memory extraction and retrieval
  • Retrieval-augmented generation over user-provided documents
  • Possibly a proper React frontend, once there's a reason to outgrow the hand-rolled one

Built one deliberate, argued-over, occasionally-yelled-at-Docker-for decision at a time.

About

An Ai powered agentic chatbot with capability of tool orchestrating, mcp and personalized responses with the use of long term memory

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages