A lightweight Retrieval-Augmented Generation (RAG) chatbot for querying medical documents. Upload one or more PDF files, extract and split their content, create embeddings, store them in a FAISS vector index, retrieve the most relevant document context, and generate an answer using an LLM.
โ ๏ธ Disclaimer: This project is intended for informational and educational purposes only. It is not a medical device and does not provide medical advice, diagnosis, or treatment. Always consult a qualified healthcare professional for medical decisions.
๐ง Project Status: Working prototype | PDF-based RAG pipeline | FAISS semantic retrieval | Streamlit interface
PDF Upload โ Text Extraction โ Chunking โ Embeddings โ FAISS Vector Store โ Similarity Retrieval โ Context โ LLM โ Answer
- ๐ PDF Upload โ Upload one or multiple medical PDF documents through the Streamlit interface
- ๐ Semantic Retrieval โ Retrieve relevant document chunks using FAISS similarity search
- ๐ค Context-Grounded Answers โ Generate responses using retrieved content from uploaded documents
- โก Real-Time Indexing โ Extract, chunk, embed, and index uploaded documents during the session
- ๐ง RAG Pipeline โ Combines document retrieval with LLM-based generation
- ๐ Retrieved Context Viewer โ Inspect the document chunks retrieved for a query
- ๐ฅ๏ธ Streamlit UI โ Clean interactive interface with sidebar document upload
- ๐ Environment-Based Secrets โ API credentials are loaded through
.env - ๐ Multi-Document Support โ Query information across multiple uploaded PDFs
| Layer | Technology |
|---|---|
| Programming Language | Python |
| UI | Streamlit |
| RAG Framework | LangChain |
| Vector Store | FAISS |
| LLM Integration | EURI API |
| Text Splitting | RecursiveCharacterTextSplitter |
| PDF Processing | PyPDF2 / pdfplumber |
| Environment Management | python-dotenv |
โโโโโโโโโโโโโโโโโโโโโโโ
โ User Uploads PDF โ
โ (Streamlit) โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ PDF Text โ
โ Extraction โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Text Chunking โ
โ 1000 / 200 overlap โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Embeddings โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ FAISS Vector โ
โ Index โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
User Question
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Similarity Search โ
โ (Top-K) โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Retrieved Context โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ LLM Generation โ
โ (EURI) โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Final Answer โ
โ (Streamlit) โ
โโโโโโโโโโโโโโโโโโโโโโโ
- The user uploads one or more PDF documents.
- Text is extracted from the uploaded files.
- Documents are divided into overlapping chunks.
- Chunks are converted into vector embeddings.
- Embeddings are stored in an in-memory FAISS index.
- The user enters a question.
- FAISS performs semantic similarity search to retrieve relevant chunks.
- Retrieved content is supplied to the LLM as context.
- The LLM generates a response based on the retrieved document context.
- The application displays both the answer and the retrieved context.
Traditional LLM applications rely primarily on the model's pretrained knowledge. That approach makes it difficult to query information contained in user-provided documents.
This project uses Retrieval-Augmented Generation to connect an LLM with external document knowledge.
RAG provides the ability to:
- ๐ Retrieve relevant information from uploaded documents
- ๐ Perform semantic search rather than relying only on keywords
- ๐ง Give the LLM additional context at inference time
- ๐ Query private or domain-specific documents
- ๐ Inspect the retrieved context used to generate an answer
- ๐ Update the knowledge available to the application by uploading new documents
- End-to-end Retrieval-Augmented Generation pipeline
- PDF text extraction and preprocessing
- Recursive document chunking
- Embedding-based semantic retrieval
- FAISS vector similarity search
- Context injection into an LLM prompt
- Streamlit-based interactive frontend
- Multiple PDF document handling
- Runtime FAISS index construction
- Environment-variable based API authentication
- Retrieved-context inspection for improved transparency
Make sure you have:
- Python 3.9 or higher
- An EURI API key
- Git installed on your system
git clone https://github.com/VENKATRAM027/Medical-RAG-Chatbot.git
cd Medical-RAG-Chatbotpython -m venv venv
venv\Scripts\activatepython3 -m venv venv
source venv/bin/activatepip install -r requirements.txtCreate a .env file in the project root.
EURI_API_KEY=your_actual_api_key_hereYou can also copy the provided example configuration:
cp .env.example .envThen edit .env and add your API key.
๐ Security: Never commit your
.envfile or expose your API key publicly. The repository should contain only.env.examplewith placeholder values.
streamlit run main.pyThe application will normally be available at:
http://localhost:8501
Medical-RAG-Chatbot/
โ
โโโ app/
โ โโโ __init__.py
โ โโโ chat_utils.py
โ โโโ config.py
โ โโโ pdf_utils.py
โ โโโ ui.py
โ โโโ vectorstore_utils.py
โ
โโโ assets/
โ โโโ demo.gif
โ โโโ chatbot-ui.png
โ โโโ retrieved-context.png
โ โโโ architecture.png
โ
โโโ main.py
โโโ requirements.txt
โโโ .env.example
โโโ .gitignore
โโโ LICENSE
โโโ README.md
| File | Purpose |
|---|---|
main.py |
Application entry point and overall orchestration |
app/chat_utils.py |
LLM initialization and response generation |
app/config.py |
Environment variable and application configuration |
app/pdf_utils.py |
PDF text extraction and processing |
app/ui.py |
Streamlit UI components |
app/vectorstore_utils.py |
FAISS index creation and similarity search |
requirements.txt |
Python package dependencies |
.env.example |
Environment-variable template |
Important parameters can be configured in the application code.
| Parameter | Location | Default | Description |
|---|---|---|---|
chunk_size |
main.py |
1000 |
Maximum number of characters per chunk |
chunk_overlap |
main.py |
200 |
Number of overlapping characters between chunks |
EURI_API_KEY |
.env |
โ | Authentication key for EURI API |
Retrieval depth and LLM-related parameters can be adjusted in the corresponding files under app/.
Upload one or more medical PDF documents using the sidebar.
The application extracts the text, splits it into chunks, creates embeddings, and builds the FAISS index.
Enter a question related to the uploaded documents.
Example:
What are the contraindications mentioned in the uploaded document?
The application performs semantic similarity search and retrieves the most relevant document chunks.
The retrieved chunks are passed to the LLM as contextual information.
Use View Retrieved Context to inspect the document content retrieved for the query.
Medical PDF
โ
Text Extraction
โ
Recursive Chunking
โ
Embedding Generation
โ
FAISS Vector Index
โ
User Query
โ
Query Embedding
โ
Similarity Search
โ
Top Relevant Chunks
โ
Context Injection
โ
EURI LLM
โ
Generated Response
This project is currently designed as a lightweight prototype and has some limitations:
- FAISS indices are created in memory during the session
- Documents may need to be re-indexed after restarting the application
- Source-level citations such as PDF filename and page number are not yet implemented
- Retrieval quality depends on document quality, chunking strategy, embeddings, and retrieval parameters
- LLM responses may still contain inaccuracies
- The system should not be used for real-world diagnosis or treatment decisions
- Persistent FAISS Index โ Save and reload vector indices between sessions
- Source Citations โ Display PDF filename and page number for retrieved chunks
- Improved Retrieval โ Experiment with hybrid search, reranking, and better retrieval strategies
- Chat History โ Support multi-turn conversations with conversation context
- Multi-Model Support โ Add configurable LLM providers such as EURI, OpenAI, or local models
- Document Preview โ Preview uploaded PDFs inside the application
- Metadata Filtering โ Filter retrieval results by document or other metadata
- Evaluation Pipeline โ Measure retrieval and answer quality using RAG evaluation metrics
- Docker Support โ Containerized deployment for easier setup
- Cloud Deployment โ Deploy the application for public demonstration
Through this project, I gained hands-on experience with:
- Retrieval-Augmented Generation
- Vector databases and similarity search
- FAISS
- Document preprocessing
- Embeddings
- LangChain
- LLM integration
- Prompt construction
- Streamlit application development
- Environment-variable based secret management
- Building an end-to-end GenAI application
This application is a technical demonstration of RAG and document-based question answering.
It should not be treated as:
- A medical diagnosis system
- A clinical decision-support system
- A substitute for a doctor
- A source of personalized medical treatment
- A validated medical device
Always verify important medical information with qualified healthcare professionals and authoritative medical sources.
Contributions, suggestions, and improvements are welcome.
# Fork the repository
git clone https://github.com/VENKATRAM027/Medical-RAG-Chatbot.git
cd Medical-RAG-Chatbot
git checkout -b feature/your-feature
# Make your changes
git add .
git commit -m "Add your feature"
git push origin feature/your-featureThen open a Pull Request on GitHub.
This project is licensed under the MIT License.
See the LICENSE file for more information.
- LangChain โ RAG and LLM application framework
- FAISS โ Efficient vector similarity search
- Streamlit โ Interactive Python web application framework
- EURI โ LLM API integration
Venkatram
GitHub: @VENKATRAM027
Project: Medical RAG Chatbot