Skip to content

Repository files navigation

BidRAG

BidRAG

Open-source RAG extraction for bidding and tender documents.

Upload an RFQ, tender, or contract. BidRAG parses its layout, tables and signatures, embeds it into a vector store, and answers a configurable list of commercial and technical questions -- with citations back to the source page.

Everything runs locally and free except the language model. Layout analysis, OCR, embeddings, the vector store, file storage, authentication and email all work with no account and no bill. You bring one API key or none, if you run a model with Ollama.

Background: BidRAG is an open-source reimplementation of Bidify the internal application I worked on during my internship at Siemens, described in this write-up. It is built entirely on open-source technology, covers the same feature set, and runs lighter than the closed-source original. It is ofcourse not as technologically deep in every place, but feature for feature, function for function, and in its logic it matches the original and in places improves on it.

The presentations I gave during the internship are in Internship_Presentations and the evaluation write-ups and measured results are in Internship_Reports_and_Results_from_Evaluation.

How it all works, in depth: WORKING.md. Third-party licenses and attribution are in NOTICE.md.


See it working

Upload a tender, get sourced answers

BidRAG: log in, upload an RFQ, and read the extracted answers

Log in, drop in an RFQ, and BidRAG returns the commercial and technical fields already filled in rated voltage, busbar current, panel counts, delivery schedule each one traceable back to the page it came from.

Full resolution, full speed (MP4, 3.4 MB)

Every step visible in Langfuse

Langfuse: cost dashboard, ingestion and answer-generation traces, chunk references

The same run seen from the inside: the cost dashboard, the ingestion trace with every chunk and its stable reference, the refined queries, the exact prompt sent to the model, and which chunks the answer actually cited.

Full resolution, full speed (MP4, 3.3 MB)

Both clips are screen recordings played at 2x, with no audio. The document in them is a synthetic RFQ every company, person, address and bank detail in it is invented.


Quick start

git clone https://github.com/spearb0lt/BidRAG.git bidrag && cd bidrag
cp .env.example .env

Set the two required values in .env:

# 1. Token signing key
python -c "import secrets; print(secrets.token_urlsafe(48))"   # paste into JWT_SECRET_KEY

# 2. An LLM. Free options: Groq, Gemini, or Ollama (no key at all).
LLM_PROVIDER=groq
LLM_MODEL=llama-3.3-70b-versatile
LLM_API_KEY=gsk_...

Then:

cd docker-compose
docker compose --env-file ../.env up --build

Open http://localhost:4200 and register. The first account created becomes the admin.

The first run downloads the embedding and layout models (~1–2 GB) into a Docker volume, so it takes a few minutes. Later runs start immediately.

Full instructions running from source, S3/MinIO, OIDC, SMTP, troubleshooting are in docs/SETUP.md.

Zero-cost setup, no API key at all

ollama pull qwen3:8b
LLM_PROVIDER=ollama
LLM_MODEL=qwen3:8b

Everything then runs on your machine, offline. It is slower than a hosted API -- expect tens of seconds per question on CPU -- so raise LLM_ANSWER_TIMEOUT_SECONDS.


What you can plug in

Set one variable per layer. Bold is the default.

Language model -- LLM_PROVIDER

Value Talks to Cost
openai_compatible Any POST /chat/completions endpoint - corporate gateways, OpenRouter, Together, DeepInfra, Fireworks, Cerebras, vLLM, LM Studio, llama.cpp. Set LLM_BASE_URL. varies
openai api.openai.com paid
gemini Google Generative Language API free tier
groq api.groq.com free tier, very fast
huggingface HF inference router free tier
ollama Local daemon free, offline
huggingface_local transformers in-process free, offline
bedrock AWS Bedrock -- the original stack paid

LLM_QUERY_MODEL can point query expansion at a cheaper model than answering. Qwen-family models have visible reasoning disabled automatically, because the extraction prompts require JSON-only replies.

Embeddings -- EMBEDDING_PROVIDER

Value Uses Cost
local sentence-transformers, default BAAI/bge-small-en-v1.5 (384-dim, ~130 MB) free, offline
ollama Local Ollama daemon, e.g. nomic-embed-text free, offline
openai / openai_compatible Any /embeddings endpoint paid per chunk
gemini Google embeddings free tier
bedrock Titan paid

EMBEDDING_DIM is not just a setting. pgvector fixes the column width when the tables are created. Switching model families later means dropping the embeddings and hierarchy_chunks tables and re-ingesting every document. Decide before your first upload.

Document extraction -- EXTRACTION_BACKEND

Value Uses Cost
docling Local layout analysis, table-structure recognition, OCR free, offline
textract Amazon Textract; needs STORAGE_BACKEND=s3 paid

Textract keeps two advantages: it detects signatures, and it reports per-element confidence. Docling does neither. Everything else is equivalent -- see How extraction stays swappable.

File storage -- STORAGE_BACKEND

Value Uses Cost
local Files under STORAGE_LOCAL_PATH, on a Docker volume free
s3 AWS S3, or self-hosted MinIO / R2 / B2 / Wasabi / Ceph free with MinIO

Try the S3 path with no cloud account: docker compose --profile s3 up.

Authentication -- AUTH_PROVIDER

Value Behaviour
local Email + password, bcrypt-hashed, JWT tokens. No provider needed.
oidc Validates RS256 tokens against any issuer's JWKS -- Keycloak, Google, Auth0, Entra ID.
dev Every request is a fixed dev user. Never expose this.

Notifications -- NOTIFICATION_CHANNEL

none · log · smtp · webhook. For SMTP without a paid provider: docker compose --profile mail up, then read the mail at http://localhost:8025.


Chunking and retrieval strategies

RAG quality is mostly decided by two choices: how a document is cut up, and how passages are found again. Both are switchable in config.yaml at the repo root, with no code changes.

chunking:
  strategy: legacy      # legacy | recursive | semantic
retrieval:
  strategy: legacy      # legacy | bm25 | hybrid_rrf | rerank | hyde | hyde_rerank
  top_k: 15
Chunking What it does
legacy Layout-aware: chunk boundaries follow the document's own headings and sections.
recursive Fixed-size character splitting with overlap, falling back through paragraph → line → sentence → word.
semantic Splits where consecutive passages stop being similar, measured with the same embedding model used for retrieval.
Retrieval What it does
legacy Dense vector search per refined query, then dedup.
bm25 Okapi BM25 keyword search.
hybrid_rrf Dense + BM25 fused by Reciprocal Rank Fusion.
rerank Dense retrieval, then a cross-encoder reorders the candidates.
hyde The LLM writes a hypothetical answer; that gets embedded and searched with.
hyde_rerank Both.

Tables, images, signatures and form groups stay atomic under every chunking strategy splitting a rate table by character count destroys it. Only narrative text is affected.

Chunking applies at ingest (re-upload to change it); retrieval applies at query time (effective on the next question). legacy on both keys reproduces exactly the behaviour from before strategies existed, so an absent or empty file changes nothing.

# after editing config.yaml
cd docker-compose && docker compose --env-file ../.env restart ai_service

Every option, every parameter, and the measured results of running each strategy through the real pipeline: WORKING.md.


Architecture

                       ┌──────────────┐
  browser ────────────▶│   frontend   │  Angular 19 + @siemens/element-ng (MIT)
                       │    :4200     │
                       └──────┬───────┘
                              │ REST + JWT
                       ┌──────▼───────┐
                       │   backend    │  FastAPI: projects, RBAC, auth, uploads
                       │    :9090     │
                       └──┬────────┬──┘
                          │        │
             ┌────────────▼──┐  ┌──▼──────────────┐
             │  ai-service   │  │  doc-conversion │  LibreOffice headless
             │    :9091      │  │      :3003      │  .docx ──▶ .pdf
             └───┬───────┬───┘  └─────────────────┘
                 │       │
   ┌─────────────▼─┐  ┌──▼────────┐   ┌──────────────────┐
   │   postgres    │  │   redis   │   │  documents volume│
   │  + pgvector   │  │  celery   │   │  (or S3/MinIO)   │
   └───────────────┘  └───────────┘   └──────────────────┘

Database setup is split deliberately: sql_scripts/init/ runs in Postgres' initdb hook (the vector extension only nothing that needs application tables, because they do not exist yet and a failing init script kills the container), while sql_scripts/seed/ is applied by the backend at startup, after SQLAlchemy has created the tables.

The pipeline, end to end:

  1. Upload -- the browser POSTs to the backend (local storage), or PUTs to a presigned URL (S3). GET /storage/capabilities tells the client which.
  2. Extract -- Docling produces layout blocks, table structure and OCR text.
  3. Chunk -- hierarchical chunks follow the document's heading structure; tables, figures and form fields stay atomic.
  4. Embed -- chunks go into Postgres with an HNSW index over cosine distance.
  5. Retrieve -- the query is expanded and refined by an LLM, then matched against both the chunk index and a heading/taxonomy index, with a lower-threshold fallback when a query returns too little.
  6. Answer -- a per-question extraction prompt runs against the retrieved context and must return JSON, with [i] markers citing source lines.

How extraction stays swappable

The chunking pipeline is ~2,700 lines written against Amazon Textract's block graph -- flat Blocks with Id, BlockType, Page, Geometry and Relationships edges. That is good, provider-neutral document understanding, and rewriting it would have meant re-tuning extraction quality from scratch.

So it was not rewritten. Instead every backend emits the same block graph: Textract's output passes through untouched, and Docling's document model is translated into that shape in bidrag/core/extraction/blocks.py. Downstream code cannot tell the difference.

That contract is pinned by tests, because a drift in block shape would not crash anything, it would silently stop finding tables or headings:

cd bidrag-ai-service && python -m pytest tests/core -q     # 81 tests
cd backend && python -m pytest tests/utils/test_auth.py -q # 34 tests

Verified end to end

The full stack has been run and checked, not just built:

Step Result
Postgres + pgvector via the initdb hook extension vector 0.8.6 installed
Backend startup seeding RBAC applied, 33 extraction questions seeded
Vector schema from EMBEDDING_DIM vector(384) on both tables, HNSW indexes built
Register / login / refresh works; refresh token correctly rejected as an access token; identical error for unknown-email and wrong-password
First registered account promoted to admin in both the RBAC tables and the legacy column
LLM provider (OpenAI-compatible gateway) live call returned {"ok": true}
Local embeddings (bge-small-en-v1.5) 384-dim; query/document prefix asymmetry gives 0.75 vs 0.38 cosine on a relevant/irrelevant pair
Upload on local storage multipart POST /upload; presigned path correctly returns 409
Docling PDF ingestion COMPLETED -- 6 hierarchy chunks, 5 embeddings, table preserved
Extraction over 33 questions ANALYSIS_COMPLETE; 13 substantive answers with citations, 20 correctly "Not specified"
Angular build dev 7.96 MB, prod 3.35 MB / 648 kB gzipped

Extracted values were checked against the source document -- customer names, payment schedule (20/60/20, net 45), notice periods (90/30 days), liquidated damages (0.5%/week capped at 10%), and switchgear ratings read out of a table.


Running without Docker

You need Postgres 16+ with pgvector, and Redis.

# Backend
cd backend && pip install -e '.[dev]' && python start_local.py

# AI service + workers
cd bidrag-ai-service && pip install -e '.[dev]' && python start_local.py
python worker_launcher.py

# Converter (needs libreoffice on PATH)
cd doc-conversion && pip install -e '.[dev]' && uvicorn src.main:app --port 3003

# Frontend
cd frontend && npm install && npm start

Optional extras, installed only if you want them:

pip install -e '.[aws]'       # S3 storage, Bedrock, Textract
pip install -e '.[hf-local]'  # in-process transformers inference
pip install -e '.[tracing]'   # Datadog APM

Configuration

Deployment settings -- where is the database, which model, which bucket -- are environment variables, resolved in this order:

  1. Process environment / .env
  2. The [$APP_ENV] section of each service's config.properties (non-secret service URLs only)
  3. Built-in defaults

Pipeline settings -- which chunking and retrieval algorithm -- live separately in config.yaml, because they are a different kind of decision. See Chunking and retrieval strategies.

.env.example documents every option. Secrets belong only in .env, which is git-ignored.

On startup the AI service logs its resolved providers - model, embedding dimension, extraction backend, storage backend with no secrets. That log line is the fastest way to diagnose "why is it using the wrong model".

The prompts are the product

sql_scripts/seed/extraction-prompts.sql seeds the question set: direct customer, end customer, payment terms, termination conditions, technical scope, and so on. Each row carries an extraction_prompt with a search strategy, extraction rules and a required JSON output shape.

This is where extraction quality lives. Adapt these to your domain and the system follows; leave them generic and results will be generic. sql_scripts/prompt-history/ keeps the original tuning iterations as a record of what was tried.


Tracing and observability

A RAG pipeline is a black box until you can see inside it -- so watch it not be one. BidRAG ships optional tracing to Langfuse, self-hosted from their published Docker images -- no Langfuse Cloud account, no data leaving your machine, and no source vendored into this repository. Each ingestion and each answered question becomes a readable trace:

  • Ingestion -- what a document was cut into, every chunk with its full text and a stable reference (C-12-0007), plus what was embedded.
  • Answering -- question → expanded query → hierarchy chunks → generated refined queries → the chunks each one retrieved with similarity scores → deduplication → the exact prompt → the answer.
  • Attribution -- which chunks the answer actually cited, which were shown and ignored, and which never reached the model.
  • Cost -- token usage per call, priced per model, filterable by project or question.

Off by default, and never on the critical path: with LANGFUSE_ENABLED=false, or Langfuse simply not running, the service behaves exactly as it does without it.

cd observability
Copy-Item .env.langfuse.example .env.langfuse   # then fill in the secrets
docker compose -f docker-compose.langfuse.yml --env-file .env.langfuse up -d
python seed_langfuse_models.py                  # register model pricing

Then set LANGFUSE_ENABLED=true plus the key pair in .env and restart ai_service. Full walkthrough, port map and troubleshooting: observability/OBSERVABILITY.md.

Verified against Langfuse server 4.1.0 with Python SDK >=3.0,<4.0. The image tags in observability/docker-compose.langfuse.yml are what decide the version you actually get; see NOTICE.md for why the SDK and server majors differ.


Security

  • No secret is ever committed; .gitignore covers .env and friends.
  • Nothing secret reaches the browser bundle. The old build shipped an OAuth client secret and an analytics token to every visitor; both are gone.
  • Storage keys are normalised and path-traversal is rejected before any filesystem access, since keys derive from client filenames.
  • Refresh tokens carry a typ claim and cannot be used as access tokens.
  • Login returns one identical error for unknown-email and wrong-password, so it cannot be used to enumerate accounts.
  • The app refuses to start if JWT_SECRET_KEY is a placeholder or under 32 characters.

Before exposing an instance publicly: set AUTH_ALLOW_REGISTRATION=false once your users exist, change POSTGRES_PASSWORD, put TLS in front, and narrow the CORS origins in each service's main.py (they currently allow *).


Contributing

Issues and pull requests are welcome. Please run the tests and keep the block contract green. sql_scripts/prompt-history/ shows the level of prompt-tuning detail this project values.

License

Apache-2.0 -- see LICENSE, and read NOTICE.md first.

BidRAG is not affiliated with, sponsored by, or endorsed by Siemens AG. The Angular UI is built on Siemens Element, a publicly released, MIT-licensed component library installed from npm like any other dependency. The MIT license requires that attribution be kept.

About

Open-source RAG extraction for bidding and tender documents. Local-first: Docling layout analysis, pgvector, local embeddings. Bring one LLM key, or none with Ollama.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages