An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries
Developed for the Smart India Hackathon 2026
Remote-sensing imagery is widely used for agricultural monitoring, disaster management, urban planning, forest monitoring, water-resource assessment, infrastructure mapping, and environmental analysis. However, most existing remote-sensing AI solutions are developed as isolated applications for a single predefined task, such as land-cover classification, object detection, visual question answering, or change detection.
These systems often require users to understand satellite-data characteristics, GIS workflows, model selection, and task-specific parameters. Consequently, non-expert users may find it difficult to obtain meaningful information from satellite imagery through simple natural-language queries.
SatQuery AI addresses this by providing an agentic, query-driven framework that automatically selects and executes suitable remote-sensing specialist models, validates inputs, combines outputs, and returns evidence-grounded responses.
SatQuery AI is a software-based agentic vision-language assistant for analysing single and paired remote-sensing images through natural-language queries. Single-image understanding is a mandatory baseline, while the principal focus is joint reasoning over paired cross-modal and multitemporal imagery.
| Input Type | Description |
|---|---|
| Single Image | One optical/multispectral or SAR image for captioning, VQA, and text-guided region grounding |
| Cross-Modal Pair | Co-registered optical/multispectral and SAR images for joint information extraction |
| Bi-Temporal Pair | Two spatially corresponding images from different times for change detection and change-based VQA |
| Supported Formats | GeoTIFF (.tif), TIFF (.tiff). PNG and JPEG accepted for prescribed public benchmark datasets |
| Capability | Description | Status |
|---|---|---|
| Remote-Sensing Adaptation | Vision-language model fine-tuned using BigEarthNet.txt with LoRA for domain adaptation | ✅ |
| Single-Image VQA | Visual question answering on optical or SAR satellite imagery | ✅ |
| Single-Image Captioning | Automated scene description and land-cover captioning | ✅ |
| Bi-Temporal Change Analysis | Change description and change-based VQA from bi-temporal image pairs | ✅ |
| Cross-Modal Pair Analysis | Joint information extraction from co-registered optical–SAR image pairs | ✅ |
| Agentic Orchestration | Automatic task classification, model selection, execution sequencing, and output integration via LangGraph | ✅ |
| Interactive GUI | Analysis console: GeoTIFF upload with in-browser validation, grounded evidence on the image and a real map, optical–SAR fusion workspace, change compare/timeline, charts, PDF / GeoJSON / JSON reports | ✅ |
| Execution Summaries | Auditable execution trace with selected task, model/tool names, and key parameters | ✅ |
- "Describe the land-cover and major objects visible in this image."
- "What is the primary land cover shown in this image?"
- "What changed between these two dates, and where did the change occur?"
- "Use the optical and SAR images together to identify built-up and water-covered regions."
- "Has the built-up area increased, decreased, or remained unchanged?"
┌────────────────────────────────────────────────────────────────┐
│ SatQuery AI │
├────────────────┬───────────────────────────────────────────────┤
│ Frontend │ Next.js 16 · TypeScript · MapLibre GL │
│ (Vercel) │ geotiff.js · jsPDF · SVG charts │
├────────────────┼───────────────────────────────────────────────┤
│ API Layer │ FastAPI · REST · CORS · File Upload │
├────────────────┼───────────────────────────────────────────────┤
│ Agentic │ LangGraph StateGraph │
│ Controller │ Router → Executor → Response │
│ │ Task classification · Model selection │
│ │ Input validation · Execution trace │
├────────────────┼───────────────────────────────────────────────┤
│ Specialist │ Single-Image VQA (Qwen2-VL + LoRA) │
│ Tools │ Single-Image Captioning (Qwen2-VL + LoRA) │
│ │ Cross-Modal VQA (Optical + SAR) │
│ │ Bi-Temporal Change VQA │
├────────────────┼───────────────────────────────────────────────┤
│ Data Layer │ Rasterio GeoTIFF processing │
│ │ Sentinel-2 RGB · Sentinel-1 VV/VH │
│ │ Normalization and tensor conversion │
└────────────────┴───────────────────────────────────────────────┘
| Component | Technology |
|---|---|
| Vision-Language Model | Qwen2-VL-2B-Instruct |
| Fine-Tuning | LoRA (PEFT) — domain adaptation on BigEarthNet.txt |
| Agent Framework | LangGraph (StateGraph) |
| Backend | Python · FastAPI · PyTorch · Transformers |
| GeoTIFF Processing | Rasterio · NumPy |
| Frontend | Next.js 16 · TypeScript · React 19 · Tailwind CSS 4 |
| Maps & rasters | MapLibre GL (Esri World Imagery basemap) · geotiff.js (in-browser GeoTIFF parsing) |
| Reports | jsPDF (PDF) · GeoJSON · JSON execution trace |
| Evaluation Metrics | BLEU-1→4 · ROUGE-L · Exact Match |
| Benchmarks | VRSBench · RSVQA · CDVQA |
| Dataset | Purpose | Reference |
|---|---|---|
| BigEarthNet.txt | Primary dataset for remote-sensing adaptation using co-registered Sentinel-1 SAR, Sentinel-2 multispectral imagery, and diverse text annotations | arxiv.org/abs/2603.29630 |
| VRSBench | Evaluation benchmark for single-image captioning, grounding, and VQA | Public benchmark |
| RSVQA | Evaluation benchmark for single-image visual question answering | Public benchmark |
| CDVQA | Evaluation benchmark for multi-temporal change-based VQA | Public benchmark |
- Python 3.10+
- Node.js 18+
- ~4GB disk space (for base VLM model download)
- GPU recommended (CUDA or Apple MPS) but CPU works
# Clone the repository
git clone https://github.com/your-username/SatQueryAI.git
cd SatQueryAI
# Create Python environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txtcd frontend
npm installStart both the FastAPI backend and the Next.js frontend:
Terminal 1 — AI Backend:
source venv/bin/activate
uvicorn src.api.main:app --reload --port 8050Terminal 2 — Frontend:
cd frontend
npm run devOpen http://localhost:3000 in your browser.
The frontend also runs on its own: the analysis console ships five real Sentinel-1/2 test scenes (frontend/public/demo/samples/) with cached specialist outputs, so every capability can be demonstrated offline. Set NEXT_PUBLIC_API_URL in frontend/.env.local to show an Offline / Live switch that sends queries to the FastAPI backend.
Recording the demo video: follow frontend/DEMO_SCRIPT.md and use npm run build && npm start for a clean production build.
frontend/scripts/build_demo_assets.py rebuilds every test scene from open Copernicus data on Microsoft Planetary Computer (no account needed):
| Scene | Data | Used for |
|---|---|---|
| Hyderabad — Hussain Sagar | Sentinel-2 L2A, 07 Jan 2025 | Captioning, VQA, text-guided grounding |
| Mumbai — Bandra · Kurla · CSMIA | Sentinel-2 L2A (06 Jan 2025) + Sentinel-1 RTC VV/VH (07 Jan 2025, 06:33 IST) | SAR VQA, optical–SAR fusion |
| Navi Mumbai International Airport | Sentinel-2 L2A, 03 Jan 2017 → 16 Jan 2026 + 10 yearly epochs | Change description, change-VQA, trend |
It writes GeoTIFF samples, preview/overlay images and the evidence JSON (areas, polygons, histograms, time series) that the console's answers are built from, using spectral indices, Otsu SAR thresholding, rule-based optical–SAR fusion and post-classification change detection.
cd frontend
npm run demo:assets # = ../venv/bin/python scripts/build_demo_assets.pypython scripts/setup_bigearthnet_data.py --num_samples 500This downloads remote-sensing imagery, converts it to Sentinel-compatible GeoTIFF format (BGR, 0-4000 reflectance), and generates VQA training pairs.
python scripts/train_lora.py --epochs 3 --batch_size 4 --accumulation_steps 4Features:
- Automatic Mixed Precision (AMP) training
- Cosine annealing learning rate scheduler
- Gradient accumulation for effective larger batch sizes
- Automatic checkpointing to
data/processed/lora_weights/
# Individual benchmarks
python scripts/evaluate.py --benchmark vrsbench --max_samples 100
python scripts/evaluate.py --benchmark rsvqa --variant LR --max_samples 100
python scripts/evaluate.py --benchmark cdvqa --max_samples 50
# All benchmarks
python scripts/evaluate.py --benchmark all --max_samples 100Metrics computed: Exact Match Accuracy, BLEU-1 through BLEU-4, ROUGE-L (F1), per-type accuracy breakdown.
SatQueryAI/
├── src/
│ ├── api/
│ │ └── main.py # FastAPI server with /analyze endpoint
│ ├── agent/
│ │ ├── orchestrator.py # LangGraph agentic controller
│ │ └── tools_registry.py # Specialist AI tools (VQA, Captioning, Change, Cross-Modal)
│ ├── models/
│ │ └── vlm_manager.py # Qwen2-VL base model + LoRA adapter injection
│ └── data_prep/
│ ├── dataset.py # PyTorch Dataset for S1+S2+text triplets
│ └── geotiff_loader.py # Rasterio GeoTIFF → normalized tensor
├── scripts/
│ ├── setup_bigearthnet_data.py # BigEarthNet.txt data preparation
│ ├── setup_dev_data.py # Development data generation
│ ├── train_lora.py # LoRA fine-tuning pipeline
│ ├── evaluate.py # Benchmark evaluation runner
│ ├── metrics.py # VQA metrics (BLEU, ROUGE-L, EM)
│ └── benchmark_loaders/ # VRSBench, RSVQA, CDVQA dataset loaders
├── frontend/
│ ├── DEMO_SCRIPT.md # Shot list for the demo video
│ ├── scripts/
│ │ ├── build_demo_assets.py # Real Sentinel-1/2 test scenes + evidence JSON
│ │ └── copy-maplibre-worker.mjs
│ ├── public/demo/ # Sample GeoTIFFs, previews, overlays
│ └── src/
│ ├── app/
│ │ ├── page.tsx # Overview (landing) page
│ │ ├── analysis/page.tsx # Analysis console
│ │ └── api-docs/page.tsx # API reference + playground
│ ├── components/
│ │ ├── console/ # Inputs, canvas views, agent panel, trace
│ │ ├── charts/ # SVG charts (dataviz-validated palette)
│ │ ├── map/ # MapLibre real-map component
│ │ ├── home/ api/ site/
│ └── lib/
│ ├── engine/ # Intent router, tool registry, agent run loop
│ ├── geo/ # GeoTIFF parsing, UTM, pair compatibility
│ ├── demo/ # Scenes, cached outputs, generated evidence
│ └── report/ # PDF / GeoJSON / JSON exports
├── data/
│ ├── raw/ # Sentinel-1 & Sentinel-2 GeoTIFFs
│ ├── processed/ # Training JSON & LoRA weights
│ └── test_samples/ # Sample imagery for testing
├── test_samples/ # Quick-test GeoTIFFs
├── requirements.txt
└── README.md
Final evaluation uses prescribed public benchmark test subsets and an ISRO/SAC evaluation dataset. Scores are normalised before combining different metrics.
| Benchmark | Task | Metrics |
|---|---|---|
| VRSBench | Single-image captioning, grounding, VQA | BLEU, ROUGE-L, EM |
| RSVQA | Single-image visual question answering | Exact Match Accuracy |
| CDVQA | Multi-temporal change-based VQA | Exact Match Accuracy, BLEU |
| ISRO/SAC | Cartosat-2S + RISAT image pairs | Task-specific (undisclosed) |
- Frontend: Deployed to Vercel (static + SSR)
- Backend: Local FastAPI server with GPU inference
For Vercel deployment:
cd frontend
npx vercel --prodSet the environment variable NEXT_PUBLIC_API_URL in the Vercel dashboard to point to your backend server.
This project was developed for the Smart India Hackathon under the Indian Space Research Organisation (ISRO) problem statement.
Built with ❤️ for advancing remote-sensing intelligence through natural-language interaction.