Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
53 changes: 41 additions & 12 deletions REQUIREMENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

> **Note:** This document is automatically generated and verified against the live test suite by `scripts/generate_requirements.py` and `tests/backend/test_requirements_sync.py`.

**Test Verification Baseline:** **1003 Automated Tests** (658 Pytest Backend + 295 Vitest Frontend + 50 Playwright E2E).
**Test Verification Baseline:** **1032 Automated Tests** (679 Pytest Backend + 303 Vitest Frontend + 50 Playwright E2E).

---

Expand Down Expand Up @@ -460,7 +460,7 @@ classDiagram
- `test_split_by_length`
- `test_chunk_markdown`

#### `tests/backend/test_chunker_languages.py` (16 tests)
#### `tests/backend/test_chunker_languages.py` (18 tests)
- `test_language_detection`
- `test_get_tree_sitter_parser_caching_and_fallbacks`
- `test_extract_symbols_unsupported_language`
Expand All @@ -477,6 +477,8 @@ classDiagram
- `test_get_file_outline_helper`
- `test_markdown_chunking_with_subchunks`
- `test_markdown_chunking_with_nested_headings_and_empty`
- `test_call_extraction_constructors_and_generics`
- `test_toplevel_call_source_symbol_preservation`

#### `tests/backend/test_db_and_tools.py` (17 tests)
- `test_db_path_and_init`
Expand Down Expand Up @@ -626,13 +628,16 @@ classDiagram
- `test_sync_single_git_repo_vector_upsert_failure`
- `test_sync_local_paths_vector_upsert_failure`

#### `tests/backend/test_mcp_v2.py` (6 tests)
#### `tests/backend/test_mcp_v2.py` (9 tests)
- `test_fastmcp_tools_registered`
- `test_fastmcp_resources_and_prompts`
- `test_fastmcp_tool_execution`
- `test_fastmcp_resource_read`
- `test_fastmcp_prompt_get`
- `test_fastmcp_streamable_http_transport`
- `test_search_code_with_dense_weight_and_score_breakdown`
- `test_search_code_empty_and_missing_ast_boundaries`
- `test_search_docs_with_dense_weight_and_score_breakdown`

#### `tests/backend/test_multi_git_providers.py` (9 tests)
- `test_detect_git_provider`
Expand All @@ -657,7 +662,7 @@ classDiagram
- `test_api_get_omni_search_symbols_and_files`
- `test_navigator_tree_has_no_empty_folder_root` - _Verify get_navigator_tree sanitizes URI schemes and never produces empty name root folders._

#### `tests/backend/test_navigator_service.py` (11 tests)
#### `tests/backend/test_navigator_service.py` (12 tests)
- `test_db`
- `test_navigator_tree_construction`
- `test_navigator_tree_all_repos`
Expand All @@ -668,18 +673,26 @@ classDiagram
- `test_symbol_impact_retrieval`
- `test_symbol_impact_not_found`
- `test_no_outgoing_calls_in_callers`
- `test_class_symbol_impact_aggregation`
- `test_real_codebase_symbol_extraction_and_navigation`

#### `tests/backend/test_schemas.py` (2 tests)
#### `tests/backend/test_schemas.py` (4 tests)
- `test_code_symbol_creation`
- `test_search_request_defaults`
- `test_search_request_dense_weight_boundaries`
- `test_search_request_search_mode_boundaries`

#### `tests/backend/test_search.py` (4 tests)
#### `tests/backend/test_search.py` (9 tests)
- `test_execute_hybrid_search_empty_query`
- `test_execute_hybrid_search_delegation`
- `test_execute_hybrid_search_ast_enrichment`
- `test_execute_hybrid_search_ast_enrichment_interval_fallback`
- `test_execute_hybrid_search_ast_enrichment_resilience`
- `test_execute_hybrid_search_ast_enrichment_skipped_for_docs`
- `test_execute_hybrid_search_exception`
- `test_execute_hybrid_search_end_to_end_real` - _Validates REAL hybrid retrieval without mocking get_vector_store or execute_hybrid_search.
Asserts both Markdown and PDF docs are matched under doc_type='doc'._
- `test_api_test_search_endpoint`

#### `tests/backend/test_tools.py` (3 tests)
- `test_dynamic_catalog_description`
Expand Down Expand Up @@ -793,7 +806,7 @@ Asserts both Markdown and PDF docs are matched under doc_type='doc'._
- `test_test_connection_active_embedded` - _Verify test_connection succeeds on active embedded Qdrant store without file lock conflict._
- `test_switch_same_embedded_directory` - _Verify switch_vector_store succeeds when switching collection on the same embedded Qdrant directory._

#### `tests/backend/test_vector_store_qdrant.py` (32 tests)
#### `tests/backend/test_vector_store_qdrant.py` (40 tests)
- `TestQdrantVectorStoreInit::test_init_in_memory_or_embedded`
- `TestQdrantVectorStoreInit::test_init_remote_success`
- `TestQdrantVectorStoreInit::test_init_remote_fallback_to_embedded_on_connection_error`
Expand All @@ -806,6 +819,10 @@ Asserts both Markdown and PDF docs are matched under doc_type='doc'._
- `TestQdrantVectorStoreOperations::test_search_dense_and_hybrid_rrf`
- `TestQdrantVectorStoreOperations::test_search_weighted_score_fusion_range_and_boost`
- `TestQdrantVectorStoreOperations::test_search_weighted_score_fusion_alpha_weighting`
- `TestQdrantVectorStoreOperations::test_search_explicit_dense_weight_and_score_decomposition`
- `TestQdrantVectorStoreOperations::test_search_modes_case_insensitivity_and_whitespace`
- `TestQdrantVectorStoreOperations::test_search_modes_semantic_and_lexical`
- `TestQdrantVectorStoreOperations::test_search_mode_lexical_empty_sparse_no_dense_fallback`
- `TestQdrantVectorStoreOperations::test_search_dense_fallback_without_sparse`
- `TestQdrantVectorStoreOperations::test_delete_by_path`
- `TestQdrantVectorStoreOperations::test_delete_by_repo`
Expand All @@ -822,6 +839,10 @@ Asserts both Markdown and PDF docs are matched under doc_type='doc'._
- `test_search_dense_and_hybrid_rrf`
- `test_search_weighted_score_fusion_range_and_boost`
- `test_search_weighted_score_fusion_alpha_weighting`
- `test_search_explicit_dense_weight_and_score_decomposition`
- `test_search_modes_case_insensitivity_and_whitespace`
- `test_search_modes_semantic_and_lexical`
- `test_search_mode_lexical_empty_sparse_no_dense_fallback`
- `test_search_dense_fallback_without_sparse`
- `test_delete_by_path`
- `test_delete_by_repo`
Expand Down Expand Up @@ -1203,12 +1224,13 @@ and leaves the prior indexed state intact without data loss._
- renders vector database health badge in header when vector_db_status is present
- renders ChromaDB provider and unhealthy status badge in header

#### `CodeNavigator.test.tsx` (5 tests)
#### `CodeNavigator.test.tsx` (6 tests)
- renders toolbar, hero layout, and fetches initial tree data
- handles density mode switching and persists to localStorage
- loads file outline on file selection and symbol impact on symbol selection
- supports caller click-through navigation jumping to caller file and symbol
- handles repo switcher change and re-fetches tree
- handles external navigation and permits subsequent repo changes without loop

#### `DiagnosticsViewer.test.tsx` (10 tests)
- renders log records, badges, and controls
Expand Down Expand Up @@ -1307,7 +1329,7 @@ and leaves the prior indexed state intact without data loss._
- calls onSelectCallee when a clickable callee is clicked for cross-file navigation
- renders loading state when loading is true

#### `NavigatorOmniSearch.test.tsx` (8 tests)
#### `NavigatorOmniSearch.test.tsx` (10 tests)
- renders omni-search input with placeholder
- fetches matches when user types and displays floating overlay
- navigates with keyboard and selects on Enter
Expand All @@ -1316,6 +1338,8 @@ and leaves the prior indexed state intact without data loss._
- handles fetch error gracefully without crashing
- closes dropdown when clicking outside the container
- supports ArrowUp navigation within bounds
- renders container prefix and highlights query match in symbol and path
- focuses search input when pressing Ctrl+K

#### `NavigatorOutline.test.tsx` (9 tests)
- renders empty placeholder when outline is null or empty
Expand Down Expand Up @@ -1405,13 +1429,18 @@ and leaves the prior indexed state intact without data loss._
- cancels sync by calling /admin/api/repos/{id}/cancel-sync
- closes EventSource on unmount

#### `SearchInspector.test.tsx` (6 tests)
#### `SearchInspector.test.tsx` (11 tests)
- renders initial prompt and inputs
- performs search and renders matching hit cards
- switches search mode and updates UI controls
- adjusts hybrid split slider and presets
- performs search and renders matching hit cards with score breakdowns and signature
- invokes onOpenInNavigator when Open in Navigator button is clicked
- displays empty results message when no hits found
- handles search API failure with error display
- performs doc search with repo filter and renders documentation hits
- renders search query form and hit card headers with responsive classes
- copies code snippet when Copy button is clicked
- renders hits with missing AST metadata and null scores without error
- handles network error during search gracefully

#### `Settings.test.tsx` (29 tests)
- renders vector database panel, auto-sync panel, multi-provider token boxes, rate limits, and host vault list
Expand Down
20 changes: 18 additions & 2 deletions app/api/routers/repositories.py
Original file line number Diff line number Diff line change
Expand Up @@ -293,19 +293,35 @@ async def api_test_search(payload: SearchRequest):
if not query:
return JSONResponse(status_code=400, content={"error": "Query required"})

search_mode = payload.search_mode or "hybrid"
hits = search_service.execute_hybrid_search(
query_text=query,
doc_type=payload.type,
repo=payload.repo,
limit=payload.limit or 6
language=payload.language,
category=payload.category,
tag=payload.tag,
limit=payload.limit or 6,
dense_weight=payload.dense_weight,
search_mode=search_mode
)
results = []
for h in hits:
results.append({
"score": round(getattr(h, "score", 0.0), 4),
"dense_score": getattr(h, "dense_score", None),
"sparse_score": getattr(h, "sparse_score", None),
"dense_rank": getattr(h, "dense_rank", None),
"sparse_rank": getattr(h, "sparse_rank", None),
"payload": getattr(h, "payload", {})
})
return {"query": query, "type": payload.type, "results": results}
return {
"query": query,
"type": payload.type,
"search_mode": search_mode,
"dense_weight": payload.dense_weight,
"results": results
}
except Exception as e:
logger.error(f"Error testing search: {e}")
return JSONResponse(status_code=500, content={"error": "Failed to execute search test."})
Expand Down
72 changes: 58 additions & 14 deletions app/mcp/handlers/search_handlers.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,30 +23,56 @@ async def handle_search_code(
query: Annotated[str, Field(description="Natural language question or code concept (e.g. 'JWT token authentication handler').")],
repo: Annotated[Optional[str], Field(description="Optional repository name/alias to filter by.")] = None,
language: Annotated[Optional[str], Field(description="Optional language filter (e.g. 'python', 'typescript', 'go').")] = None,
limit: Annotated[int, Field(description="Max number of code blocks to return (default 5).")] = 5
limit: Annotated[int, Field(description="Max number of code blocks to return (default 5).")] = 5,
dense_weight: Annotated[Optional[float], Field(description="Weight between 0.0 (pure lexical BM25) and 1.0 (pure semantic vector). Default is 0.5 balanced.")] = None,
mode: Annotated[Optional[str], Field(description="Search mode: 'hybrid' (default), 'semantic', or 'lexical'.")] = "hybrid"
) -> str:
"""Hybrid semantic and BM25 search over code functions, classes, and logic snippets with line numbers and GitHub links."""
"""Hybrid semantic and BM25 search over code functions, classes, and logic snippets with line numbers, AST signatures, and GitHub links."""
query = query.strip() if query else ""
if not query:
return "Error: search query cannot be empty."

try:
from app.services.auth import enforce_tool_permission, Role
enforce_tool_permission(Role.VIEWER)
hits = _get_tools_attr("execute_hybrid_search", execute_hybrid_search)(query_text=query, doc_type="code", repo=repo, language=language, limit=limit)
hits = _get_tools_attr("execute_hybrid_search", execute_hybrid_search)(
query_text=query,
doc_type="code",
repo=repo,
language=language,
limit=limit,
dense_weight=dense_weight,
search_mode=mode or "hybrid"
)
if not hits:
return f"No matching code snippets found for query: '{query}'."

formatted = []
for hit in hits:
p = hit.payload
header = f"### [{p.get('repo')}] {p.get('rel_path')} (Lines {p.get('start_line')}-{p.get('end_line')})"
if p.get("symbol"):
header += f" - Symbol: `{p.get('symbol')}`"
header = f"### [{p.get('repo')}] {p.get('rel_path')} (Lines {p.get('start_line')}-{p.get('end_line')})\n"

sym = p.get("full_symbol") or p.get("symbol")
kind = p.get("kind")
if sym:
kind_str = f" (`{kind}`)" if kind else ""
header += f"- **Symbol**: `{sym}`{kind_str}\n"

sig = p.get("signature")
if sig:
header += f"- **Signature**: `{sig}`\n"

link_url = p.get("permalink_url") or p.get("github_url")
if link_url:
header += f"\nSource Link: {link_url}"
header += f"\nRelevance Score: {hit.score:.4f} ({hit.score * 100:.1f}%)\n"
header += f"- **Source Link**: {link_url}\n"

score_val = float(hit.score) if isinstance(getattr(hit, "score", None), (int, float)) else 0.0
score_str = f"{score_val:.4f} ({score_val * 100:.1f}%)"
d_val = getattr(hit, "dense_score", None)
s_val = getattr(hit, "sparse_score", None)
if isinstance(d_val, (int, float)) and isinstance(s_val, (int, float)):
score_str += f" [Semantic: {float(d_val) * 100:.1f}% | Lexical: {float(s_val) * 100:.1f}%]"
header += f"- Relevance Score: {score_str}\n\n"

lang = p.get("language", "")
block = f"{header}```{lang}\n{p.get('content')}\n```"
Expand All @@ -63,7 +89,9 @@ async def handle_search_docs(
repo: Annotated[Optional[str], Field(description="Optional repository/vault filter.")] = None,
category: Annotated[Optional[str], Field(description="Optional category filter.")] = None,
tag: Annotated[Optional[str], Field(description="Optional tag filter.")] = None,
limit: Annotated[int, Field(description="Max documents to return (default 5).")] = 5
limit: Annotated[int, Field(description="Max documents to return (default 5).")] = 5,
dense_weight: Annotated[Optional[float], Field(description="Weight between 0.0 (pure lexical BM25) and 1.0 (pure semantic vector). Default is 0.5 balanced.")] = None,
mode: Annotated[Optional[str], Field(description="Search mode: 'hybrid' (default), 'semantic', or 'lexical'.")] = "hybrid"
) -> str:
"""Hybrid search across system documentation, markdown notes, architectural decisions, and runbooks."""
query = query.strip() if query else ""
Expand All @@ -73,7 +101,16 @@ async def handle_search_docs(
try:
from app.services.auth import enforce_tool_permission, Role
enforce_tool_permission(Role.VIEWER)
hits = _get_tools_attr("execute_hybrid_search", execute_hybrid_search)(query_text=query, doc_type="doc", repo=repo, category=category, tag=tag, limit=limit)
hits = _get_tools_attr("execute_hybrid_search", execute_hybrid_search)(
query_text=query,
doc_type="doc",
repo=repo,
category=category,
tag=tag,
limit=limit,
dense_weight=dense_weight,
search_mode=mode or "hybrid"
)
if not hits:
return f"No matching documentation found for query: '{query}'."

Expand All @@ -84,17 +121,24 @@ async def handle_search_docs(
header = f"### [{p.get('repo')}] {p.get('rel_path')}"
if p.get("heading") and p.get("heading") != "Root":
header += f" -> {p.get('heading')}"
header += "\n"
if tags_str:
header += f"\nTags: {tags_str}"
header += f"- **Tags**: {tags_str}\n"
link_url = p.get("permalink_url") or p.get("github_url")
if link_url:
header += f"\nSource Link: {link_url}"
header += f"\nRelevance Score: {hit.score:.4f} ({hit.score * 100:.1f}%)\n"
header += f"- **Source Link**: {link_url}\n"

score_val = float(hit.score) if isinstance(getattr(hit, "score", None), (int, float)) else 0.0
score_str = f"{score_val:.4f} ({score_val * 100:.1f}%)"
d_val = getattr(hit, "dense_score", None)
s_val = getattr(hit, "sparse_score", None)
if isinstance(d_val, (int, float)) and isinstance(s_val, (int, float)):
score_str += f" [Semantic: {float(d_val) * 100:.1f}% | Lexical: {float(s_val) * 100:.1f}%]"
header += f"- Relevance Score: {score_str}\n\n"

block = f"{header}---\n{p.get('content')}"
formatted.append(block)


return "\n\n========================\n\n".join(formatted)
except Exception as e:
logger.error(f"search_docs failed: {e}")
Expand Down
4 changes: 3 additions & 1 deletion app/models/schemas.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
from pydantic import BaseModel, Field
from typing import List, Optional, Dict, Any, Tuple
from typing import List, Optional, Dict, Any, Tuple, Literal

# Database & Sync Models
class RepoConfig(BaseModel):
Expand Down Expand Up @@ -117,6 +117,8 @@ class SearchRequest(BaseModel):
tag: Optional[str] = None
limit: int = 5
exact: bool = True
dense_weight: Optional[float] = Field(default=None, ge=0.0, le=1.0)
search_mode: Optional[Literal["hybrid", "semantic", "lexical"]] = "hybrid"

class SyncRequest(BaseModel):
repo: Optional[str] = None
Expand Down
8 changes: 4 additions & 4 deletions app/services/chunking/symbol_extractor.py
Original file line number Diff line number Diff line change
Expand Up @@ -154,12 +154,12 @@ def traverse(node, parent_symbol: Optional[str] = None):

elif node.type in CALL_NODE_TYPES:
target = extract_target_from_call_node(node, source_bytes)
if target and target not in ("self", "this", "super"):
if target and target not in ("self", "this", "super", "new", "var"):
# Clean method prefix if full_symbol has parent
active_src = parent_symbol if parent_symbol else file_symbol
# if current active symbol is a method like Foo.bar, extract just bar or Foo.bar
if active_src and "." in active_src:
active_src_name = active_src.split(".")[-1]
# If current active symbol is a method like Foo.bar, extract just bar
if parent_symbol and "." in parent_symbol:
active_src_name = parent_symbol.split(".")[-1]
else:
active_src_name = active_src
relationships.append({
Expand Down
Loading
Loading