From f47ab4f9337705d57eb0dffa09411ae7d96362c1 Mon Sep 17 00:00:00 2001 From: "github-actions[bot]" Date: Fri, 18 Sep 2026 01:04:22 +0000 Subject: [PATCH 1/4] =?UTF-8?q?docs:=20Initialize=20=F0=9F=A6=A9=20Flaming?= =?UTF-8?q?o=20Code=20Documentation=20run=20[skip=20ci]?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .flamingo-ai-technical-writer-status.md | 5 +++++ 1 file changed, 5 insertions(+) create mode 100644 .flamingo-ai-technical-writer-status.md diff --git a/.flamingo-ai-technical-writer-status.md b/.flamingo-ai-technical-writer-status.md new file mode 100644 index 00000000..132db0fa --- /dev/null +++ b/.flamingo-ai-technical-writer-status.md @@ -0,0 +1,5 @@ +# 🦩 Flamingo Code Documentation: Started + +Run ID: doc-orchestrator-1789693445469 +Status: In Progress +Started: 2026-09-18 01:04:22 UTC From 3626223ffe806ee84c241420f4f0c980e3f04a73 Mon Sep 17 00:00:00 2001 From: "github-actions[bot]" Date: Fri, 18 Sep 2026 01:04:27 +0000 Subject: [PATCH 2/4] chore(docs): Clean slate - remove all documentation (153 files) [skip ci] --- docs/README.md | 84 ------ docs/development/.gitignore | 7 - docs/development/README.md | 23 -- docs/development/architecture/README.md | 251 ---------------- docs/development/setup/environment.md | 58 ---- docs/development/setup/local-development.md | 90 ------ docs/diagrams/architecture/.gitignore | 8 - docs/diagrams/architecture/README-2.mmd | 17 -- docs/diagrams/architecture/README-3.mmd | 18 -- docs/diagrams/architecture/README.mmd | 29 -- .../architecture/agent-tools-core-2.mmd | 17 -- .../architecture/agent-tools-core-3.mmd | 13 - .../architecture/agent-tools-core.mmd | 12 - .../architecture/analysis_pipeline-2.mmd | 67 ----- .../architecture/analysis_pipeline-3.mmd | 27 -- .../architecture/analysis_pipeline-4.mmd | 10 - .../architecture/analysis_pipeline-5.mmd | 11 - .../architecture/analysis_pipeline-6.mmd | 24 -- .../architecture/analysis_pipeline-7.mmd | 9 - .../architecture/analysis_pipeline.mmd | 12 - docs/diagrams/architecture/backend-core-2.mmd | 18 -- docs/diagrams/architecture/backend-core-3.mmd | 10 - docs/diagrams/architecture/backend-core.mmd | 20 -- .../architecture/c_family_analyzers-2.mmd | 17 -- .../architecture/c_family_analyzers-3.mmd | 50 ---- .../architecture/c_family_analyzers-4.mmd | 13 - .../architecture/c_family_analyzers.mmd | 8 - docs/diagrams/architecture/cli-core.mmd | 27 -- docs/diagrams/architecture/config-core-2.mmd | 9 - docs/diagrams/architecture/config-core-3.mmd | 10 - docs/diagrams/architecture/config-core-4.mmd | 12 - docs/diagrams/architecture/config-core-5.mmd | 21 -- docs/diagrams/architecture/config-core-6.mmd | 16 - docs/diagrams/architecture/config-core.mmd | 43 --- .../diagrams/architecture/configuration-2.mmd | 42 --- .../diagrams/architecture/configuration-3.mmd | 20 -- .../diagrams/architecture/configuration-4.mmd | 20 -- .../diagrams/architecture/configuration-5.mmd | 16 - .../diagrams/architecture/configuration-6.mmd | 8 - docs/diagrams/architecture/configuration.mmd | 16 - .../dependency-analyzer-core-2.mmd | 22 -- .../architecture/dependency-analyzer-core.mmd | 21 -- .../dependency-analyzer-models-2.mmd | 60 ---- .../dependency-analyzer-models.mmd | 14 - .../documentation-generator-2.mmd | 43 --- .../documentation-generator-3.mmd | 13 - .../architecture/documentation-generator.mmd | 11 - .../diagrams/architecture/frontend-core-2.mmd | 24 -- .../diagrams/architecture/frontend-core-3.mmd | 51 ---- .../diagrams/architecture/frontend-core-4.mmd | 8 - docs/diagrams/architecture/frontend-core.mmd | 25 -- docs/diagrams/architecture/generation-2.mmd | 26 -- docs/diagrams/architecture/generation-3.mmd | 14 - docs/diagrams/architecture/generation.mmd | 11 - .../architecture/git_integration-2.mmd | 26 -- .../architecture/git_integration-3.mmd | 13 - .../diagrams/architecture/git_integration.mmd | 18 -- .../architecture/graph_construction-2.mmd | 19 -- .../architecture/graph_construction-3.mmd | 15 - .../architecture/graph_construction-4.mmd | 8 - .../architecture/graph_construction-5.mmd | 22 -- .../architecture/graph_construction.mmd | 11 - .../architecture/html_generation-2.mmd | 19 -- .../architecture/html_generation-3.mmd | 17 -- .../diagrams/architecture/html_generation.mmd | 10 - .../diagrams/architecture/java_analyzer-2.mmd | 16 - .../diagrams/architecture/java_analyzer-3.mmd | 6 - docs/diagrams/architecture/java_analyzer.mmd | 19 -- .../javascript_typescript_analyzers-2.mmd | 45 --- .../javascript_typescript_analyzers-3.mmd | 19 -- .../javascript_typescript_analyzers-4.mmd | 17 -- .../javascript_typescript_analyzers-5.mmd | 19 -- .../javascript_typescript_analyzers.mmd | 10 - docs/diagrams/architecture/job_models-2.mmd | 7 - docs/diagrams/architecture/job_models-3.mmd | 20 -- docs/diagrams/architecture/job_models-4.mmd | 15 - docs/diagrams/architecture/job_models.mmd | 52 ---- docs/diagrams/architecture/llm-services-2.mmd | 20 -- docs/diagrams/architecture/llm-services-3.mmd | 22 -- docs/diagrams/architecture/llm-services-4.mmd | 16 - docs/diagrams/architecture/llm-services-5.mmd | 24 -- docs/diagrams/architecture/llm-services.mmd | 24 -- .../architecture/logging-config-2.mmd | 21 -- .../architecture/logging-config-3.mmd | 13 - .../architecture/logging-config-4.mmd | 12 - docs/diagrams/architecture/logging-config.mmd | 14 - docs/diagrams/architecture/php_analyzer-2.mmd | 41 --- docs/diagrams/architecture/php_analyzer-3.mmd | 20 -- docs/diagrams/architecture/php_analyzer-4.mmd | 7 - docs/diagrams/architecture/php_analyzer.mmd | 9 - .../architecture/python_analyzer-2.mmd | 4 - .../architecture/python_analyzer-3.mmd | 17 -- .../architecture/python_analyzer-4.mmd | 6 - .../diagrams/architecture/python_analyzer.mmd | 11 - .../architecture/sample_fixtures-2.mmd | 28 -- .../architecture/sample_fixtures-3.mmd | 11 - .../architecture/sample_fixtures-4.mmd | 5 - .../architecture/sample_fixtures-5.mmd | 4 - .../diagrams/architecture/sample_fixtures.mmd | 4 - .../architecture/test-clustering-2.mmd | 23 -- .../architecture/test-clustering-3.mmd | 15 - .../diagrams/architecture/test-clustering.mmd | 37 --- .../architecture/test-multi-path-2.mmd | 12 - .../diagrams/architecture/test-multi-path.mmd | 25 -- docs/diagrams/architecture/test_suites-2.mmd | 19 -- docs/diagrams/architecture/test_suites-3.mmd | 25 -- docs/diagrams/architecture/test_suites-4.mmd | 15 - docs/diagrams/architecture/test_suites.mmd | 46 --- .../architecture/tree-sitter-analyzers.mmd | 22 -- docs/diagrams/architecture/utils-2.mmd | 22 -- docs/diagrams/architecture/utils-3.mmd | 11 - docs/diagrams/architecture/utils.mmd | 36 --- docs/getting-started/.gitignore | 7 - docs/getting-started/first-steps.md | 95 ------ docs/getting-started/introduction.md | 85 ------ docs/getting-started/prerequisites.md | 71 ----- docs/getting-started/quick-start.md | 114 -------- docs/reference/architecture/.gitignore | 8 - docs/reference/architecture/README.md | 123 -------- .../agent-tools-core/agent-tools-core.md | 123 -------- .../architecture/backend-core/backend-core.md | 98 ------- .../analysis_pipeline.md | 274 ------------------ .../dependency-analyzer-core.md | 124 -------- .../graph_construction.md | 218 -------------- .../dependency-analyzer-models.md | 194 ------------- .../documentation-generator.md | 228 --------------- .../backend-core/llm-services/llm-services.md | 224 -------------- .../logging-config/logging-config.md | 181 ------------ .../c_family_analyzers.md | 260 ----------------- .../tree-sitter-analyzers/java_analyzer.md | 178 ------------ .../javascript_typescript_analyzers.md | 229 --------------- .../tree-sitter-analyzers/php_analyzer.md | 225 -------------- .../tree-sitter-analyzers/python_analyzer.md | 196 ------------- .../tree-sitter-analyzers.md | 115 -------- .../architecture/cli-core/cli-core.md | 103 ------- .../cli-core/configuration/configuration.md | 249 ---------------- .../cli-core/generation/generation.md | 167 ----------- .../git_integration/git_integration.md | 147 ---------- .../html_generation/html_generation.md | 131 --------- .../cli-core/job_models/job_models.md | 244 ---------------- .../cli-core/cli-core/utils/utils.md | 168 ----------- .../cli-core/cli_core/utils/utils.md | 168 ----------- .../architecture/cli-core/configuration.md | 249 ---------------- .../architecture/cli-core/generation.md | 167 ----------- .../architecture/cli-core/git_integration.md | 147 ---------- .../architecture/cli-core/html_generation.md | 131 --------- .../architecture/cli-core/job_models.md | 244 ---------------- .../architecture/config-core/config-core.md | 220 -------------- .../frontend-core/frontend-core.md | 257 ---------------- .../test-clustering/test-clustering.md | 208 ------------- .../test-multi-path/sample_fixtures.md | 125 -------- .../test-multi-path/test-multi-path.md | 87 ------ .../test-multi-path/test_suites.md | 261 ----------------- 153 files changed, 9433 deletions(-) delete mode 100644 docs/README.md delete mode 100644 docs/development/.gitignore delete mode 100644 docs/development/README.md delete mode 100644 docs/development/architecture/README.md delete mode 100644 docs/development/setup/environment.md delete mode 100644 docs/development/setup/local-development.md delete mode 100644 docs/diagrams/architecture/.gitignore delete mode 100644 docs/diagrams/architecture/README-2.mmd delete mode 100644 docs/diagrams/architecture/README-3.mmd delete mode 100644 docs/diagrams/architecture/README.mmd delete mode 100644 docs/diagrams/architecture/agent-tools-core-2.mmd delete mode 100644 docs/diagrams/architecture/agent-tools-core-3.mmd delete mode 100644 docs/diagrams/architecture/agent-tools-core.mmd delete mode 100644 docs/diagrams/architecture/analysis_pipeline-2.mmd delete mode 100644 docs/diagrams/architecture/analysis_pipeline-3.mmd delete mode 100644 docs/diagrams/architecture/analysis_pipeline-4.mmd delete mode 100644 docs/diagrams/architecture/analysis_pipeline-5.mmd delete mode 100644 docs/diagrams/architecture/analysis_pipeline-6.mmd delete mode 100644 docs/diagrams/architecture/analysis_pipeline-7.mmd delete mode 100644 docs/diagrams/architecture/analysis_pipeline.mmd delete mode 100644 docs/diagrams/architecture/backend-core-2.mmd delete mode 100644 docs/diagrams/architecture/backend-core-3.mmd delete mode 100644 docs/diagrams/architecture/backend-core.mmd delete mode 100644 docs/diagrams/architecture/c_family_analyzers-2.mmd delete mode 100644 docs/diagrams/architecture/c_family_analyzers-3.mmd delete mode 100644 docs/diagrams/architecture/c_family_analyzers-4.mmd delete mode 100644 docs/diagrams/architecture/c_family_analyzers.mmd delete mode 100644 docs/diagrams/architecture/cli-core.mmd delete mode 100644 docs/diagrams/architecture/config-core-2.mmd delete mode 100644 docs/diagrams/architecture/config-core-3.mmd delete mode 100644 docs/diagrams/architecture/config-core-4.mmd delete mode 100644 docs/diagrams/architecture/config-core-5.mmd delete mode 100644 docs/diagrams/architecture/config-core-6.mmd delete mode 100644 docs/diagrams/architecture/config-core.mmd delete mode 100644 docs/diagrams/architecture/configuration-2.mmd delete mode 100644 docs/diagrams/architecture/configuration-3.mmd delete mode 100644 docs/diagrams/architecture/configuration-4.mmd delete mode 100644 docs/diagrams/architecture/configuration-5.mmd delete mode 100644 docs/diagrams/architecture/configuration-6.mmd delete mode 100644 docs/diagrams/architecture/configuration.mmd delete mode 100644 docs/diagrams/architecture/dependency-analyzer-core-2.mmd delete mode 100644 docs/diagrams/architecture/dependency-analyzer-core.mmd delete mode 100644 docs/diagrams/architecture/dependency-analyzer-models-2.mmd delete mode 100644 docs/diagrams/architecture/dependency-analyzer-models.mmd delete mode 100644 docs/diagrams/architecture/documentation-generator-2.mmd delete mode 100644 docs/diagrams/architecture/documentation-generator-3.mmd delete mode 100644 docs/diagrams/architecture/documentation-generator.mmd delete mode 100644 docs/diagrams/architecture/frontend-core-2.mmd delete mode 100644 docs/diagrams/architecture/frontend-core-3.mmd delete mode 100644 docs/diagrams/architecture/frontend-core-4.mmd delete mode 100644 docs/diagrams/architecture/frontend-core.mmd delete mode 100644 docs/diagrams/architecture/generation-2.mmd delete mode 100644 docs/diagrams/architecture/generation-3.mmd delete mode 100644 docs/diagrams/architecture/generation.mmd delete mode 100644 docs/diagrams/architecture/git_integration-2.mmd delete mode 100644 docs/diagrams/architecture/git_integration-3.mmd delete mode 100644 docs/diagrams/architecture/git_integration.mmd delete mode 100644 docs/diagrams/architecture/graph_construction-2.mmd delete mode 100644 docs/diagrams/architecture/graph_construction-3.mmd delete mode 100644 docs/diagrams/architecture/graph_construction-4.mmd delete mode 100644 docs/diagrams/architecture/graph_construction-5.mmd delete mode 100644 docs/diagrams/architecture/graph_construction.mmd delete mode 100644 docs/diagrams/architecture/html_generation-2.mmd delete mode 100644 docs/diagrams/architecture/html_generation-3.mmd delete mode 100644 docs/diagrams/architecture/html_generation.mmd delete mode 100644 docs/diagrams/architecture/java_analyzer-2.mmd delete mode 100644 docs/diagrams/architecture/java_analyzer-3.mmd delete mode 100644 docs/diagrams/architecture/java_analyzer.mmd delete mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers-2.mmd delete mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers-3.mmd delete mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers-4.mmd delete mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers-5.mmd delete mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers.mmd delete mode 100644 docs/diagrams/architecture/job_models-2.mmd delete mode 100644 docs/diagrams/architecture/job_models-3.mmd delete mode 100644 docs/diagrams/architecture/job_models-4.mmd delete mode 100644 docs/diagrams/architecture/job_models.mmd delete mode 100644 docs/diagrams/architecture/llm-services-2.mmd delete mode 100644 docs/diagrams/architecture/llm-services-3.mmd delete mode 100644 docs/diagrams/architecture/llm-services-4.mmd delete mode 100644 docs/diagrams/architecture/llm-services-5.mmd delete mode 100644 docs/diagrams/architecture/llm-services.mmd delete mode 100644 docs/diagrams/architecture/logging-config-2.mmd delete mode 100644 docs/diagrams/architecture/logging-config-3.mmd delete mode 100644 docs/diagrams/architecture/logging-config-4.mmd delete mode 100644 docs/diagrams/architecture/logging-config.mmd delete mode 100644 docs/diagrams/architecture/php_analyzer-2.mmd delete mode 100644 docs/diagrams/architecture/php_analyzer-3.mmd delete mode 100644 docs/diagrams/architecture/php_analyzer-4.mmd delete mode 100644 docs/diagrams/architecture/php_analyzer.mmd delete mode 100644 docs/diagrams/architecture/python_analyzer-2.mmd delete mode 100644 docs/diagrams/architecture/python_analyzer-3.mmd delete mode 100644 docs/diagrams/architecture/python_analyzer-4.mmd delete mode 100644 docs/diagrams/architecture/python_analyzer.mmd delete mode 100644 docs/diagrams/architecture/sample_fixtures-2.mmd delete mode 100644 docs/diagrams/architecture/sample_fixtures-3.mmd delete mode 100644 docs/diagrams/architecture/sample_fixtures-4.mmd delete mode 100644 docs/diagrams/architecture/sample_fixtures-5.mmd delete mode 100644 docs/diagrams/architecture/sample_fixtures.mmd delete mode 100644 docs/diagrams/architecture/test-clustering-2.mmd delete mode 100644 docs/diagrams/architecture/test-clustering-3.mmd delete mode 100644 docs/diagrams/architecture/test-clustering.mmd delete mode 100644 docs/diagrams/architecture/test-multi-path-2.mmd delete mode 100644 docs/diagrams/architecture/test-multi-path.mmd delete mode 100644 docs/diagrams/architecture/test_suites-2.mmd delete mode 100644 docs/diagrams/architecture/test_suites-3.mmd delete mode 100644 docs/diagrams/architecture/test_suites-4.mmd delete mode 100644 docs/diagrams/architecture/test_suites.mmd delete mode 100644 docs/diagrams/architecture/tree-sitter-analyzers.mmd delete mode 100644 docs/diagrams/architecture/utils-2.mmd delete mode 100644 docs/diagrams/architecture/utils-3.mmd delete mode 100644 docs/diagrams/architecture/utils.mmd delete mode 100644 docs/getting-started/.gitignore delete mode 100644 docs/getting-started/first-steps.md delete mode 100644 docs/getting-started/introduction.md delete mode 100644 docs/getting-started/prerequisites.md delete mode 100644 docs/getting-started/quick-start.md delete mode 100644 docs/reference/architecture/.gitignore delete mode 100644 docs/reference/architecture/README.md delete mode 100644 docs/reference/architecture/backend-core/agent-tools-core/agent-tools-core.md delete mode 100644 docs/reference/architecture/backend-core/backend-core.md delete mode 100644 docs/reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md delete mode 100644 docs/reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md delete mode 100644 docs/reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md delete mode 100644 docs/reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md delete mode 100644 docs/reference/architecture/backend-core/documentation-generator/documentation-generator.md delete mode 100644 docs/reference/architecture/backend-core/llm-services/llm-services.md delete mode 100644 docs/reference/architecture/backend-core/logging-config/logging-config.md delete mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md delete mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md delete mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md delete mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md delete mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md delete mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md delete mode 100644 docs/reference/architecture/cli-core/cli-core.md delete mode 100644 docs/reference/architecture/cli-core/cli-core/configuration/configuration.md delete mode 100644 docs/reference/architecture/cli-core/cli-core/generation/generation.md delete mode 100644 docs/reference/architecture/cli-core/cli-core/git_integration/git_integration.md delete mode 100644 docs/reference/architecture/cli-core/cli-core/html_generation/html_generation.md delete mode 100644 docs/reference/architecture/cli-core/cli-core/job_models/job_models.md delete mode 100644 docs/reference/architecture/cli-core/cli-core/utils/utils.md delete mode 100644 docs/reference/architecture/cli-core/cli_core/utils/utils.md delete mode 100644 docs/reference/architecture/cli-core/configuration.md delete mode 100644 docs/reference/architecture/cli-core/generation.md delete mode 100644 docs/reference/architecture/cli-core/git_integration.md delete mode 100644 docs/reference/architecture/cli-core/html_generation.md delete mode 100644 docs/reference/architecture/cli-core/job_models.md delete mode 100644 docs/reference/architecture/config-core/config-core.md delete mode 100644 docs/reference/architecture/frontend-core/frontend-core.md delete mode 100644 docs/reference/architecture/test-clustering/test-clustering.md delete mode 100644 docs/reference/architecture/test-multi-path/sample_fixtures.md delete mode 100644 docs/reference/architecture/test-multi-path/test-multi-path.md delete mode 100644 docs/reference/architecture/test-multi-path/test_suites.md diff --git a/docs/README.md b/docs/README.md deleted file mode 100644 index 2d3847c3..00000000 --- a/docs/README.md +++ /dev/null @@ -1,84 +0,0 @@ -# CodeWiki Documentation - -Welcome to the documentation for **CodeWiki** — an AI-powered documentation generator for source code repositories. - -## 📚 Table of Contents - -### Getting Started - -New to CodeWiki? Start here: - -- [Introduction](./getting-started/introduction.md) — What CodeWiki is, key features, and target audience -- [Prerequisites](./getting-started/prerequisites.md) — Required software, versions, and environment setup -- [Quick Start](./getting-started/quick-start.md) — Install CodeWiki and generate your first documentation set in ~5 minutes -- [First Steps](./getting-started/first-steps.md) — Customizing generation, tuning token budgets, and the Git/GitHub Pages workflow - -### Development - -Contributing to CodeWiki itself: - -- [Development Overview](./development/README.md) — Quick navigation for developing, testing, and contributing -- [Environment Setup](./development/setup/environment.md) — IDE recommendations, required tools, and dev environment variables -- [Local Development](./development/setup/local-development.md) — Cloning the repo, editable installs, and running the CLI or web app locally -- [Architecture Overview](./development/architecture/README.md) — High-level module architecture, core components, and data flow - -### Reference - -Technical reference documentation generated from source-code analysis: - -- [Reference Overview](./reference/architecture/README.md) — End-to-end architecture, module relationships, and summary -- [CLI Core](./reference/architecture/cli-core/cli-core.md) — Command-line orchestration, local config, Git integration, HTML generation -- [Backend Core](./reference/architecture/backend-core/backend-core.md) — Repository analysis, dependency graphs, module clustering, LLM agent orchestration -- [Frontend Core](./reference/architecture/frontend-core/frontend-core.md) — FastAPI web application, job processing, caching, doc serving -- [Config Core](./reference/architecture/config-core/config-core.md) — Shared runtime `Config` model for paths, provider settings, and agent instructions -- [Test Multi Path](./reference/architecture/test-multi-path/test-multi-path.md) — Fixtures and checks for multi-root source analysis -- [Test Clustering](./reference/architecture/test-clustering/test-clustering.md) — Diagnostic and validation scripts for LLM-based module clustering - -#### Backend Core Subsystems - -- [Agent Tools Core](./reference/architecture/backend-core/agent-tools-core/agent-tools-core.md) — Controlled repository inspection and documentation editing tools -- [Dependency Analyzer Core](./reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md) — File discovery, AST parsing, call analysis, and graph construction - - [Graph Construction](./reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md) - - [Analysis Pipeline](./reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md) -- [Tree Sitter Analyzers](./reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md) — Language support for C, C++, C#, Java, JavaScript, TypeScript, PHP, and Python - - [Python Analyzer](./reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md) - - [PHP Analyzer](./reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md) - - [JavaScript/TypeScript Analyzers](./reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md) - - [Java Analyzer](./reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md) - - [C-Family Analyzers](./reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md) -- [Dependency Analyzer Models](./reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md) — Repository, node, relationship, and analysis-result contracts -- [Documentation Generator](./reference/architecture/backend-core/documentation-generator/documentation-generator.md) — Top-level pipeline coordination and documentation output -- [LLM Services](./reference/architecture/backend-core/llm-services/llm-services.md) — Model selection, request counting, and fallback handling -- [Logging Config](./reference/architecture/backend-core/logging-config/logging-config.md) — Shared colorized logging support - -#### CLI Core Subsystems - -- [Generation](./reference/architecture/cli-core/generation.md) — CLI adapter for staged documentation generation -- [Configuration](./reference/architecture/cli-core/configuration.md) — Persisted settings, keyring-backed credentials, and agent instructions -- [Job Models](./reference/architecture/cli-core/job_models.md) — Generation job status, statistics, and LLM configuration models -- [Git Integration](./reference/architecture/cli-core/git_integration.md) — Clean-tree checks, branch creation, commits, and remote URL handling -- [HTML Generation](./reference/architecture/cli-core/html_generation.md) — Static documentation viewer generation - -#### Test Fixtures Reference - -- [Test Suites](./reference/architecture/test-multi-path/test_suites.md) -- [Sample Fixtures](./reference/architecture/test-multi-path/sample_fixtures.md) - -### Diagrams - -Visual documentation — Mermaid architecture diagrams generated from source-code analysis are available in: - -- `./diagrams/architecture/` — includes diagrams for the CLI, backend, frontend, config core, dependency analyzers, and test modules - -## 📖 Quick Links - -- [Project README](../README.md) — Main project README -- [Contributing](../CONTRIBUTING.md) — How to contribute -- [License](../LICENSE.md) — License information - -## 💬 Community - -CodeWiki does not use GitHub Issues or Discussions. For questions and discussion, join the OpenMSP Slack community: [https://www.openmsp.ai/](https://www.openmsp.ai/) ([join link](https://join.slack.com/t/openmsp/shared_invite/zt-36bl7mx0h-3~U2nFH6nqHqoTPXMaHEHA)). - ---- -*Documentation generated by [🦩 Flamingo Code Documentation](https://flamingo.run)* diff --git a/docs/development/.gitignore b/docs/development/.gitignore deleted file mode 100644 index a5d01be5..00000000 --- a/docs/development/.gitignore +++ /dev/null @@ -1,7 +0,0 @@ -# VoltAgent temp files -temp/ - -# JSON intermediate files (except schema/config) -*.json -!*-schema.json -!*-config.json diff --git a/docs/development/README.md b/docs/development/README.md deleted file mode 100644 index 74ab77c4..00000000 --- a/docs/development/README.md +++ /dev/null @@ -1,23 +0,0 @@ -# Development Documentation - -This section covers everything you need to develop, test, and contribute to **CodeWiki** itself — the AI-powered documentation generator, not a repository you're documenting *with* CodeWiki. - -CodeWiki is a Python package (`pyproject.toml`, Python `>=3.12`) organized around a CLI (`codewiki/cli`), a backend documentation pipeline (`codewiki/src/be`), a FastAPI web application (`codewiki/src/fe`), and a shared runtime configuration model (`codewiki/src/config.py`). - -## Quick Navigation - -| Guide | Description | -|---|---| -| [Environment Setup](setup/environment.md) | IDE recommendations, required tools, and development environment variables. | -| [Local Development](setup/local-development.md) | Cloning the repo, installing in editable mode, running the CLI and web app locally, and debugging. | -| [Architecture Overview](architecture/README.md) | High-level module architecture, core components, and data flow through the documentation pipeline. | -| [Security](security/README.md) | Credential storage, safe file access, input validation, and secure coding practices used in CodeWiki. | -| [Testing](testing/README.md) | Structure of the diagnostic/validation scripts, how to run them, and expectations for new tests. | -| [Contributing Guidelines](contributing/guidelines.md) | Code style, branch naming, commit format, and the review checklist for pull requests. | - -## Where to Start - -1. Read the [Architecture Overview](architecture/README.md) to understand how the CLI, backend, and frontend modules fit together. -2. Follow [Environment Setup](setup/environment.md) and [Local Development](setup/local-development.md) to get a working local copy of CodeWiki. -3. Review [Security](security/README.md) and [Testing](testing/README.md) before submitting changes. -4. Read [Contributing Guidelines](contributing/guidelines.md) before opening a pull request. diff --git a/docs/development/architecture/README.md b/docs/development/architecture/README.md deleted file mode 100644 index db28893a..00000000 --- a/docs/development/architecture/README.md +++ /dev/null @@ -1,251 +0,0 @@ -# Architecture Overview - -CodeWiki is a layered, AI-driven documentation engine. This document describes the high-level architecture, core components, data flow, and key design decisions. - -For detailed module-level documentation, see the [Reference Architecture](../../reference/architecture/README.md). - ---- - -## High-Level Architecture - -CodeWiki is organized into six core layers, each with clearly defined responsibilities: - -```mermaid -graph TD - User["User (CLI or Web)"] --> CLICore["CLI Core"] - User --> FECore["Frontend Core (FastAPI)"] - - CLICore --> DocGen["Documentation Generator"] - FECore --> DocGen - - DocGen --> DepAnalysis["Dependency Analysis"] - DocGen --> AgentOrch["Agent Orchestration"] - - AgentOrch --> LLMSvc["LLM Services"] - AgentOrch --> Tools["Agent Tools"] - - DepAnalysis --> GraphBuilder["Dependency Graph Builder"] - GraphBuilder --> DocGen - - DocGen --> Utils["Utils (FileManager)"] - DocGen --> Output["Generated Documentation"] - - LLMSvc --> Provider["External LLM Providers"] -``` - ---- - -## Core Components - -| Component | Location | Responsibility | -|-----------|----------|---------------| -| **CLI Core** | `codewiki/cli/` | Command-line interface, config management, Git integration, HTML generation | -| **Frontend Core** | `codewiki/src/fe/` | FastAPI web app, job queue, caching, Markdown rendering | -| **Documentation Generator** | `codewiki/src/be/documentation_generator.py` | Orchestrates the full generation pipeline | -| **Agent Orchestration** | `codewiki/src/be/agent_orchestrator.py` | AI agent lifecycle: creation, tooling, execution, persistence | -| **LLM Services** | `codewiki/src/be/llm_services.py` | LLM provider factory, fallback chaining, request counting | -| **Dependency Analysis** | `codewiki/src/be/dependency_analyzer/` | AST parsing, call graph resolution, dependency graph building | -| **Utils** | `codewiki/src/utils.py` | Stateless `FileManager` for all file and JSON I/O | - ---- - -## Generation Pipeline (Data Flow) - -The complete documentation generation pipeline runs in a deterministic sequence of stages: - -```mermaid -flowchart TD - Start["Start Job"] --> DepAnalysis["1. Dependency Analysis"] - DepAnalysis --> Cluster["2. Module Clustering (LLM)"] - Cluster --> LeafDocs["3. Generate Leaf Module Docs"] - LeafDocs --> ParentDocs["4. Generate Parent Module Docs"] - ParentDocs --> Overview["5. Generate Repository Overview"] - Overview --> Metadata["6. Write metadata.json"] - Metadata --> EndNode["Done"] -``` - -### Stage Descriptions - -| Stage | Weight | Description | -|-------|--------|-------------| -| **Dependency Analysis** | 40% | Parse all source files, extract components and relationships, build dependency graph | -| **Module Clustering** | 20% | Use LLM to group related components into logical modules | -| **Documentation Generation** | 30% | Generate Markdown documentation, leaf-first | -| **HTML Generation** | 5% | Build static `index.html` for GitHub Pages (optional) | -| **Finalization** | 5% | Write `metadata.json`, commit to Git branch (optional) | - ---- - -## Dependency Analysis Pipeline - -The Dependency Analysis layer transforms raw source code into a structured dependency graph: - -```mermaid -flowchart LR - Repo["Repository"] --> RepoAnalyzer["Repo Analyzer"] - RepoAnalyzer --> FileTree["File Tree"] - FileTree --> CallGraphAnalyzer["Call Graph Analyzer"] - CallGraphAnalyzer --> LangAnalyzers["Language Analyzers"] - LangAnalyzers --> Nodes["Node Models"] - LangAnalyzers --> Relationships["CallRelationship Models"] - Nodes --> DepParser["Dependency Parser"] - Relationships --> DepParser - DepParser --> GraphBuilder["Dependency Graph Builder"] - GraphBuilder --> AnalysisResult["AnalysisResult"] -``` - -**Supported languages:** - -| Language | Analyzer | -|---------|---------| -| Python | `PythonASTAnalyzer` (native `ast` module) | -| JavaScript | `TreeSitterJSAnalyzer` | -| TypeScript | `TreeSitterTSAnalyzer` | -| Java | `TreeSitterJavaAnalyzer` | -| C# | `TreeSitterCSharpAnalyzer` | -| C | `TreeSitterCAnalyzer` | -| C++ | `TreeSitterCppAnalyzer` | -| PHP | `TreeSitterPHPAnalyzer` | - ---- - -## Agent Orchestration - -The Agent Orchestration layer manages the AI agent lifecycle for each module: - -```mermaid -flowchart TD - Start["Process Module"] --> CheckComplex{"is_complex_module()?"} - CheckComplex -->|"Yes"| ComplexAgent["Create Complex Agent"] - CheckComplex -->|"No"| LeafAgent["Create Leaf Agent"] - - ComplexAgent --> Tools1["read_code_components + str_replace_editor + generate_sub_module_documentation"] - LeafAgent --> Tools2["read_code_components + str_replace_editor"] - - Tools1 --> Execute["Run Agent (async)"] - Tools2 --> Execute - Execute --> Verify["Verify Output File"] - Verify --> Persist["Save module_tree.json"] -``` - -**Idempotency:** Documentation generation is skipped if a module's output file already exists. This allows safe re-runs and incremental updates. - ---- - -## LLM Services Architecture - -The LLM Services layer abstracts provider differences and manages three model roles: - -```mermaid -flowchart TD - Request["Generation Request"] --> TryMain["Main Model"] - TryMain -->|"Success"| ReturnMain["Return Response"] - TryMain -->|"Failure"| TryFallback["Fallback Model"] - TryFallback --> ReturnFallback["Return Fallback Response"] - - ClusterRequest["Cluster Request"] --> ClusterModel["Cluster Model"] - ClusterModel --> ClusterResponse["Cluster Response"] -``` - -| Role | Default | Purpose | -|------|---------|---------| -| **Main model** | `claude-sonnet-4` | Primary documentation generation | -| **Cluster model** | Same as main | Module grouping and structural reasoning | -| **Fallback model** | Same as main | Backup if main model fails | - ---- - -## Frontend Core Architecture - -```mermaid -sequenceDiagram - participant Browser - participant Routes as "WebRoutes" - participant Worker as "BackgroundWorker" - participant Cache as "CacheManager" - participant GitHub as "GitHubRepoProcessor" - participant DocGen as "DocumentationGenerator" - - Browser->>Routes: POST / (repo_url) - Routes->>Cache: Check cache (SHA-256 hash) - Cache-->>Routes: Hit or miss - Routes->>Worker: Enqueue job - Worker->>GitHub: Clone repository - Worker->>DocGen: Run documentation generation - DocGen-->>Worker: Write docs to output directory - Worker->>Cache: Store result - Browser->>Routes: GET /docs/{job_id} - Routes-->>Browser: Rendered HTML -``` - ---- - -## Output Directory Structure - -All generated documentation follows a deterministic hierarchical structure: - -```text -docs/ -├── README.md ← Repository overview -├── module_tree.json ← Module hierarchy -├── metadata.json ← Generation audit log -└── / - ├── .md ← Module documentation - └── / - └── .md ← Sub-module documentation -``` - -This structure mirrors the logical module hierarchy discovered from the codebase — not the raw file system layout. - ---- - -## Key Design Decisions - -### 1. Leaf-First (Bottom-Up) Generation - -Parent modules are documented *after* their children. This ensures: - -- LLM context stays bounded — each call only receives one module's context -- Parent overviews accurately summarize child documentation -- Generation is parallelizable at the leaf level - -### 2. Synthetic Module Fallback - -If the clustering LLM fails to produce any modules, CodeWiki automatically creates synthetic modules grouped by top-level directory. This prevents the system from attempting to document an entire large repository in a single LLM call. - -### 3. OS Keychain for API Keys - -API keys are never written to disk in plaintext. The CLI uses the system keychain (`keyring` library), which maps to macOS Keychain, Windows Credential Manager, or Linux Secret Service depending on the platform. - -### 4. Provider-Agnostic LLM Layer - -The LLM Services layer uses the OpenAI SDK for all providers, with configurable base URLs. This allows CodeWiki to work with Anthropic, Azure, LiteLLM proxies, and any OpenAI-compatible endpoint without code changes. - -### 5. Idempotent Generation - -The `process_module()` method skips generation if the output file already exists. This makes re-runs safe and supports incremental documentation updates. - -### 6. Deterministic Namespacing - -All component identifiers follow the pattern: - -```text -... -``` - -This allows cross-language and cross-repository dependency resolution without ambiguity. - ---- - -## Reference Documentation - -For deep-dive module documentation, see: - -- [Reference Architecture Overview](../../reference/architecture/README.md) -- [CLI Core](../../reference/architecture/cli-core/cli-core.md) -- [Agent Orchestration](../../reference/architecture/agent-orchestration/agent-orchestration.md) -- [LLM Services](../../reference/architecture/llm-services/llm-services.md) -- [Dependency Analysis](../../reference/architecture/dependency-analysis/dependency-analysis.md) -- [Documentation Generation](../../reference/architecture/documentation-generation/documentation-generation.md) -- [Frontend Core](../../reference/architecture/frontend-core/frontend-core.md) -- [Utils](../../reference/architecture/utils/utils.md) diff --git a/docs/development/setup/environment.md b/docs/development/setup/environment.md deleted file mode 100644 index 1ab4a478..00000000 --- a/docs/development/setup/environment.md +++ /dev/null @@ -1,58 +0,0 @@ -# Development Environment Setup - -This guide covers the tools and settings recommended for developing CodeWiki itself. - -## Required Development Tools - -| Tool | Version | Notes | -|---|---|---| -| Python | `>=3.12` | Matches `requires-python` in `pyproject.toml`. | -| pip | Latest | Used to install both runtime and `dev` optional dependencies. | -| Git | Any recent version | CodeWiki's own CLI and web app both shell out to `git` / use GitPython. | -| Node.js | `>=14.0.0` | Required by `mermaid-py`, which validates Mermaid diagrams embedded in generated docs during tests and generation. | -| Docker & Docker Compose | Recent version | Optional, for testing the containerized web app (`docker/docker-compose.yml`, `docker/Dockerfile`). | - -## Installing Development Dependencies - -CodeWiki defines an optional `dev` dependency group in `pyproject.toml`: - -```bash -pip install -e ".[dev]" -``` - -This installs, in addition to the runtime dependencies: - -- `pytest`, `pytest-cov`, `pytest-asyncio` — testing -- `black` — code formatting (line length 100, target `py312`) -- `mypy` — static type checking (`python_version = "3.12"`) -- `ruff` — linting - -## IDE Recommendations - -Any editor with solid Python tooling works well. If using **VS Code**, the following extensions align with the project's tooling: - -- **Python** (Microsoft) — core language support, linting, debugging -- **Pylance** — type-checking assistance (complements `mypy`) -- **Black Formatter** — matches the project's `[tool.black]` configuration (line-length 100, `py312` target) -- **Ruff** — matches the project's linter configuration -- **Mermaid Preview** — useful when reviewing generated architecture diagrams in Markdown output - -If using **PyCharm**, enable Black as the external formatter and configure the line length to 100 to match `[tool.black]` in `pyproject.toml`. - -## Development Environment Variables - -CodeWiki's core CLI configuration lives in `~/.codewiki/config.json` and the OS keyring — it does not require environment variables for normal development use. However, a few areas of the codebase do read environment variables directly: - -| Variable | Used By | Purpose | -|---|---|---| -| `OPENAI_API_KEY` | Ad-hoc clustering test scripts (`test_clustering_*.py`) | API key for OpenAI-compatible providers during manual testing. | -| `ANTHROPIC_API_KEY` | Ad-hoc clustering test scripts | API key for Anthropic providers during manual testing. | -| `MAIN_API_KEY` / `CLUSTER_API_KEY` / `FALLBACK_API_KEY` | Ad-hoc clustering test scripts | Per-role overrides used when running the standalone clustering diagnostics. | -| `PYTHONPATH` | Docker image / `run_web_app.py` path setup | Set to `/app` inside the container; locally, `run_web_app.py` inserts `codewiki/src` onto `sys.path` itself. | -| `APP_PORT` | `docker/docker-compose.yml` | Host port mapping for the containerized web app (defaults to `8000`). | - -> **Tip:** For scripts that read API keys from the environment, consider using a local `.env.local` file with `python-dotenv` (already a project dependency) rather than exporting secrets into your shell history. - -## Next Steps - -Continue to [Local Development](local-development.md) to clone the repository, install it in editable mode, and run the CLI or web app locally. diff --git a/docs/development/setup/local-development.md b/docs/development/setup/local-development.md deleted file mode 100644 index 99763148..00000000 --- a/docs/development/setup/local-development.md +++ /dev/null @@ -1,90 +0,0 @@ -# Local Development - -This guide walks through cloning CodeWiki, installing it in editable mode, and running it locally against a test repository. - -## Clone and Install - -```bash -git clone https://github.com/flamingo-stack/CodeWiki.git -cd CodeWiki - -# Editable install with development dependencies -pip install -e ".[dev]" -``` - -Editable installs (`-e`) mean changes to `codewiki/` source files take effect immediately without reinstalling — ideal for iterating on the CLI, backend pipeline, or web app. - -Verify the CLI is on your `$PATH` and pointing at your local checkout: - -```bash -codewiki --version -``` - -## Running the CLI Locally - -Once installed, you can run `codewiki` against any repository (including CodeWiki's own repository, or one of the bundled test fixtures under `test-multi-path/`): - -```bash -codewiki config set \ - --cluster-api-key "sk-..." --main-api-key "sk-..." \ - --cluster-model "your-model" --main-model "your-model" \ - --cluster-base-url "https://api.your-provider.com/v1" \ - --main-base-url "https://api.your-provider.com/v1" - -cd /path/to/some/repo -codewiki generate --verbose -``` - -You can also invoke the CLI as a module without relying on the installed console script: - -```bash -python -m codewiki --help -python -m codewiki generate -``` - -## Running the Web Application Locally - -The FastAPI web app can be started directly with Python: - -```bash -python codewiki/run_web_app.py -``` - -This inserts `codewiki/src` onto `sys.path` and delegates to `fe.web_app.main()`. By default it listens on `127.0.0.1:8000` (see `WebAppConfig.DEFAULT_HOST` / `DEFAULT_PORT` in `codewiki/src/fe/config.py`). - -Alternatively, run it in a container using Docker Compose: - -```bash -cd docker -docker compose up --build -``` - -The container maps port `8000` (overridable with the `APP_PORT` environment variable) and mounts `../output` for persistent cache/output storage, plus your `~/.ssh` directory (read-only) for cloning private repositories over SSH. - -## Hot Reload / Iterating on Code - -- **CLI changes**: Because the package is installed with `pip install -e .`, edits to any file under `codewiki/cli/` or `codewiki/src/` are picked up the next time you invoke `codewiki` — no reinstall needed. -- **Web app changes**: `run_web_app.py` does not enable an auto-reloading development server by default. If you need hot reload while iterating on FastAPI routes, restart `python codewiki/run_web_app.py` after each change, or run the underlying ASGI app through `uvicorn` with `--reload` if you invoke it that way directly. - -## Debug Configuration - -For debugging the CLI in an IDE (e.g., VS Code's Python debugger), set the program to run as a module with arguments, for example: - -```bash -python -m codewiki generate --verbose -``` - -Configure your IDE's launch configuration to run `codewiki/cli/main.py` (or `python -m codewiki`) with the working directory set to the repository you want to analyze, and pass CLI arguments like `generate --verbose` for step-by-step output. - -For debugging the web app, set breakpoints inside `codewiki/src/fe/routes.py` or `codewiki/src/fe/background_worker.py` and launch `codewiki/run_web_app.py` directly through your IDE's debugger rather than the CLI. - -## Testing Your Changes Against Sample Fixtures - -The repository includes a self-contained multi-path test fixture at `test-multi-path/` (with `main/`, `deps/`, `external/` subdirectories) specifically designed to exercise multi-root dependency analysis. Use it to sanity-check changes to the dependency analyzer without needing a large external repository: - -```bash -python test-multi-path/test_multi_path.py -python test-multi-path/integration_test.py -``` - -See the [Testing](../testing/README.md) guide for more on how these and the clustering diagnostic scripts are organized. diff --git a/docs/diagrams/architecture/.gitignore b/docs/diagrams/architecture/.gitignore deleted file mode 100644 index b4f4c7a6..00000000 --- a/docs/diagrams/architecture/.gitignore +++ /dev/null @@ -1,8 +0,0 @@ -# CodeWiki temp files (dependency graphs can be 7GB+) -temp/ -dependency_graphs/ - -# JSON intermediate files (except schema/config) -*.json -!*-schema.json -!*-config.json diff --git a/docs/diagrams/architecture/README-2.mmd b/docs/diagrams/architecture/README-2.mmd deleted file mode 100644 index 765cb17c..00000000 --- a/docs/diagrams/architecture/README-2.mmd +++ /dev/null @@ -1,17 +0,0 @@ -sequenceDiagram - participant Caller - participant Config as "Config Core" - participant Generator as "DocumentationGenerator" - participant Analyzer as "Dependency Analyzer" - participant Agents as "AgentOrchestrator" - participant Output as "Documentation Output" - - Caller->>Config: Build validated runtime configuration - Caller->>Generator: Start documentation run - Generator->>Analyzer: Analyze repository files and calls - Analyzer-->>Generator: Dependency graph and leaf components - Generator->>Generator: Cluster components into modules - Generator->>Agents: Generate leaf module documentation - Agents-->>Generator: Generated module content - Generator->>Output: Write Markdown, module tree, and metadata - Generator-->>Caller: Generation complete diff --git a/docs/diagrams/architecture/README-3.mmd b/docs/diagrams/architecture/README-3.mmd deleted file mode 100644 index a75c9430..00000000 --- a/docs/diagrams/architecture/README-3.mmd +++ /dev/null @@ -1,18 +0,0 @@ -flowchart LR - Config["Config Core"] --> CLI["CLI Core"] - Config --> Frontend["Frontend Core"] - Config --> Backend["Backend Core"] - - CLI --> Backend - Frontend --> Backend - - Backend --> AgentTools["Agent Tools Core"] - Backend --> Analyzer["Dependency Analyzer Core"] - Backend --> Language["Tree-sitter Analyzers"] - Backend --> LLM["LLM Services"] - - TestPaths["Test Multi Path"] --> Config - TestPaths --> Analyzer - - TestCluster["Test Clustering"] --> Config - TestCluster --> Backend diff --git a/docs/diagrams/architecture/README.mmd b/docs/diagrams/architecture/README.mmd deleted file mode 100644 index 3b6ad602..00000000 --- a/docs/diagrams/architecture/README.mmd +++ /dev/null @@ -1,29 +0,0 @@ -flowchart TD - User["Developer or Web User"] --> Entry{{"Choose entry point?"}} - Entry -->|CLI| CLI["CLI Core"] - Entry -->|Web| Frontend["Frontend Core"] - - CLI --> RuntimeConfig["Config Core"] - Frontend --> RuntimeConfig - - CLI --> GitOps["Git Integration"] - Frontend --> RepoProcessor["GitHub Repository Processor"] - - GitOps --> Source["Source Repository"] - RepoProcessor --> Source - - RuntimeConfig --> Generator["DocumentationGenerator"] - Source --> Generator - - Generator --> Analysis["Dependency Analysis"] - Analysis --> Parsers["Language Parsers"] - Parsers --> Graph["Dependency Graph"] - - Graph --> Clustering["Module Clustering"] - Clustering --> Agents["LLM Agent Orchestration"] - Agents --> Docs["Markdown Documentation"] - - Docs --> Metadata["Module Tree and Metadata"] - Metadata --> HTML["Optional HTML Viewer"] - Docs --> Output["Generated Documentation Output"] - HTML --> Output diff --git a/docs/diagrams/architecture/agent-tools-core-2.mmd b/docs/diagrams/architecture/agent-tools-core-2.mmd deleted file mode 100644 index d8ffbd0e..00000000 --- a/docs/diagrams/architecture/agent-tools-core-2.mmd +++ /dev/null @@ -1,17 +0,0 @@ -sequenceDiagram - participant Agent - participant Tool as "str_replace_editor" - participant Edit as "EditTool" - participant FS as "Filesystem" - participant Validator as "Mermaid Validator" - - Agent->>Tool: command="create", working_dir="docs", path="module.md" - Tool->>Tool: resolve absolute_path under docs root - Tool->>Edit: EditTool(registry, docs_path) - Edit->>Edit: validate_path(command, path) - Edit->>FS: create parent dirs + write_file - FS-->>Edit: written - Edit-->>Tool: success log - Tool->>Validator: validate_mermaid_diagrams(path) - Validator-->>Tool: validation result - Tool-->>Agent: combined result string diff --git a/docs/diagrams/architecture/agent-tools-core-3.mmd b/docs/diagrams/architecture/agent-tools-core-3.mmd deleted file mode 100644 index 3c61cd9f..00000000 --- a/docs/diagrams/architecture/agent-tools-core-3.mmd +++ /dev/null @@ -1,13 +0,0 @@ -flowchart LR - Edit["Edit .md file"] --> Check{{"path ends with .md?"}} - Check -->|"no"| Done["Return result"] - Check -->|"yes"| Attempts["Read mermaid_attempts counter"] - Attempts --> Limit{{"attempts >= MAX?"}} - Limit -->|"yes"| Skip["Skip validation, warn agent"] - Limit -->|"no"| Validate["validate_mermaid_diagrams"] - Validate --> HasError{{"errors found?"}} - HasError -->|"yes"| Increment["Increment counter in registry"] - HasError -->|"no"| Reset["Reset counter to 0"] - Increment --> Done - Reset --> Done - Skip --> Done diff --git a/docs/diagrams/architecture/agent-tools-core.mmd b/docs/diagrams/architecture/agent-tools-core.mmd deleted file mode 100644 index f7eb85cd..00000000 --- a/docs/diagrams/architecture/agent-tools-core.mmd +++ /dev/null @@ -1,12 +0,0 @@ -flowchart TD - Orchestrator["AgentOrchestrator"] -->|"builds"| Deps["CodeWikiDeps"] - Orchestrator -->|"invokes tool with"| ToolCall["str_replace_editor Tool Call"] - ToolCall -->|"uses ctx.deps"| Deps - ToolCall -->|"instantiates"| EditTool["EditTool"] - EditTool -->|"large .py view"| Filemap["Filemap"] - EditTool -->|"view/edit windows"| WindowExpander["WindowExpander"] - EditTool -->|"reads"| RepoFS[("Repository Files")] - EditTool -->|"reads/writes"| DocsFS[("Generated Docs Files")] - ToolCall -->|"post-edit .md check"| MermaidValidator["validate_mermaid_diagrams"] - Deps -->|"references"| Config["Config"] - Deps -->|"references"| Node["Node (component registry)"] diff --git a/docs/diagrams/architecture/analysis_pipeline-2.mmd b/docs/diagrams/architecture/analysis_pipeline-2.mmd deleted file mode 100644 index 9488e900..00000000 --- a/docs/diagrams/architecture/analysis_pipeline-2.mmd +++ /dev/null @@ -1,67 +0,0 @@ -classDiagram - class AnalysisService { - +call_graph_analyzer: CallGraphAnalyzer - +analyze_local_repository(repo_path, max_files, languages) dict - +analyze_repository_full(github_url, include_patterns, exclude_patterns) AnalysisResult - +analyze_repository_structure_only(github_url, include_patterns, exclude_patterns) dict - +cleanup_all() - -_clone_repository(github_url) str - -_analyze_structure(repo_dir, include, exclude) dict - -_analyze_call_graph(file_tree, repo_dir) dict - -_read_readme_file(repo_dir) str - -_cleanup_repository(temp_dir) - } - class RepoAnalyzer { - +include_patterns: list - +exclude_patterns: list - +analyze_repository_structure(repo_dir) dict - -_build_file_tree(repo_dir) dict - -_should_exclude_path(path, filename) bool - -_should_include_file(path, filename) bool - -_count_files(tree) int - -_calculate_size(tree) float - } - class CallGraphAnalyzer { - +functions: dict~str, Node~ - +call_relationships: list~CallRelationship~ - +analyze_code_files(code_files, base_dir) dict - +extract_code_files(file_tree) list - -_analyze_code_file(repo_dir, file_info) - -_resolve_call_relationships() - -_deduplicate_relationships() - -_generate_visualization_data() dict - +generate_llm_format() dict - -_select_most_connected_nodes(target_count) - } - class AnalysisResult { - +repository: Repository - +functions: list - +relationships: list - +file_tree: dict - +summary: dict - +visualization: dict - +readme_content: str - } - class Node { - +id: str - +name: str - +node_type: str - +file_path: str - +component_id: str - +docstring: str - +parameters: list - } - class CallRelationship { - +caller: str - +callee: str - +call_line: int - +is_resolved: bool - } - - AnalysisService --> RepoAnalyzer : uses - AnalysisService --> CallGraphAnalyzer : uses - AnalysisService --> AnalysisResult : produces - CallGraphAnalyzer --> Node : produces - CallGraphAnalyzer --> CallRelationship : produces - AnalysisResult --> Node : contains - AnalysisResult --> CallRelationship : contains diff --git a/docs/diagrams/architecture/analysis_pipeline-3.mmd b/docs/diagrams/architecture/analysis_pipeline-3.mmd deleted file mode 100644 index bb86a740..00000000 --- a/docs/diagrams/architecture/analysis_pipeline-3.mmd +++ /dev/null @@ -1,27 +0,0 @@ -sequenceDiagram - participant Caller - participant AS as AnalysisService - participant Clone as "Cloning Utility" - participant RA as RepoAnalyzer - participant CGA as CallGraphAnalyzer - participant FS as Filesystem - - Caller->>AS: analyze_repository_full(github_url) - AS->>Clone: clone_repository(github_url) - Clone-->>AS: temp_dir - AS->>AS: parse_github_url(github_url) - AS->>RA: analyze_repository_structure(temp_dir) - RA->>FS: walk directory tree - RA-->>AS: file_tree + summary - AS->>CGA: extract_code_files(file_tree) - CGA-->>AS: code_files - AS->>CGA: analyze_code_files(code_files, temp_dir) - CGA->>CGA: dispatch per language, parse each file - CGA->>CGA: resolve_call_relationships() - CGA->>CGA: deduplicate_relationships() - CGA->>CGA: generate_visualization_data() - CGA-->>AS: functions, relationships, call_graph, visualization - AS->>FS: read README file - AS->>AS: build AnalysisResult - AS->>Clone: cleanup_repository(temp_dir) - AS-->>Caller: AnalysisResult diff --git a/docs/diagrams/architecture/analysis_pipeline-4.mmd b/docs/diagrams/architecture/analysis_pipeline-4.mmd deleted file mode 100644 index d720e0bf..00000000 --- a/docs/diagrams/architecture/analysis_pipeline-4.mmd +++ /dev/null @@ -1,10 +0,0 @@ -flowchart LR - subgraph Full["analyze_repository_full"] - F1["Clone"] --> F2["Structure Analysis"] --> F3["Call Graph Analysis"] --> F4["README Read"] --> F5["AnalysisResult"] - end - subgraph StructureOnly["analyze_repository_structure_only"] - S1["Clone"] --> S2["Structure Analysis"] --> S3["Structure Dict"] - end - subgraph Local["analyze_local_repository"] - L1["Local Path"] --> L2["Structure Analysis"] --> L3["Filter by language / max_files"] --> L4["Call Graph Analysis"] --> L5["Simplified Dict"] - end diff --git a/docs/diagrams/architecture/analysis_pipeline-5.mmd b/docs/diagrams/architecture/analysis_pipeline-5.mmd deleted file mode 100644 index c961c238..00000000 --- a/docs/diagrams/architecture/analysis_pipeline-5.mmd +++ /dev/null @@ -1,11 +0,0 @@ -flowchart TD - Input["repo_dir: str or list[str]"] --> Check{{"Single path?"}} - Check -->|"Yes"| Single["_build_file_tree(repo_dir)"] - Check -->|"No"| Multi["_analyze_multiple_repositories(repo_dirs)"] - Multi --> Loop["For each repo_dir: compute namespace, build tree"] - Loop --> Wrap["Wrap tree with namespace prefix"] - Wrap --> Merge["Merge into single root tree"] - Single --> Summary1["Compute total_files, total_size_kb"] - Merge --> Summary2["Aggregate totals across repositories"] - Summary1 --> Out["file_tree + summary"] - Summary2 --> Out diff --git a/docs/diagrams/architecture/analysis_pipeline-6.mmd b/docs/diagrams/architecture/analysis_pipeline-6.mmd deleted file mode 100644 index 6595ab6f..00000000 --- a/docs/diagrams/architecture/analysis_pipeline-6.mmd +++ /dev/null @@ -1,24 +0,0 @@ -flowchart TD - FT["file_tree"] --> Extract["extract_code_files()"] - Extract --> CodeFiles["code_files list"] - CodeFiles --> Dispatch{{"Dispatch by language"}} - Dispatch -->|"python"| PyA["Python AST Analyzer"] - Dispatch -->|"javascript"| JsA["JS Tree-Sitter Analyzer"] - Dispatch -->|"typescript"| TsA["TS Tree-Sitter Analyzer"] - Dispatch -->|"java"| JavaA["Java Tree-Sitter Analyzer"] - Dispatch -->|"csharp"| CsA["C# Tree-Sitter Analyzer"] - Dispatch -->|"c"| CA["C Tree-Sitter Analyzer"] - Dispatch -->|"cpp"| CppA["C++ Tree-Sitter Analyzer"] - Dispatch -->|"php"| PhpA["PHP Tree-Sitter Analyzer"] - PyA --> Agg["functions dict, call_relationships list"] - JsA --> Agg - TsA --> Agg - JavaA --> Agg - CsA --> Agg - CA --> Agg - CppA --> Agg - PhpA --> Agg - Agg --> Resolve["_resolve_call_relationships()"] - Resolve --> Dedup["_deduplicate_relationships()"] - Dedup --> Viz["_generate_visualization_data()"] - Viz --> Result["functions, relationships, call_graph, visualization"] diff --git a/docs/diagrams/architecture/analysis_pipeline-7.mmd b/docs/diagrams/architecture/analysis_pipeline-7.mmd deleted file mode 100644 index 625b8c6d..00000000 --- a/docs/diagrams/architecture/analysis_pipeline-7.mmd +++ /dev/null @@ -1,9 +0,0 @@ -flowchart TD - Start["callee: str"] --> Direct{{"callee in func_lookup?"}} - Direct -->|"Yes"| Resolved["Mark resolved, rewrite callee to func_id"] - Direct -->|"No"| HasDot{{"Contains '.'?"}} - HasDot -->|"No"| Unresolved["Leave unresolved"] - HasDot -->|"Yes"| MethodName["Extract trailing segment after last '.'"] - MethodName --> MethodLookup{{"method_name in func_lookup?"}} - MethodLookup -->|"Yes"| Resolved - MethodLookup -->|"No"| Unresolved diff --git a/docs/diagrams/architecture/analysis_pipeline.mmd b/docs/diagrams/architecture/analysis_pipeline.mmd deleted file mode 100644 index 3ee233ea..00000000 --- a/docs/diagrams/architecture/analysis_pipeline.mmd +++ /dev/null @@ -1,12 +0,0 @@ -flowchart TD - Client["Caller (e.g. Documentation Generator)"] --> AS["AnalysisService"] - AS -->|"clone_repository()"| Clone["Repository Cloning (analysis.cloning)"] - AS -->|"analyze_repository_structure()"| RA["RepoAnalyzer"] - AS -->|"analyze_code_files()"| CGA["CallGraphAnalyzer"] - RA -->|"file_tree"| CGA - CGA -->|"extract_code_files()"| Extract["Code File Extraction"] - CGA -->|"dispatch by language"| Analyzers["Language Analyzers"] - Analyzers --> TS["Tree-Sitter Analyzers module"] - CGA -->|"Node / CallRelationship"| Models["Dependency Analyzer Models module"] - AS -->|"AnalysisResult"| Models - AS -->|"Repository"| Models diff --git a/docs/diagrams/architecture/backend-core-2.mmd b/docs/diagrams/architecture/backend-core-2.mmd deleted file mode 100644 index 6e055d43..00000000 --- a/docs/diagrams/architecture/backend-core-2.mmd +++ /dev/null @@ -1,18 +0,0 @@ -sequenceDiagram - participant Caller - participant Generator as "DocumentationGenerator" - participant Builder as "DependencyGraphBuilder" - participant Agent as "AgentOrchestrator" - participant Tools as "Agent Tools" - participant Output as "Documentation Files" - - Caller->>Generator: run() - Generator->>Builder: build_dependency_graph() - Builder-->>Generator: components and leaf nodes - Generator->>Generator: cluster modules and order leaves first - Generator->>Agent: process leaf module - Agent->>Tools: inspect repository and write docs - Tools-->>Agent: tool results - Agent-->>Generator: module documentation complete - Generator->>Output: write parent overviews and metadata - Generator-->>Caller: documentation complete diff --git a/docs/diagrams/architecture/backend-core-3.mmd b/docs/diagrams/architecture/backend-core-3.mmd deleted file mode 100644 index cdcb1dfd..00000000 --- a/docs/diagrams/architecture/backend-core-3.mmd +++ /dev/null @@ -1,10 +0,0 @@ -flowchart LR - Models["Dependency Analyzer Models"] --> Analyzer["Dependency Analyzer Core"] - Language["Tree-sitter Analyzers"] --> Analyzer - Logging["Logging Config"] -.-> Analyzer - - Analyzer --> Generator["Documentation Generator"] - LLM["LLM Services"] --> Generator - Generator --> Orchestrator["AgentOrchestrator"] - Orchestrator --> Tools["Agent Tools Core"] - Tools --> Output["Markdown Documentation"] diff --git a/docs/diagrams/architecture/backend-core.mmd b/docs/diagrams/architecture/backend-core.mmd deleted file mode 100644 index 41fed11d..00000000 --- a/docs/diagrams/architecture/backend-core.mmd +++ /dev/null @@ -1,20 +0,0 @@ -flowchart TD - Input["Source Repository"] --> Generator["DocumentationGenerator"] - Generator --> GraphBuilder["DependencyGraphBuilder"] - GraphBuilder --> Parser["DependencyParser"] - Parser --> Analysis["AnalysisService"] - Analysis --> RepoAnalyzer["RepoAnalyzer"] - Analysis --> CallAnalyzer["CallGraphAnalyzer"] - CallAnalyzer --> LanguageAnalyzers["Tree-sitter and Python Analyzers"] - LanguageAnalyzers --> Models["Dependency Analyzer Models"] - - GraphBuilder --> Components["Component Dependency Graph"] - Components --> Generator - - Generator --> Orchestrator["AgentOrchestrator"] - Orchestrator --> AgentTools["Agent Tools Core"] - Orchestrator --> LLM["CountingFallbackModel"] - LLM --> Provider["Configured LLM Provider"] - - AgentTools --> Docs["Generated Markdown Documentation"] - Generator --> Docs diff --git a/docs/diagrams/architecture/c_family_analyzers-2.mmd b/docs/diagrams/architecture/c_family_analyzers-2.mmd deleted file mode 100644 index 2e0228b8..00000000 --- a/docs/diagrams/architecture/c_family_analyzers-2.mmd +++ /dev/null @@ -1,17 +0,0 @@ -sequenceDiagram - participant Caller as "analyze_cpp_file()" - participant Analyzer as "TreeSitterCppAnalyzer" - participant TS as "tree-sitter Parser" - participant Pass1 as "_extract_nodes()" - participant Pass2 as "_extract_relationships()" - - Caller->>Analyzer: "__init__(file_path, content, repo_path)" - Analyzer->>TS: "parse(content)" - TS-->>Analyzer: "syntax tree" - Analyzer->>Pass1: "traverse(root, top_level_nodes, lines)" - Pass1->>Pass1: "detect class/struct/function/namespace/variable" - Pass1-->>Analyzer: "populated top_level_nodes + nodes list" - Analyzer->>Pass2: "traverse(root, top_level_nodes)" - Pass2->>Pass2: "detect calls, inheritance, instantiation, usage" - Pass2-->>Analyzer: "call_relationships list" - Analyzer-->>Caller: "nodes, call_relationships" diff --git a/docs/diagrams/architecture/c_family_analyzers-3.mmd b/docs/diagrams/architecture/c_family_analyzers-3.mmd deleted file mode 100644 index 9d42be91..00000000 --- a/docs/diagrams/architecture/c_family_analyzers-3.mmd +++ /dev/null @@ -1,50 +0,0 @@ -classDiagram - class TreeSitterCAnalyzer { - +file_path - +content - +repo_path - +nodes - +call_relationships - -_analyze() - -_extract_nodes() - -_extract_relationships() - } - class TreeSitterCppAnalyzer { - +file_path - +content - +repo_path - +nodes - +call_relationships - -_analyze() - -_extract_nodes() - -_extract_relationships() - } - class TreeSitterCSharpAnalyzer { - +file_path - +content - +repo_path - +nodes - +call_relationships - -_analyze() - -_extract_nodes() - -_extract_relationships() - } - class Node { - +id - +name - +component_type - +file_path - +source_code - } - class CallRelationship { - +caller - +callee - +call_line - +is_resolved - } - TreeSitterCAnalyzer --> Node : produces - TreeSitterCAnalyzer --> CallRelationship : produces - TreeSitterCppAnalyzer --> Node : produces - TreeSitterCppAnalyzer --> CallRelationship : produces - TreeSitterCSharpAnalyzer --> Node : produces - TreeSitterCSharpAnalyzer --> CallRelationship : produces diff --git a/docs/diagrams/architecture/c_family_analyzers-4.mmd b/docs/diagrams/architecture/c_family_analyzers-4.mmd deleted file mode 100644 index 843a0ffd..00000000 --- a/docs/diagrams/architecture/c_family_analyzers-4.mmd +++ /dev/null @@ -1,13 +0,0 @@ -flowchart LR - Repo["Repository Source Files"] --> Parser["Dependency Parser"] - Parser -->|".c / .h"| CAnalyzer["TreeSitterCAnalyzer"] - Parser -->|".cpp / .hpp"| CppAnalyzer["TreeSitterCppAnalyzer"] - Parser -->|".cs"| CSharpAnalyzer["TreeSitterCSharpAnalyzer"] - CAnalyzer --> RawNodes["Raw Nodes + Relationships"] - CppAnalyzer --> RawNodes - CSharpAnalyzer --> RawNodes - RawNodes --> GraphBuilder["Dependency Graph Builder"] - RawNodes --> CallGraph["Call Graph Analyzer"] - CallGraph --> Resolved["Resolved Cross-File Relationships"] - GraphBuilder --> FinalGraph["Repository Dependency Graph"] - Resolved --> FinalGraph diff --git a/docs/diagrams/architecture/c_family_analyzers.mmd b/docs/diagrams/architecture/c_family_analyzers.mmd deleted file mode 100644 index e144dcff..00000000 --- a/docs/diagrams/architecture/c_family_analyzers.mmd +++ /dev/null @@ -1,8 +0,0 @@ -flowchart TD - Source["Source File Content"] --> Parser["tree-sitter Parser"] - Parser --> Tree["Concrete Syntax Tree"] - Tree --> Extract["Node Extraction Pass"] - Extract --> TopLevel["Top-Level Node Registry"] - TopLevel --> Rel["Relationship Extraction Pass"] - Rel --> Nodes["List of Node objects"] - Rel --> Calls["List of CallRelationship objects"] diff --git a/docs/diagrams/architecture/cli-core.mmd b/docs/diagrams/architecture/cli-core.mmd deleted file mode 100644 index 9eda5120..00000000 --- a/docs/diagrams/architecture/cli-core.mmd +++ /dev/null @@ -1,27 +0,0 @@ -flowchart TD - User["CLI User"] -->|"codewiki generate"| ConfigMgr["ConfigManager"] - ConfigMgr -->|"loads/saves"| ConfigFile[("~/.codewiki/config.json")] - ConfigMgr -->|"stores API keys"| Keyring[("OS Keyring")] - ConfigMgr -->|"provides"| Config["Configuration + AgentInstructions"] - - Config -->|"to_backend_config()"| CLIDocGen["CLIDocumentationGenerator"] - GitMgr["GitManager"] -->|"branch/commit info"| CLIDocGen - - CLIDocGen -->|"drives"| BackendGen["Backend DocumentationGenerator"] - CLIDocGen -->|"tracks progress via"| Progress["ProgressTracker"] - CLIDocGen -->|"builds"| Job["DocumentationJob"] - CLIDocGen -->|"optional"| HTMLGen["HTMLGenerator"] - - HTMLGen -->|"reads"| ModuleTree[("module_tree.json")] - HTMLGen -->|"reads"| Metadata[("metadata.json")] - HTMLGen -->|"writes"| IndexHTML[("index.html")] - - Job -->|"serializes to"| Metadata - - Logger["CLILogger"] -.->|"used by"| CLIDocGen - Logger -.->|"used by"| ConfigMgr - Logger -.->|"used by"| GitMgr - - BackendGen -->|"belongs to"| BackendCore["Backend Core module"] - - style BackendCore fill:#eee,stroke:#999,stroke-dasharray: 5 5 diff --git a/docs/diagrams/architecture/config-core-2.mmd b/docs/diagrams/architecture/config-core-2.mmd deleted file mode 100644 index 5a7d4885..00000000 --- a/docs/diagrams/architecture/config-core-2.mmd +++ /dev/null @@ -1,9 +0,0 @@ -flowchart TD - A["Config instance"] --> B["to_dict(include_secrets=False)"] - B --> C["Plain dict, no API keys"] - A --> D["to_dict(include_secrets=True)"] - D --> E["Plain dict, includes API keys"] - C --> F["from_dict(data)"] - E --> F - F -->|"missing secrets"| G["TypeError: missing required field"] - F -->|"secrets present"| H["Reconstructed Config"] diff --git a/docs/diagrams/architecture/config-core-3.mmd b/docs/diagrams/architecture/config-core-3.mmd deleted file mode 100644 index b1b39a47..00000000 --- a/docs/diagrams/architecture/config-core-3.mmd +++ /dev/null @@ -1,10 +0,0 @@ -flowchart TD - AI["agent_instructions dict"] --> IP["include_patterns"] - AI --> EP["exclude_patterns"] - AI --> FM["focus_modules"] - AI --> DT["doc_type"] - AI --> CI["custom_instructions"] - DT --> GPA["get_prompt_addition()"] - FM --> GPA - CI -->|"escape_format_braces"| GPA - GPA --> Prompt["Combined prompt-addition string"] diff --git a/docs/diagrams/architecture/config-core-4.mmd b/docs/diagrams/architecture/config-core-4.mmd deleted file mode 100644 index 5cf96da4..00000000 --- a/docs/diagrams/architecture/config-core-4.mmd +++ /dev/null @@ -1,12 +0,0 @@ -flowchart TD - Start["Config.validate_source_paths()"] --> CheckPrimary{{"repo_path exists and is dir?"}} - CheckPrimary -->|"no"| Err1["raise ValueError"] - CheckPrimary -->|"yes"| HasAdditional{{"additional_source_paths set?"}} - HasAdditional -->|"no"| Done["Validation OK single-path mode"] - HasAdditional -->|"yes"| Loop["For each additional path"] - Loop --> CheckExists{{"path exists and is dir?"}} - CheckExists -->|"no"| Err2["raise ValueError"] - CheckExists -->|"yes"| CheckRead{{"path readable?"}} - CheckRead -->|"no"| Err3["raise OSError"] - CheckRead -->|"yes"| Loop - Loop --> Done2["Validation OK multi-path mode"] diff --git a/docs/diagrams/architecture/config-core-5.mmd b/docs/diagrams/architecture/config-core-5.mmd deleted file mode 100644 index 584684bd..00000000 --- a/docs/diagrams/architecture/config-core-5.mmd +++ /dev/null @@ -1,21 +0,0 @@ -flowchart TD - subgraph CLIFlow["CLI Core entry point"] - CM["ConfigManager: persisted JSON plus keyring"] - CM -->|"from_config_manager"| FC1["Config.from_cli(...)"] - end - subgraph WebFlow["Frontend Core entry point"] - BW["BackgroundWorker for web job"] - BW -->|"from_web_job"| FA1["Config.from_args wrapping Namespace"] - end - subgraph EnvFlow["Environment-driven CLI entry point"] - ArgParse["argparse.Namespace from CLI arguments"] - ArgParse -->|"from_args"| FA2["Reads MAIN_MODEL, FALLBACK_MODEL, CLUSTER_API_KEY, MAIN_API_KEY, FALLBACK_API_KEY from environment"] - end - subgraph DirectFlow["Direct parameter entry point"] - Caller["Any caller with explicit parameters"] - Caller -->|"from_cli"| Validate["Validation block: keys, urls, types, ranges, max_token_field enum"] - Validate --> VSP["validate_source_paths()"] - VSP --> Instance["Config instance"] - end - FA1 --> FA2 - FC1 --> Validate diff --git a/docs/diagrams/architecture/config-core-6.mmd b/docs/diagrams/architecture/config-core-6.mmd deleted file mode 100644 index 76efd71d..00000000 --- a/docs/diagrams/architecture/config-core-6.mmd +++ /dev/null @@ -1,16 +0,0 @@ -sequenceDiagram - participant CLI as "CLI Core (ConfigManager)" - participant Config as "Config.from_config_manager" - participant FromCli as "Config.from_cli" - participant Validate as "validate_source_paths" - - CLI->>Config: from_config_manager(manager, repo_path, output_dir) - Config->>CLI: manager.get_config() - Config->>CLI: get_cluster_api_key / get_main_api_key / get_fallback_api_key - Config->>Config: check models and keys are present - Config->>FromCli: from_cli(repo_path, output_dir, models, keys, urls, tokens, temps) - FromCli->>FromCli: validation block keys urls types ranges enums - FromCli->>Validate: validate_source_paths() - Validate-->>FromCli: OK or raises ValueError or OSError - FromCli-->>Config: Config instance - Config-->>CLI: Config instance diff --git a/docs/diagrams/architecture/config-core.mmd b/docs/diagrams/architecture/config-core.mmd deleted file mode 100644 index 8a87bba8..00000000 --- a/docs/diagrams/architecture/config-core.mmd +++ /dev/null @@ -1,43 +0,0 @@ -classDiagram - class Config { - +str repo_path - +str output_dir - +str dependency_graph_dir - +str docs_dir - +int max_depth - +str main_model - +str cluster_model - +str fallback_model - +str cluster_api_key - +str main_api_key - +str fallback_api_key - +Optional~str~ cluster_base_url - +Optional~str~ main_base_url - +Optional~str~ fallback_base_url - +int cluster_max_tokens - +int main_max_tokens - +int fallback_max_tokens - +int max_token_per_module - +int max_token_per_leaf_module - +float cluster_temperature - +float main_temperature - +float fallback_temperature - +Optional~Dict~ agent_instructions - +Optional~str~ diagrams_dir - +Optional~List~str~~ additional_source_paths - +to_dict(include_secrets) Dict - +from_dict(data)$ Config - +include_patterns() Optional~List~ - +exclude_patterns() Optional~List~ - +focus_modules() Optional~List~ - +doc_type() Optional~str~ - +custom_instructions() Optional~str~ - +all_source_paths() List~str~ - +validate_source_paths() void - +is_multi_path_mode() bool - +get_prompt_addition() str - +from_args(args)$ Config - +from_web_job(repo_path, docs_dir)$ Config - +from_cli(...)$ Config - +from_config_manager(manager, repo_path, output_dir)$ Config - } diff --git a/docs/diagrams/architecture/configuration-2.mmd b/docs/diagrams/architecture/configuration-2.mmd deleted file mode 100644 index 07b6b7d6..00000000 --- a/docs/diagrams/architecture/configuration-2.mmd +++ /dev/null @@ -1,42 +0,0 @@ -classDiagram - class Configuration { - +str main_model - +str cluster_model - +str fallback_model - +str default_output - +str cluster_base_url - +str main_base_url - +str fallback_base_url - +int cluster_max_tokens - +int main_max_tokens - +int fallback_max_tokens - +float cluster_temperature - +float main_temperature - +float fallback_temperature - +bool cluster_temperature_supported - +bool main_temperature_supported - +bool fallback_temperature_supported - +int max_token_per_module - +int max_token_per_leaf_module - +int max_depth - +AgentInstructions agent_instructions - +validate() - +to_dict() dict - +from_dict(data) Configuration - +is_complete() bool - +to_backend_config(...) Config - } - - class AgentInstructions { - +List~str~ include_patterns - +List~str~ exclude_patterns - +List~str~ focus_modules - +str doc_type - +str custom_instructions - +to_dict() dict - +from_dict(data) AgentInstructions - +is_empty() bool - +get_prompt_addition() str - } - - Configuration "1" *-- "1" AgentInstructions : agent_instructions diff --git a/docs/diagrams/architecture/configuration-3.mmd b/docs/diagrams/architecture/configuration-3.mmd deleted file mode 100644 index 8b326c48..00000000 --- a/docs/diagrams/architecture/configuration-3.mmd +++ /dev/null @@ -1,20 +0,0 @@ -classDiagram - class ConfigManager { - -Optional~str~ _cluster_api_key - -Optional~str~ _main_api_key - -Optional~str~ _fallback_api_key - -Optional~Configuration~ _config - -bool _keyring_available - +load() bool - +save(...) void - +get_cluster_api_key() Optional~str~ - +get_main_api_key() Optional~str~ - +get_fallback_api_key() Optional~str~ - +get_config() Optional~Configuration~ - +is_configured() bool - +delete_api_keys() void - +clear() void - +keyring_available bool - +config_file_path Path - } - ConfigManager --> Configuration : loads/saves diff --git a/docs/diagrams/architecture/configuration-4.mmd b/docs/diagrams/architecture/configuration-4.mmd deleted file mode 100644 index 7ecb115e..00000000 --- a/docs/diagrams/architecture/configuration-4.mmd +++ /dev/null @@ -1,20 +0,0 @@ -sequenceDiagram - participant Caller - participant CM as ConfigManager - participant FS as "config.json" - participant KR as "System Keyring" - - Caller->>CM: load() - CM->>FS: check exists / read - alt file missing - FS-->>CM: not found - CM-->>Caller: False - else file present - FS-->>CM: JSON content - CM->>CM: Configuration.from_dict(data) - CM->>KR: get_password(cluster_api_key) - CM->>KR: get_password(main_api_key) - CM->>KR: get_password(fallback_api_key) - KR-->>CM: key values (or None) - CM-->>Caller: True - end diff --git a/docs/diagrams/architecture/configuration-5.mmd b/docs/diagrams/architecture/configuration-5.mmd deleted file mode 100644 index 14c17fb6..00000000 --- a/docs/diagrams/architecture/configuration-5.mmd +++ /dev/null @@ -1,16 +0,0 @@ -sequenceDiagram - participant Caller - participant CM as ConfigManager - participant FS as "config.json" - participant KR as "System Keyring" - - Caller->>CM: save(fields..., api_keys...) - CM->>FS: ensure_directory(CONFIG_DIR) - alt no in-memory config - CM->>CM: load() existing or create default Configuration - end - CM->>CM: apply provided field updates - CM->>CM: Configuration.validate() (if models set) - CM->>KR: set_password(...) for each provided API key - CM->>FS: write JSON (version + Configuration.to_dict()) - CM-->>Caller: done (or raises ConfigurationError) diff --git a/docs/diagrams/architecture/configuration-6.mmd b/docs/diagrams/architecture/configuration-6.mmd deleted file mode 100644 index 05559834..00000000 --- a/docs/diagrams/architecture/configuration-6.mmd +++ /dev/null @@ -1,8 +0,0 @@ -flowchart LR - A["CLI Command"] --> B["ConfigManager.load()"] - B --> C["Configuration"] - C --> D["Configuration.to_backend_config()"] - D --> E["ConfigManager (keyring lookup for missing keys)"] - D --> F["Merge AgentInstructions (runtime over persisted)"] - D --> G["Config.from_cli(...)"] - G --> H["Backend Config"] diff --git a/docs/diagrams/architecture/configuration.mmd b/docs/diagrams/architecture/configuration.mmd deleted file mode 100644 index aa110208..00000000 --- a/docs/diagrams/architecture/configuration.mmd +++ /dev/null @@ -1,16 +0,0 @@ -flowchart TD - subgraph ConfigModule["Configuration Module"] - CM["ConfigManager"] - Cfg["Configuration"] - AI["AgentInstructions"] - end - - FS["~/.codewiki/config.json"] - KR["System Keyring"] - Backend["Backend Config"] - - CM -->|"load() / save()"| FS - CM -->|"get/set API keys"| KR - CM -->|"holds"| Cfg - Cfg -->|"has one"| AI - Cfg -->|"to_backend_config()"| Backend diff --git a/docs/diagrams/architecture/dependency-analyzer-core-2.mmd b/docs/diagrams/architecture/dependency-analyzer-core-2.mmd deleted file mode 100644 index cc2c308c..00000000 --- a/docs/diagrams/architecture/dependency-analyzer-core-2.mmd +++ /dev/null @@ -1,22 +0,0 @@ -sequenceDiagram - participant Builder as "DependencyGraphBuilder" - participant Parser as "DependencyParser" - participant Service as "AnalysisService" - participant RepoA as "RepoAnalyzer" - participant CallA as "CallGraphAnalyzer" - participant TS as "Tree-sitter Analyzers" - - Builder->>Parser: parse_repository() - Parser->>Service: _analyze_structure(repo_path) - Service->>RepoA: analyze_repository_structure() - RepoA-->>Service: file_tree + summary - Parser->>Service: _analyze_call_graph(file_tree, repo_path) - Service->>CallA: extract_code_files() / analyze_code_files() - CallA->>TS: analyze_python_file / analyze_javascript_file_treesitter / ... - TS-->>CallA: functions, relationships - CallA-->>Service: functions, relationships, visualization - Service-->>Parser: call_graph_result - Parser->>Parser: build namespaced Node components - Parser-->>Builder: components (Dict[str, Node]) - Builder->>Builder: build_graph_from_components / validate / filter leaves - Builder-->>Builder: (components, leaf_nodes) diff --git a/docs/diagrams/architecture/dependency-analyzer-core.mmd b/docs/diagrams/architecture/dependency-analyzer-core.mmd deleted file mode 100644 index 706eba5f..00000000 --- a/docs/diagrams/architecture/dependency-analyzer-core.mmd +++ /dev/null @@ -1,21 +0,0 @@ -flowchart TD - subgraph AnalysisPipeline["Analysis Pipeline"] - direction TB - RepoAnalyzer["RepoAnalyzer"] -->|"file_tree"| AnalysisService["AnalysisService"] - CallGraphAnalyzer["CallGraphAnalyzer"] -->|"functions + relationships"| AnalysisService - end - - subgraph GraphConstruction["Graph Construction"] - direction TB - DependencyParser["DependencyParser"] -->|"components (Node)"| DependencyGraphBuilder["DependencyGraphBuilder"] - end - - Config["Config"] -.->|"repo paths, patterns"| DependencyGraphBuilder - DependencyGraphBuilder -->|"drives"| DependencyParser - DependencyParser -->|"uses"| AnalysisService - CallGraphAnalyzer -->|"delegates per-language"| TreeSitterAnalyzers["Tree-sitter Analyzers"] - AnalysisService -->|"produces"| Models["Node / CallRelationship / AnalysisResult"] - DependencyGraphBuilder -->|"writes"| GraphFile[("dependency_graph.json")] - - click TreeSitterAnalyzers "../tree-sitter-analyzers/tree-sitter-analyzers.md" - click Models "../dependency-analyzer-models/dependency-analyzer-models.md" diff --git a/docs/diagrams/architecture/dependency-analyzer-models-2.mmd b/docs/diagrams/architecture/dependency-analyzer-models-2.mmd deleted file mode 100644 index 2cff6eb2..00000000 --- a/docs/diagrams/architecture/dependency-analyzer-models-2.mmd +++ /dev/null @@ -1,60 +0,0 @@ -classDiagram - class Node { - +str id - +str name - +str component_type - +str file_path - +str relative_path - +Set~str~ depends_on - +Optional~str~ source_code - +int start_line - +int end_line - +bool has_docstring - +str docstring - +Optional~List~str~~ parameters - +Optional~str~ node_type - +Optional~List~str~~ base_classes - +Optional~str~ class_name - +Optional~str~ display_name - +Optional~str~ component_id - +str short_id - +str namespace - +bool is_from_deps - +get_display_name() str - } - - class CallRelationship { - +str caller - +str callee - +Optional~int~ call_line - +bool is_resolved - } - - class Repository { - +str url - +str name - +str clone_path - +str analysis_id - } - - class AnalysisResult { - +Repository repository - +List~Node~ functions - +List~CallRelationship~ relationships - +Dict file_tree - +Dict summary - +Dict visualization - +Optional~str~ readme_content - } - - class NodeSelection { - +List~str~ selected_nodes - +bool include_relationships - +Dict~str,str~ custom_names - } - - AnalysisResult "1" --> "1" Repository : describes - AnalysisResult "1" --> "*" Node : contains - AnalysisResult "1" --> "*" CallRelationship : contains - CallRelationship "*" --> "1" Node : caller/callee reference (by id) - NodeSelection "1" --> "*" Node : references (by id) diff --git a/docs/diagrams/architecture/dependency-analyzer-models.mmd b/docs/diagrams/architecture/dependency-analyzer-models.mmd deleted file mode 100644 index 2aeb4b4e..00000000 --- a/docs/diagrams/architecture/dependency-analyzer-models.mmd +++ /dev/null @@ -1,14 +0,0 @@ -flowchart TD - Analyzers["Tree-sitter Analyzers"] -->|"produce"| Node["Node"] - Analyzers -->|"produce"| CallRel["CallRelationship"] - Parser["Dependency Parser / Graph Builder"] -->|"assembles"| Node - Parser -->|"assembles"| CallRel - RepoModel["Repository"] -->|"describes source of"| Node - AnalysisSvc["Analysis Service / Repo Analyzer / Call Graph Analyzer"] -->|"aggregates"| Node - AnalysisSvc -->|"aggregates"| CallRel - AnalysisSvc -->|"aggregates"| RepoModel - AnalysisSvc -->|"produces"| Result["AnalysisResult"] - Selection["NodeSelection"] -->|"filters"| Result - Result -->|"consumed by"| DocGen["Documentation Generator"] - Result -->|"consumed by"| Agent["Agent Orchestrator"] - Result -->|"consumed by"| CLIGen["CLI Documentation Generator"] diff --git a/docs/diagrams/architecture/documentation-generator-2.mmd b/docs/diagrams/architecture/documentation-generator-2.mmd deleted file mode 100644 index 6fefa10b..00000000 --- a/docs/diagrams/architecture/documentation-generator-2.mmd +++ /dev/null @@ -1,43 +0,0 @@ -sequenceDiagram - participant Caller as "CLI/Web Caller" - participant DG as "DocumentationGenerator" - participant GB as "DependencyGraphBuilder" - participant Cluster as "cluster_modules" - participant AO as "AgentOrchestrator" - participant LLM as "call_llm" - participant FS as "file_manager" - - Caller->>DG: run() - DG->>GB: build_dependency_graph() - GB-->>DG: components, leaf_nodes - alt "module tree cache exists" - DG->>FS: load_json(first_module_tree.json) - else "no cache" - DG->>Cluster: cluster_modules(leaf_nodes, components, config) - Cluster-->>DG: module_tree - DG->>FS: save_json(first_module_tree.json) - end - alt "module_tree empty but leaf_nodes exist" - DG->>DG: build synthetic modules by top-level directory - DG->>FS: save_json(synthetic module_tree) - end - DG->>FS: save_json(module_tree.json) - DG->>DG: generate_module_documentation(components, leaf_nodes) - loop "for each module in leaf-first order" - alt "is_leaf_module" - DG->>AO: process_module(name, components, ids, path, dir) - AO->>LLM: agent.run(user_prompt) - LLM-->>AO: markdown docs - AO->>FS: save module_name.md - else "parent module" - DG->>DG: build_overview_structure(...) - DG->>LLM: call_llm(MODULE_OVERVIEW_PROMPT) - LLM-->>DG: "..." - DG->>FS: save parent module_name.md - end - end - DG->>DG: generate_parent_module_docs([], working_dir) - DG->>LLM: call_llm(REPO_OVERVIEW_PROMPT) - DG->>FS: save overview.md - DG->>DG: create_documentation_metadata(working_dir, components, len(leaf_nodes)) - DG-->>Caller: "documentation complete" diff --git a/docs/diagrams/architecture/documentation-generator-3.mmd b/docs/diagrams/architecture/documentation-generator-3.mmd deleted file mode 100644 index 64ae9844..00000000 --- a/docs/diagrams/architecture/documentation-generator-3.mmd +++ /dev/null @@ -1,13 +0,0 @@ -flowchart LR - subgraph Tree["Module Tree Example"] - Root["Backend"] --> Auth["Authentication"] - Root --> API["API Layer"] - Auth --> JWT["JWT"] - Auth --> OAuth["OAuth"] - end - subgraph Order["Processing Order"] - O1["1: JWT (leaf)"] --> O2["2: OAuth (leaf)"] - O2 --> O3["3: Authentication (parent)"] - O3 --> O4["4: API Layer (leaf/parent)"] - O4 --> O5["5: Backend (parent)"] - end diff --git a/docs/diagrams/architecture/documentation-generator.mmd b/docs/diagrams/architecture/documentation-generator.mmd deleted file mode 100644 index de7a1fdb..00000000 --- a/docs/diagrams/architecture/documentation-generator.mmd +++ /dev/null @@ -1,11 +0,0 @@ -flowchart TD - CLI["CLI Core: CLIDocumentationGenerator"] --> DG["Documentation Generator: DocumentationGenerator"] - Web["Frontend Core: BackgroundWorker"] --> DG - DG --> GraphBuilder["Dependency Analyzer Core: DependencyGraphBuilder"] - DG --> Cluster["cluster_modules"] - DG --> Orchestrator["Backend Core: AgentOrchestrator"] - DG --> LLM["LLM Services: call_llm / CountingFallbackModel"] - GraphBuilder --> Analyzers["Tree-sitter Analyzers"] - GraphBuilder --> Models["Dependency Analyzer Models"] - Orchestrator --> AgentTools["Agent Tools Core"] - DG --> ConfigCore["Config Core: Config"] diff --git a/docs/diagrams/architecture/frontend-core-2.mmd b/docs/diagrams/architecture/frontend-core-2.mmd deleted file mode 100644 index 9aba09d9..00000000 --- a/docs/diagrams/architecture/frontend-core-2.mmd +++ /dev/null @@ -1,24 +0,0 @@ -sequenceDiagram - participant Browser - participant Routes as "WebRoutes" - participant Worker as "BackgroundWorker" - participant Cache as "CacheManager" - participant Git as "GitHubRepoProcessor" - participant DocGen as "DocumentationGenerator" - - Browser->>Routes: POST / (repo_url, commit_id) - Routes->>Git: is_valid_github_url / get_repo_info - Routes->>Cache: get_cached_docs(repo_url) - alt Cache hit - Cache-->>Routes: docs_path - Routes-->>Browser: Render success message - else Cache miss - Routes->>Worker: add_job(job_id, JobStatus) - Worker-->>Routes: queued - Routes-->>Browser: Render "queued" message - Worker->>Git: clone_repository(clone_url, temp_dir, commit_id) - Worker->>DocGen: run() (async documentation generation) - DocGen-->>Worker: docs_dir populated - Worker->>Cache: add_to_cache(repo_url, docs_path) - Worker->>Worker: save_job_statuses() - end diff --git a/docs/diagrams/architecture/frontend-core-3.mmd b/docs/diagrams/architecture/frontend-core-3.mmd deleted file mode 100644 index e4ac9525..00000000 --- a/docs/diagrams/architecture/frontend-core-3.mmd +++ /dev/null @@ -1,51 +0,0 @@ -classDiagram - class WebAppConfig { - +CACHE_DIR - +TEMP_DIR - +QUEUE_SIZE - +CACHE_EXPIRY_DAYS - +CLONE_TIMEOUT - +ensure_directories() - } - class GitHubRepoProcessor { - +is_valid_github_url(url) - +get_repo_info(url) - +clone_repository(clone_url, target_dir, commit_id) - } - class CacheManager { - +cache_index - +get_cached_docs(repo_url) - +add_to_cache(repo_url, docs_path) - +remove_from_cache(repo_url) - } - class BackgroundWorker { - +job_status - +processing_queue - +start() - +add_job(job_id, job) - +get_job_status(job_id) - } - class WebRoutes { - +index_get(request) - +index_post(request, repo_url, commit_id) - +get_job_status(job_id) - +serve_generated_docs(job_id, filename) - } - class JobStatus - class CacheEntry - class RepositorySubmission - class JobStatusResponse - class StringTemplateLoader - - WebRoutes --> BackgroundWorker - WebRoutes --> CacheManager - WebRoutes --> GitHubRepoProcessor - WebRoutes --> StringTemplateLoader - BackgroundWorker --> CacheManager - BackgroundWorker --> GitHubRepoProcessor - BackgroundWorker --> JobStatus - CacheManager --> CacheEntry - WebRoutes --> JobStatusResponse - GitHubRepoProcessor --> WebAppConfig - CacheManager --> WebAppConfig - BackgroundWorker --> WebAppConfig diff --git a/docs/diagrams/architecture/frontend-core-4.mmd b/docs/diagrams/architecture/frontend-core-4.mmd deleted file mode 100644 index 3be8bfef..00000000 --- a/docs/diagrams/architecture/frontend-core-4.mmd +++ /dev/null @@ -1,8 +0,0 @@ -stateDiagram-v2 - [*] --> queued: "add_job()" - queued --> processing: "worker picks up job" - processing --> completed: "cache hit OR generation succeeds" - processing --> failed: "clone or generation error" - failed --> queued: "resubmission after cooldown" - completed --> [*] - failed --> [*] diff --git a/docs/diagrams/architecture/frontend-core.mmd b/docs/diagrams/architecture/frontend-core.mmd deleted file mode 100644 index e36980d6..00000000 --- a/docs/diagrams/architecture/frontend-core.mmd +++ /dev/null @@ -1,25 +0,0 @@ -flowchart TD - User["Browser / API Client"] -->|"submit repo_url"| Routes["WebRoutes"] - Routes -->|"validate URL"| GitProc["GitHubRepoProcessor"] - Routes -->|"check cache"| Cache["CacheManager"] - Routes -->|"enqueue job"| Worker["BackgroundWorker"] - Routes -->|"render HTML"| Templates["StringTemplateLoader / render_template"] - - Worker -->|"clone repository"| GitProc - Worker -->|"build Config.from_web_job"| ConfigCore["Config (config-core)"] - Worker -->|"generate docs"| DocGen["DocumentationGenerator (backend-core)"] - Worker -->|"store result path"| Cache - Worker -->|"persist status"| JobsFile[("jobs.json")] - - Cache -->|"persist index"| CacheFile[("cache_index.json")] - - subgraph models_group["Data Models"] - JobStatus["JobStatus"] - CacheEntry["CacheEntry"] - RepositorySubmission["RepositorySubmission"] - JobStatusResponse["JobStatusResponse"] - end - - Routes --> models_group - Worker --> models_group - Cache --> models_group diff --git a/docs/diagrams/architecture/generation-2.mmd b/docs/diagrams/architecture/generation-2.mmd deleted file mode 100644 index 670f2390..00000000 --- a/docs/diagrams/architecture/generation-2.mmd +++ /dev/null @@ -1,26 +0,0 @@ -sequenceDiagram - participant Caller as "CLI Command" - participant CDG as "CLIDocumentationGenerator" - participant BC as "BackendConfig" - participant DG as "DocumentationGenerator" - participant CM as "cluster_modules" - participant HG as "HTMLGenerator" - participant Job as "DocumentationJob" - - Caller->>CDG: generate() - CDG->>Job: start() - CDG->>BC: Config.from_cli(...) - CDG->>DG: _run_backend_generation(backend_config) - DG->>DG: graph_builder.build_dependency_graph() - DG->>CM: cluster_modules(leaf_nodes, components, config) - Note over DG,CM: Synthetic module fallback if tree is empty - DG->>DG: generate_module_documentation(components, leaf_nodes) - opt diagrams_dir configured - DG->>DG: extract_and_save_mermaid_diagrams() - end - opt generate_html is true - CDG->>HG: generate(output_path, ...) - end - CDG->>CDG: _finalize_job() - CDG->>Job: complete() - CDG-->>Caller: DocumentationJob diff --git a/docs/diagrams/architecture/generation-3.mmd b/docs/diagrams/architecture/generation-3.mmd deleted file mode 100644 index 28f2c2c3..00000000 --- a/docs/diagrams/architecture/generation-3.mmd +++ /dev/null @@ -1,14 +0,0 @@ -flowchart LR - subgraph CLIConfig["CLI config dict"] - A1["main_model / cluster_model / fallback_model"] - A2["*_api_key"] - A3["*_base_url"] - A4["*_api_version"] - A5["*_max_tokens / *_temperature"] - A6["max_token_per_module / max_depth"] - A7["agent_instructions"] - A8["additional_paths"] - end - CLIConfig --> FromCLI["Config.from_cli()"] - FromCLI --> BackendConfig["Backend Config instance"] - BackendConfig --> DG2["DocumentationGenerator"] diff --git a/docs/diagrams/architecture/generation.mmd b/docs/diagrams/architecture/generation.mmd deleted file mode 100644 index 352a888d..00000000 --- a/docs/diagrams/architecture/generation.mmd +++ /dev/null @@ -1,11 +0,0 @@ -flowchart TD - CLI["CLI Command Entry Point"] --> CDG["CLIDocumentationGenerator"] - CDG --> PT["ProgressTracker (utils)"] - CDG --> Job["DocumentationJob (job_models)"] - CDG --> BC["BackendConfig (config-core)"] - CDG --> DG["DocumentationGenerator (backend-core)"] - DG --> GB["DependencyGraphBuilder"] - DG --> AO["AgentOrchestrator"] - CDG --> CM["cluster_modules (backend-core)"] - CDG --> HG["HTMLGenerator (html_generation)"] - CDG --> LOG["ColoredFormatter (backend-core logging)"] diff --git a/docs/diagrams/architecture/git_integration-2.mmd b/docs/diagrams/architecture/git_integration-2.mmd deleted file mode 100644 index 2339eeb2..00000000 --- a/docs/diagrams/architecture/git_integration-2.mmd +++ /dev/null @@ -1,26 +0,0 @@ -sequenceDiagram - participant CLI as "CLI Entry Point" - participant GM as "GitManager" - participant DocGen as "CLIDocumentationGenerator" - participant Repo as "git.Repo" - - CLI->>GM: "GitManager(repo_path)" - GM->>Repo: "git.Repo(repo_path)" - CLI->>GM: "check_clean_working_directory()" - GM-->>CLI: "(is_clean, status_message)" - alt "Working directory dirty and not forced" - GM-->>CLI: "raise RepositoryError" - else "Clean or forced" - CLI->>GM: "create_documentation_branch(force)" - GM->>Repo: "create_head(branch_name)" - GM->>Repo: "checkout()" - GM-->>CLI: "branch_name" - CLI->>DocGen: "generate()" - DocGen-->>CLI: "DocumentationJob" - CLI->>GM: "commit_documentation(docs_path, message)" - GM->>Repo: "index.add([docs_path])" - GM->>Repo: "index.commit(message)" - GM-->>CLI: "commit_hash" - CLI->>GM: "get_github_pr_url(branch_name)" - GM-->>CLI: "pr_url or None" - end diff --git a/docs/diagrams/architecture/git_integration-3.mmd b/docs/diagrams/architecture/git_integration-3.mmd deleted file mode 100644 index da45ed9e..00000000 --- a/docs/diagrams/architecture/git_integration-3.mmd +++ /dev/null @@ -1,13 +0,0 @@ -flowchart TD - Start["create_documentation_branch(force)"] --> CheckForce{{force?}} - CheckForce -->|"false"| CheckClean["check_clean_working_directory()"] - CheckForce -->|"true"| GenName["Generate timestamped branch name"] - CheckClean --> IsClean{{"is_clean?"}} - IsClean -->|"false"| RaiseErr["raise RepositoryError with remediation steps"] - IsClean -->|"true"| GenName - GenName --> CheckExists{{"branch name exists?"}} - CheckExists -->|"true"| AppendCounter["Append incrementing counter suffix"] - AppendCounter --> CheckExists - CheckExists -->|"false"| CreateBranch["repo.create_head(branch_name)"] - CreateBranch --> Checkout["new_branch.checkout()"] - Checkout --> ReturnName["return branch_name"] diff --git a/docs/diagrams/architecture/git_integration.mmd b/docs/diagrams/architecture/git_integration.mmd deleted file mode 100644 index b8d0ca9d..00000000 --- a/docs/diagrams/architecture/git_integration.mmd +++ /dev/null @@ -1,18 +0,0 @@ -classDiagram - class GitManager { - +Path repo_path - +Repo repo - +__init__(repo_path) - +check_clean_working_directory() Tuple - +create_documentation_branch(force) str - +commit_documentation(docs_path, message) str - +get_remote_url(remote_name) str - +get_current_branch() str - +get_commit_hash() str - +branch_exists(branch_name) bool - +get_github_pr_url(branch_name) str - } - class RepositoryError { - +__init__(message) - } - GitManager ..> RepositoryError : raises diff --git a/docs/diagrams/architecture/graph_construction-2.mmd b/docs/diagrams/architecture/graph_construction-2.mmd deleted file mode 100644 index a6ed0636..00000000 --- a/docs/diagrams/architecture/graph_construction-2.mmd +++ /dev/null @@ -1,19 +0,0 @@ -flowchart TD - Start["parse_repository()"] --> Check{{"len(repo_paths) == 1?"}} - Check -->|"Yes"| Single["_parse_single_repository()"] - Check -->|"No"| Multi["_parse_multiple_repositories()"] - - Single --> S1["analysis_service._analyze_structure()"] - S1 --> S2["analysis_service._analyze_call_graph()"] - S2 --> S3["_build_components_from_analysis()"] - S3 --> SResult["self.components"] - - Multi --> M1["For each repo path: compute namespace"] - M1 --> M2["analysis_service._analyze_structure()"] - M2 --> M3["analysis_service._analyze_call_graph()"] - M3 --> M4["_build_namespaced_components()"] - M4 --> M5["Merge into all_components"] - M5 --> M6{{"More paths?"}} - M6 -->|"Yes"| M1 - M6 -->|"No"| M7["_resolve_cross_namespace_dependencies()"] - M7 --> MResult["self.components"] diff --git a/docs/diagrams/architecture/graph_construction-3.mmd b/docs/diagrams/architecture/graph_construction-3.mmd deleted file mode 100644 index 1fc2338c..00000000 --- a/docs/diagrams/architecture/graph_construction-3.mmd +++ /dev/null @@ -1,15 +0,0 @@ -flowchart TD - A["build_dependency_graph()"] --> B["Ensure dependency_graph_dir exists"] - B --> C["Compute sanitized dependency_graph_path"] - C --> D["Resolve include/exclude patterns from Config"] - D --> E["Build repo_paths from config.all_source_paths"] - E --> F["Instantiate DependencyParser"] - F --> G["parser.parse_repository()"] - G --> H["Log component type breakdown"] - H --> I["parser.save_dependency_graph(path)"] - I --> J["build_graph_from_components(components)"] - J --> K["validate_graph_completeness(components, graph)"] - K --> L["get_leaf_nodes(graph, components)"] - L --> M["Determine valid leaf types from available component types"] - M --> N["Filter leaf_nodes: skip invalid / wrong-type / not-found"] - N --> O["Return (components, keep_leaf_nodes)"] diff --git a/docs/diagrams/architecture/graph_construction-4.mmd b/docs/diagrams/architecture/graph_construction-4.mmd deleted file mode 100644 index 5e5d67a9..00000000 --- a/docs/diagrams/architecture/graph_construction-4.mmd +++ /dev/null @@ -1,8 +0,0 @@ -flowchart TD - Start["For each leaf_node in leaf_nodes"] --> V1{{"Is leaf_node a valid non-empty string without error keywords?"}} - V1 -->|"No"| SkipInvalid["skipped_invalid += 1"] - V1 -->|"Yes"| V2{{"leaf_node in components?"}} - V2 -->|"No"| SkipNotFound["skipped_not_found += 1"] - V2 -->|"Yes"| V3{{"component_type in valid_types?"}} - V3 -->|"No"| SkipType["skipped_type += 1"] - V3 -->|"Yes"| Keep["keep_leaf_nodes.append(leaf_node)"] diff --git a/docs/diagrams/architecture/graph_construction-5.mmd b/docs/diagrams/architecture/graph_construction-5.mmd deleted file mode 100644 index 68967328..00000000 --- a/docs/diagrams/architecture/graph_construction-5.mmd +++ /dev/null @@ -1,22 +0,0 @@ -sequenceDiagram - participant Caller as "Pipeline Caller" - participant Builder as "DependencyGraphBuilder" - participant Parser as "DependencyParser" - participant Svc as "AnalysisService" - participant FS as "File System" - - Caller->>Builder: build_dependency_graph() - Builder->>Parser: new DependencyParser(repo_paths, patterns) - Builder->>Parser: parse_repository() - Parser->>Svc: _analyze_structure(repo_path) - Svc-->>Parser: structure_result - Parser->>Svc: _analyze_call_graph(file_tree, repo_path) - Svc-->>Parser: call_graph_result - Parser->>Parser: _build_components_from_analysis() / _build_namespaced_components() - Parser-->>Builder: components (Dict[str, Node]) - Builder->>Parser: save_dependency_graph(path) - Parser->>FS: write JSON - Builder->>Builder: build_graph_from_components(components) - Builder->>Builder: validate_graph_completeness(components, graph) - Builder->>Builder: get_leaf_nodes(graph, components) - Builder-->>Caller: (components, leaf_nodes) diff --git a/docs/diagrams/architecture/graph_construction.mmd b/docs/diagrams/architecture/graph_construction.mmd deleted file mode 100644 index 32fb052a..00000000 --- a/docs/diagrams/architecture/graph_construction.mmd +++ /dev/null @@ -1,11 +0,0 @@ -flowchart TD - Config["Config"] --> Builder["DependencyGraphBuilder"] - Builder -->|"instantiates"| Parser["DependencyParser"] - Parser -->|"uses"| AnalysisSvc["AnalysisService"] - AnalysisSvc -->|"structure + call graph"| Parser - Parser -->|"produces"| Nodes["Node objects (Dict[str, Node])"] - Builder -->|"build_graph_from_components()"| TraversalGraph["In-memory traversal graph"] - Builder -->|"validate_graph_completeness()"| Validation["Graph validation"] - Builder -->|"get_leaf_nodes()"| LeafFilter["Leaf node filtering"] - Parser -->|"save_dependency_graph()"| JSONFile["dependency_graph.json"] - LeafFilter --> Output["(components, leaf_nodes)"] diff --git a/docs/diagrams/architecture/html_generation-2.mmd b/docs/diagrams/architecture/html_generation-2.mmd deleted file mode 100644 index 12aad20b..00000000 --- a/docs/diagrams/architecture/html_generation-2.mmd +++ /dev/null @@ -1,19 +0,0 @@ -sequenceDiagram - participant CLIGen as "CLIDocumentationGenerator" - participant HTMLGen as "HTMLGenerator" - participant FS as "safe_read / safe_write" - participant Git as "GitPython" - - CLIGen->>HTMLGen: HTMLGenerator() - CLIGen->>HTMLGen: detect_repository_info(repo_path) - HTMLGen->>Git: Repo(repo_path) - Git-->>HTMLGen: remote URL, name - HTMLGen-->>CLIGen: name, url, github_pages_url - CLIGen->>HTMLGen: generate(output_path, title, repository_url, docs_dir) - HTMLGen->>FS: safe_read(module_tree.json) - HTMLGen->>FS: safe_read(metadata.json) - HTMLGen->>FS: safe_read(viewer_template.html) - HTMLGen->>HTMLGen: _build_info_content(metadata) - HTMLGen->>HTMLGen: substitute placeholders - HTMLGen->>FS: safe_write(index.html) - HTMLGen-->>CLIGen: index.html written diff --git a/docs/diagrams/architecture/html_generation-3.mmd b/docs/diagrams/architecture/html_generation-3.mmd deleted file mode 100644 index d4be53b1..00000000 --- a/docs/diagrams/architecture/html_generation-3.mmd +++ /dev/null @@ -1,17 +0,0 @@ -flowchart TD - Start["generate() called"] --> CheckDocsDir{{"docs_dir provided?"}} - CheckDocsDir -->|"yes"| AutoLoadTree["load_module_tree(docs_dir)"] - CheckDocsDir -->|"yes"| AutoLoadMeta["load_metadata(docs_dir)"] - CheckDocsDir -->|"no"| UseProvided["use provided module_tree/metadata"] - AutoLoadTree --> Defaults - AutoLoadMeta --> Defaults - UseProvided --> Defaults - Defaults["Apply defaults for module_tree/config"] --> LoadTemplate["Load viewer_template.html"] - LoadTemplate -->|"missing"| RaiseErr["raise FileSystemError"] - LoadTemplate -->|"found"| BuildInfo["_build_info_content(metadata)"] - BuildInfo --> BuildRepoLink["Build repository link HTML"] - BuildRepoLink --> ComputeBasePath["Compute docs_base_path"] - ComputeBasePath --> SerializeJson["Serialize config/module_tree/metadata to JSON"] - SerializeJson --> Replace["Replace template placeholders"] - Replace --> WriteOut["safe_write(output_path)"] - WriteOut --> Done["index.html generated"] diff --git a/docs/diagrams/architecture/html_generation.mmd b/docs/diagrams/architecture/html_generation.mmd deleted file mode 100644 index fe3dd6cf..00000000 --- a/docs/diagrams/architecture/html_generation.mmd +++ /dev/null @@ -1,10 +0,0 @@ -flowchart TD - Docs["Documentation Output Directory"] --> ModuleTreeJson["module_tree.json"] - Docs --> MetadataJson["metadata.json"] - Template["viewer_template.html"] --> Generator["HTMLGenerator"] - ModuleTreeJson --> Generator - MetadataJson --> Generator - RepoPath["Repository Path"] --> DetectInfo["detect_repository_info()"] - DetectInfo --> Generator - Generator -->|"safe_write()"| IndexHtml["index.html"] - Generator -->|"on failure"| FSError["FileSystemError"] diff --git a/docs/diagrams/architecture/java_analyzer-2.mmd b/docs/diagrams/architecture/java_analyzer-2.mmd deleted file mode 100644 index aca32f56..00000000 --- a/docs/diagrams/architecture/java_analyzer-2.mmd +++ /dev/null @@ -1,16 +0,0 @@ -sequenceDiagram - participant Caller as "analyze_java_file()" - participant Analyzer as "TreeSitterJavaAnalyzer" - participant Parser as "tree_sitter.Parser" - participant NodesPass as "_extract_nodes()" - participant RelsPass as "_extract_relationships()" - - Caller->>Analyzer: "__init__(file_path, content, repo_path)" - Analyzer->>Analyzer: "_analyze()" - Analyzer->>Parser: "parse(content)" - Parser-->>Analyzer: "AST root node" - Analyzer->>NodesPass: "_extract_nodes(root, top_level_nodes, lines)" - NodesPass-->>Analyzer: "self.nodes populated" - Analyzer->>RelsPass: "_extract_relationships(root, top_level_nodes)" - RelsPass-->>Analyzer: "self.call_relationships populated" - Analyzer-->>Caller: "nodes, call_relationships" diff --git a/docs/diagrams/architecture/java_analyzer-3.mmd b/docs/diagrams/architecture/java_analyzer-3.mmd deleted file mode 100644 index ccd70e87..00000000 --- a/docs/diagrams/architecture/java_analyzer-3.mmd +++ /dev/null @@ -1,6 +0,0 @@ -flowchart LR - A["class_declaration
with superclass"] -->|"1: Inheritance"| R1["CallRelationship
class extends BaseClass"] - B["class/enum/record
with super_interfaces"] -->|"2: Interface Implementation"| R2["CallRelationship
class implements Interface"] - C["field_declaration"] -->|"3: Field Type Use"| R3["CallRelationship
class has field of Type"] - D["method_invocation"] -->|"4: Method Call"| R4["CallRelationship
caller calls object.method()"] - E["object_creation_expression"] -->|"5: Object Creation"| R5["CallRelationship
class creates new Type()"] diff --git a/docs/diagrams/architecture/java_analyzer.mmd b/docs/diagrams/architecture/java_analyzer.mmd deleted file mode 100644 index 0f8c5faf..00000000 --- a/docs/diagrams/architecture/java_analyzer.mmd +++ /dev/null @@ -1,19 +0,0 @@ -flowchart TD - subgraph Pipeline["Dependency Analysis Pipeline"] - CGA["CallGraphAnalyzer"] -->|"dispatches .java files"| JA["TreeSitterJavaAnalyzer"] - end - - subgraph JavaAnalyzerModule["Java Analyzer"] - JA -->|"1: parse source"| TS["tree-sitter-java grammar"] - TS -->|"AST"| EN["_extract_nodes()"] - TS -->|"AST"| ER["_extract_relationships()"] - EN -->|"appends"| Nodes["self.nodes: List[Node]"] - ER -->|"appends"| Rels["self.call_relationships: List[CallRelationship]"] - end - - Nodes -->|"consumed by"| DP["DependencyParser"] - Rels -->|"consumed by"| DP - DP -->|"builds"| Graph["Dependency Graph"] - - Node["Node model"] -.->|"defines schema for"| Nodes - CallRel["CallRelationship model"] -.->|"defines schema for"| Rels diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers-2.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers-2.mmd deleted file mode 100644 index 62ab2cab..00000000 --- a/docs/diagrams/architecture/javascript_typescript_analyzers-2.mmd +++ /dev/null @@ -1,45 +0,0 @@ -classDiagram - class TreeSitterJSAnalyzer { - +file_path: Path - +content: str - +repo_path: str - +nodes: List~Node~ - +call_relationships: List~CallRelationship~ - +top_level_nodes: dict - +analyze() None - -_extract_functions(node) None - -_extract_call_relationships(node) None - -_extract_jsdoc_type_dependencies(node, caller) None - } - class TreeSitterTSAnalyzer { - +file_path: Path - +content: str - +repo_path: str - +nodes: List~Node~ - +call_relationships: List~CallRelationship~ - +top_level_nodes: dict - +analyze() None - -_extract_all_entities(node, all_entities, depth) None - -_filter_top_level_declarations(all_entities) None - -_extract_all_relationships(node, all_entities) None - } - class Node { - +id: str - +name: str - +component_type: str - +source_code: str - +start_line: int - +end_line: int - +node_type: str - +base_classes: List~str~ - } - class CallRelationship { - +caller: str - +callee: str - +call_line: int - +is_resolved: bool - } - TreeSitterJSAnalyzer --> Node : produces - TreeSitterJSAnalyzer --> CallRelationship : produces - TreeSitterTSAnalyzer --> Node : produces - TreeSitterTSAnalyzer --> CallRelationship : produces diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers-3.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers-3.mmd deleted file mode 100644 index 82f1d2c5..00000000 --- a/docs/diagrams/architecture/javascript_typescript_analyzers-3.mmd +++ /dev/null @@ -1,19 +0,0 @@ -flowchart TD - Start["analyze()"] --> Parse["Parser.parse(content)"] - Parse --> ExtractFn["_extract_functions(root_node)"] - ExtractFn --> TraverseFn["_traverse_for_functions(node)"] - TraverseFn -->|"class/abstract/interface"| ExtractClass["_extract_class_declaration()"] - ExtractClass --> ExtractMethods["_extract_methods_from_class()"] - TraverseFn -->|"function_declaration"| ExtractFunc["_extract_function_declaration()"] - TraverseFn -->|"export_statement"| ExtractExport["_extract_exported_function()"] - TraverseFn -->|"lexical_declaration"| ExtractArrow["_extract_arrow_function_from_declaration()"] - ExtractFn --> ExtractCalls["_extract_call_relationships(root_node)"] - ExtractCalls --> TraverseCalls["_traverse_for_calls(node, current_top_level)"] - TraverseCalls -->|"call_expression"| CallRel["_extract_call_from_node()"] - TraverseCalls -->|"new_expression"| NewRel["Constructor relationship"] - TraverseCalls -->|"class_heritage"| InheritRel["Inheritance relationship"] - TraverseCalls -->|"JSDoc comment"| JsdocRel["_parse_jsdoc_types()"] - CallRel --> AddRel["_add_relationship() (dedup)"] - NewRel --> AddRel - InheritRel --> AddRel - JsdocRel --> AddRel diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers-4.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers-4.mmd deleted file mode 100644 index 26d20004..00000000 --- a/docs/diagrams/architecture/javascript_typescript_analyzers-4.mmd +++ /dev/null @@ -1,17 +0,0 @@ -flowchart TD - A["analyze()"] --> B["Pass 1: _extract_all_entities()
builds all_entities dict"] - B --> C["Pass 2: _filter_top_level_declarations()
_is_actually_top_level() check"] - C --> D["Node objects appended to self.nodes
and self.top_level_nodes"] - D --> E["_extract_constructor_dependencies()
(classes only)"] - B --> F["Pass 3: _extract_all_relationships()"] - F --> G["_traverse_for_relationships()
tracks current_top_level"] - G -->|"call_expression"| H["_extract_call_relationship()"] - G -->|"new_expression"| I["_extract_new_relationship()"] - G -->|"member_expression"| J["_extract_member_relationship()"] - G -->|"type_annotation / type_arguments"| K["_extract_type_relationship()"] - G -->|"extends_clause / implements_clause"| L["_extract_inheritance_relationship()"] - H --> M["_add_relationship()"] - I --> M - J --> M - K --> M - L --> M diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers-5.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers-5.mmd deleted file mode 100644 index eb338103..00000000 --- a/docs/diagrams/architecture/javascript_typescript_analyzers-5.mmd +++ /dev/null @@ -1,19 +0,0 @@ -sequenceDiagram - participant DP as "DependencyParser" - participant JS as "TreeSitterJSAnalyzer" - participant TS as "TreeSitterTSAnalyzer" - participant DGB as "DependencyGraphBuilder" - - DP->>JS: analyze_javascript_file_treesitter(path, content, repo_path) - JS->>JS: Parser.parse(content) - JS->>JS: _extract_functions() / _extract_call_relationships() - JS-->>DP: (nodes, call_relationships) - - DP->>TS: analyze_typescript_file_treesitter(path, content, repo_path) - TS->>TS: Parser.parse(content) - TS->>TS: _extract_all_entities() / _filter_top_level_declarations() / _extract_all_relationships() - TS-->>DP: (nodes, call_relationships) - - DP->>DGB: aggregated nodes + relationships - DGB->>DGB: resolve callee ids against repository Node index - DGB-->>DP: Repository dependency graph diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers.mmd deleted file mode 100644 index 47fbe960..00000000 --- a/docs/diagrams/architecture/javascript_typescript_analyzers.mmd +++ /dev/null @@ -1,10 +0,0 @@ -flowchart LR - DP["DependencyParser"] -->|"dispatches .js/.jsx/.mjs/.cjs files"| JSA["TreeSitterJSAnalyzer"] - DP -->|"dispatches .ts/.tsx files"| TSA["TreeSitterTSAnalyzer"] - JSA -->|"produces"| Nodes["Node objects"] - JSA -->|"produces"| Rels["CallRelationship objects"] - TSA -->|"produces"| Nodes - TSA -->|"produces"| Rels - Nodes --> DGB["DependencyGraphBuilder"] - Rels --> DGB - DGB --> Graph["Repository Dependency Graph"] diff --git a/docs/diagrams/architecture/job_models-2.mmd b/docs/diagrams/architecture/job_models-2.mmd deleted file mode 100644 index fc6bbd1a..00000000 --- a/docs/diagrams/architecture/job_models-2.mmd +++ /dev/null @@ -1,7 +0,0 @@ -stateDiagram-v2 - [*] --> PENDING : "DocumentationJob() created" - PENDING --> RUNNING : "start()" - RUNNING --> COMPLETED : "complete()" - RUNNING --> FAILED : "fail(error_message)" - COMPLETED --> [*] - FAILED --> [*] diff --git a/docs/diagrams/architecture/job_models-3.mmd b/docs/diagrams/architecture/job_models-3.mmd deleted file mode 100644 index 5cc40d39..00000000 --- a/docs/diagrams/architecture/job_models-3.mmd +++ /dev/null @@ -1,20 +0,0 @@ -sequenceDiagram - participant Caller as "CLI Component" - participant Job as "DocumentationJob" - participant Coerce as "Coercion Helpers" - - Caller->>Job: "DocumentationJob(...)" - Job-->>Caller: "job instance (status=PENDING)" - Caller->>Job: "job.start()" - Caller->>Job: "job.to_dict() / job.to_json()" - Job-->>Caller: "dict / JSON string" - Note over Caller: "Persisted to disk or transmitted" - - Caller->>Job: "DocumentationJob.from_dict(raw_data)" - Job->>Coerce: "_coerce_job_status(raw_data.status)" - Job->>Coerce: "_coerce_int(raw_data.module_count)" - Job->>Coerce: "_coerce_generation_options(raw_data.generation_options)" - Job->>Coerce: "_coerce_llm_config(raw_data.llm_config)" - Job->>Coerce: "_coerce_statistics(raw_data.statistics)" - Coerce-->>Job: "validated nested objects" - Job-->>Caller: "reconstructed DocumentationJob" diff --git a/docs/diagrams/architecture/job_models-4.mmd b/docs/diagrams/architecture/job_models-4.mmd deleted file mode 100644 index 0fc4c4ce..00000000 --- a/docs/diagrams/architecture/job_models-4.mmd +++ /dev/null @@ -1,15 +0,0 @@ -flowchart TD - JobModels["Job Models"] - CliCore["Cli Core (parent)"] - Generation["Generation"] - Configuration["Configuration"] - GitIntegration["Git Integration"] - HtmlGeneration["Html Generation"] - Utils["Utils"] - - CliCore --> JobModels - Generation -->|"creates & drives"| JobModels - Configuration -->|"populates LLMConfig"| JobModels - GitIntegration -->|"populates commit_hash / branch_name"| JobModels - HtmlGeneration -->|"reads files_generated"| JobModels - Utils -->|"reports on status/statistics"| JobModels diff --git a/docs/diagrams/architecture/job_models.mmd b/docs/diagrams/architecture/job_models.mmd deleted file mode 100644 index 80bc5699..00000000 --- a/docs/diagrams/architecture/job_models.mmd +++ /dev/null @@ -1,52 +0,0 @@ -classDiagram - class JobStatus { - <> - PENDING - RUNNING - COMPLETED - FAILED - } - class GenerationOptions { - +bool create_branch - +bool github_pages - +bool no_cache - +str custom_output - } - class LLMConfig { - +str main_model - +str cluster_model - +str base_url - } - class JobStatistics { - +int total_files_analyzed - +int leaf_nodes - +int max_depth - +int total_tokens_used - } - class DocumentationJob { - +str job_id - +str repository_path - +str repository_name - +str output_directory - +str commit_hash - +str branch_name - +str timestamp_start - +str timestamp_end - +JobStatus status - +str error_message - +List files_generated - +int module_count - +GenerationOptions generation_options - +LLMConfig llm_config - +JobStatistics statistics - +start() - +complete() - +fail(error_message) - +to_dict() dict - +to_json() str - +from_dict(data) DocumentationJob - } - DocumentationJob --> JobStatus : "status" - DocumentationJob --> GenerationOptions : "generation_options" - DocumentationJob --> LLMConfig : "llm_config (optional)" - DocumentationJob --> JobStatistics : "statistics" diff --git a/docs/diagrams/architecture/llm-services-2.mmd b/docs/diagrams/architecture/llm-services-2.mmd deleted file mode 100644 index 9606e040..00000000 --- a/docs/diagrams/architecture/llm-services-2.mmd +++ /dev/null @@ -1,20 +0,0 @@ -flowchart LR - subgraph LlmServices["Llm Services"] - Factory["Model factories
(main / fallback / cluster)"] - Counting["CountingFallbackModel"] - DirectCall["call_llm()"] - end - - subgraph DocGen["Documentation Generator"] - DG["DocumentationGenerator"] - Orchestrator["AgentOrchestrator"] - end - - ConfigCore["Config"] --> Factory - Factory --> Counting - Orchestrator -->|"create_fallback_models(config)"| Factory - Orchestrator -->|"used by pydantic-ai Agent"| Counting - DG -->|"call_llm(prompt, config)"| DirectCall - - DirectCall --> OpenAISDK["openai.OpenAI client"] - Counting --> PydanticAI["pydantic-ai FallbackModel / OpenAIProvider"] diff --git a/docs/diagrams/architecture/llm-services-3.mmd b/docs/diagrams/architecture/llm-services-3.mmd deleted file mode 100644 index 7d804584..00000000 --- a/docs/diagrams/architecture/llm-services-3.mmd +++ /dev/null @@ -1,22 +0,0 @@ -sequenceDiagram - participant Caller - participant Factory as "create_main_model() / create_fallback_model() / create_cluster_model()" - participant Config as "Config" - participant Provider as "OpenAIProvider" - participant Model as "OpenAIModel" - - Caller->>Factory: create_X_model(config) - Factory->>Config: getattr(config, "X_max_tokens", None) - Factory->>Config: getattr(config, "X_temperature", 0.0) - Factory->>Config: getattr(config, "X_temperature_supported", True) - Factory->>Config: getattr(config, "X_base_url", None) - alt base_url missing - Factory-->>Caller: raise ValueError("X_base_url is required...") - end - Factory->>Config: getattr(config, "X_api_key", None) - alt api_key missing - Factory-->>Caller: raise ValueError("X_api_key is required...") - end - Factory->>Provider: OpenAIProvider(base_url, api_key) - Factory->>Model: OpenAIModel(model_name, provider, settings) - Model-->>Caller: configured OpenAIModel diff --git a/docs/diagrams/architecture/llm-services-4.mmd b/docs/diagrams/architecture/llm-services-4.mmd deleted file mode 100644 index 2a21e453..00000000 --- a/docs/diagrams/architecture/llm-services-4.mmd +++ /dev/null @@ -1,16 +0,0 @@ -sequenceDiagram - participant Orchestrator as "AgentOrchestrator" - participant LlmServices as "llm_services" - participant MainModel as "OpenAIModel (main)" - participant FallbackModelObj as "OpenAIModel (fallback)" - participant CFM as "CountingFallbackModel" - - Orchestrator->>LlmServices: create_fallback_models(config) - LlmServices->>LlmServices: create_main_model(config) - LlmServices-->>MainModel: main - LlmServices->>LlmServices: create_fallback_model(config) - LlmServices-->>FallbackModelObj: fallback - LlmServices->>CFM: CountingFallbackModel(main, fallback) - CFM-->>Orchestrator: fallback_models - - Note over Orchestrator,CFM: Passed as the model backend
for pydantic-ai Agent diff --git a/docs/diagrams/architecture/llm-services-5.mmd b/docs/diagrams/architecture/llm-services-5.mmd deleted file mode 100644 index e86222d9..00000000 --- a/docs/diagrams/architecture/llm-services-5.mmd +++ /dev/null @@ -1,24 +0,0 @@ -sequenceDiagram - participant Caller - participant CallLLM as "call_llm()" - participant ClientFactory as "create_openai_client()" - participant Config as "Config" - participant OpenAISDK as "openai.OpenAI" - - Caller->>CallLLM: call_llm(prompt, config, model, temperature) - CallLLM->>CallLLM: resolve stage (main / cluster / fallback) - CallLLM->>ClientFactory: create_openai_client(config, model) - ClientFactory->>Config: resolve base_url / api_key / api_version for stage - alt base_url or api_key missing - ClientFactory-->>CallLLM: raise ValueError(...) - end - ClientFactory->>OpenAISDK: OpenAI(base_url, api_key, default_headers) - OpenAISDK-->>CallLLM: client - CallLLM->>CallLLM: get_model_max_token_field(stage) - CallLLM->>OpenAISDK: client.chat.completions.create(**kwargs) - alt success - OpenAISDK-->>CallLLM: response - CallLLM-->>Caller: response_content - else OpenAIError - CallLLM-->>Caller: raise RuntimeError(context-wrapped) - end diff --git a/docs/diagrams/architecture/llm-services.mmd b/docs/diagrams/architecture/llm-services.mmd deleted file mode 100644 index aa28dfd3..00000000 --- a/docs/diagrams/architecture/llm-services.mmd +++ /dev/null @@ -1,24 +0,0 @@ -flowchart TD - Config["Config"] --> CreateMain["create_main_model()"] - Config --> CreateFallback["create_fallback_model()"] - Config --> CreateCluster["create_cluster_model()"] - Config --> CreateClient["create_openai_client()"] - - CreateMain --> MainModel["OpenAIModel (main)"] - CreateFallback --> FallbackModel_["OpenAIModel (fallback)"] - - MainModel --> CreateChain["create_fallback_models()"] - FallbackModel_ --> CreateChain - - CreateChain --> CountingModel["CountingFallbackModel"] - CountingModel --> IncCounter["increment_request_counter()"] - IncCounter --> Counter[("_request_counter")] - - CountingModel --> Agent["pydantic-ai Agent"] - - CreateCluster --> ClusterModel["OpenAIModel (cluster)"] - ClusterModel --> ClusterAgent["Clustering LLM calls"] - - CreateClient --> OpenAIClient["openai.OpenAI client"] - OpenAIClient --> CallLLM["call_llm()"] - CallLLM --> Response["LLM response text"] diff --git a/docs/diagrams/architecture/logging-config-2.mmd b/docs/diagrams/architecture/logging-config-2.mmd deleted file mode 100644 index 9387e9d5..00000000 --- a/docs/diagrams/architecture/logging-config-2.mmd +++ /dev/null @@ -1,21 +0,0 @@ -sequenceDiagram - participant Caller as "Backend Component" - participant Setup as "setup_module_logging()" - participant Logger as "logging.Logger" - participant Handler as "StreamHandler(stdout)" - participant Fmt as "ColoredFormatter" - - Caller->>Setup: setup_module_logging("my_module", level) - Setup->>Logger: logging.getLogger("my_module") - Setup->>Handler: create StreamHandler(sys.stdout) - Setup->>Fmt: create ColoredFormatter() - Setup->>Handler: setFormatter(Fmt) - Setup->>Logger: handlers.clear() - Setup->>Logger: addHandler(Handler) - Setup->>Logger: propagate = False - Setup-->>Caller: return configured logger - Caller->>Logger: logger.info("message") - Logger->>Handler: emit(record) - Handler->>Fmt: format(record) - Fmt-->>Handler: colored log line - Handler-->>Caller: printed to stdout diff --git a/docs/diagrams/architecture/logging-config-3.mmd b/docs/diagrams/architecture/logging-config-3.mmd deleted file mode 100644 index f5e7ee0e..00000000 --- a/docs/diagrams/architecture/logging-config-3.mmd +++ /dev/null @@ -1,13 +0,0 @@ -classDiagram - class Formatter { - <> - +format(record) - +formatTime(record, datefmt) - +formatException(exc_info) - } - class ColoredFormatter { - +COLORS : dict - +COMPONENT_COLORS : dict - +format(record) str - } - Formatter <|-- ColoredFormatter diff --git a/docs/diagrams/architecture/logging-config-4.mmd b/docs/diagrams/architecture/logging-config-4.mmd deleted file mode 100644 index 9d71e38e..00000000 --- a/docs/diagrams/architecture/logging-config-4.mmd +++ /dev/null @@ -1,12 +0,0 @@ -graph TD - Backend["Backend Core"] --> LoggingConfig["Logging Config"] - Backend --> DepAnalyzer["Dependency Analyzer Core"] - Backend --> TreeSitter["Tree-Sitter Analyzers"] - Backend --> DepModels["Dependency Analyzer Models"] - Backend --> DocGen["Documentation Generator"] - Backend --> LLMServices["LLM Services"] - Backend --> AgentTools["Agent Tools Core"] - - DepAnalyzer -.->|"colored console output"| LoggingConfig - TreeSitter -.->|"colored console output"| LoggingConfig - DocGen -.->|"colored console output"| LoggingConfig diff --git a/docs/diagrams/architecture/logging-config.mmd b/docs/diagrams/architecture/logging-config.mmd deleted file mode 100644 index 1a14fc08..00000000 --- a/docs/diagrams/architecture/logging-config.mmd +++ /dev/null @@ -1,14 +0,0 @@ -flowchart TD - A["logging.LogRecord"] --> B["ColoredFormatter.format(record)"] - B --> C["Look up level color in COLORS"] - B --> D["Format timestamp HH:MM:SS"] - D --> E["Wrap timestamp in blue"] - C --> F["Wrap levelname in level color"] - C --> G["Wrap message in level color"] - E --> H["Concatenate: timestamp + level + message"] - F --> H - G --> H - H --> I{"record.exc_info present?"} - I -->|"Yes"| J["Append formatException() output"] - I -->|"No"| K["Return colored log line"] - J --> K diff --git a/docs/diagrams/architecture/php_analyzer-2.mmd b/docs/diagrams/architecture/php_analyzer-2.mmd deleted file mode 100644 index d4697603..00000000 --- a/docs/diagrams/architecture/php_analyzer-2.mmd +++ /dev/null @@ -1,41 +0,0 @@ -classDiagram - class NamespaceResolver { - +string current_namespace - +dict use_map - +register_namespace(ns) - +register_use(fqn, alias) - +resolve(name) string - } - class TreeSitterPHPAnalyzer { - +Path file_path - +string content - +string repo_path - +List nodes - +List call_relationships - +NamespaceResolver namespace_resolver - +_analyze() - +_extract_namespace_info(node, depth) - +_extract_nodes(node, lines, depth, parent_class) - +_extract_relationships(node, depth) - +_is_template_file() bool - +_get_component_id(name, parent_class) string - } - class Node { - +string id - +string name - +string component_type - +string file_path - +Set depends_on - +string source_code - +int start_line - +int end_line - } - class CallRelationship { - +string caller - +string callee - +int call_line - +bool is_resolved - } - TreeSitterPHPAnalyzer --> NamespaceResolver : "delegates name resolution" - TreeSitterPHPAnalyzer --> Node : "creates" - TreeSitterPHPAnalyzer --> CallRelationship : "creates" diff --git a/docs/diagrams/architecture/php_analyzer-3.mmd b/docs/diagrams/architecture/php_analyzer-3.mmd deleted file mode 100644 index 3ee94d5f..00000000 --- a/docs/diagrams/architecture/php_analyzer-3.mmd +++ /dev/null @@ -1,20 +0,0 @@ -sequenceDiagram - participant Caller as "Caller" - participant Analyzer as "TreeSitterPHPAnalyzer" - participant Parser as "tree_sitter Parser" - participant Resolver as "NamespaceResolver" - - Caller->>Analyzer: "__init__(file_path, content, repo_path)" - Analyzer->>Analyzer: "_is_template_file()" - alt is template file - Analyzer-->>Caller: "return (skip analysis)" - else not a template - Analyzer->>Parser: "parse(content)" - Parser-->>Analyzer: "AST root node" - Analyzer->>Analyzer: "_extract_namespace_info(root)" - Analyzer->>Resolver: "register_namespace() / register_use()" - Analyzer->>Analyzer: "_extract_nodes(root, lines)" - Analyzer->>Analyzer: "_extract_relationships(root)" - Analyzer->>Resolver: "resolve(type_name)" - Analyzer-->>Caller: "nodes, call_relationships populated" - end diff --git a/docs/diagrams/architecture/php_analyzer-4.mmd b/docs/diagrams/architecture/php_analyzer-4.mmd deleted file mode 100644 index e3ac6b68..00000000 --- a/docs/diagrams/architecture/php_analyzer-4.mmd +++ /dev/null @@ -1,7 +0,0 @@ -flowchart LR - File["PHP source file"] --> Analyzer["TreeSitterPHPAnalyzer"] - Analyzer -->|"nodes"| NodeList["List of Node objects
(classes, methods, functions, etc.)"] - Analyzer -->|"relationships"| RelList["List of CallRelationship objects
(extends, implements, new, static call)"] - NodeList --> Parser["DependencyParser
(Graph Construction)"] - RelList --> Parser - Parser --> Graph["Dependency Graph
(components + depends_on edges)"] diff --git a/docs/diagrams/architecture/php_analyzer.mmd b/docs/diagrams/architecture/php_analyzer.mmd deleted file mode 100644 index b02e83ed..00000000 --- a/docs/diagrams/architecture/php_analyzer.mmd +++ /dev/null @@ -1,9 +0,0 @@ -flowchart TD - Caller["Dependency Analysis Pipeline"] -->|"analyze_php_file(path, content, repo_path)"| Analyze["analyze_php_file()"] - Analyze --> Analyzer["TreeSitterPHPAnalyzer"] - Analyzer -->|"uses"| Resolver["NamespaceResolver"] - Analyzer -->|"parses with"| TreeSitter["tree_sitter_php / Parser"] - Analyzer -->|"produces"| Nodes["List of Node"] - Analyzer -->|"produces"| Rels["List of CallRelationship"] - Nodes --> Models["Dependency Analyzer Models"] - Rels --> Models diff --git a/docs/diagrams/architecture/python_analyzer-2.mmd b/docs/diagrams/architecture/python_analyzer-2.mmd deleted file mode 100644 index 364a74b2..00000000 --- a/docs/diagrams/architecture/python_analyzer-2.mmd +++ /dev/null @@ -1,4 +0,0 @@ -flowchart LR - FilePath["file_path"] -->|"os.path.relpath"| RelPath["relative_path"] - RelPath -->|"strip .py/.pyx, replace separators with dots"| ModulePath["module.path"] - ModulePath -->|"+ '::' + ComponentName"| ComponentID["component_id"] diff --git a/docs/diagrams/architecture/python_analyzer-3.mmd b/docs/diagrams/architecture/python_analyzer-3.mmd deleted file mode 100644 index b6853911..00000000 --- a/docs/diagrams/architecture/python_analyzer-3.mmd +++ /dev/null @@ -1,17 +0,0 @@ -sequenceDiagram - participant Visitor as PythonASTAnalyzer - participant AST as ast.NodeVisitor - participant Nodes as top_level_nodes - participant Rels as call_relationships - - AST->>Visitor: visit_ClassDef(node) - Visitor->>Visitor: extract base_classes - Visitor->>Nodes: register class Node - Visitor->>Rels: append inheritance CallRelationship (if resolved) - Visitor->>Visitor: set current_class_name - Visitor->>AST: generic_visit(node) - AST->>Visitor: visit_Call(node) [inside class/function body] - Visitor->>Visitor: _get_call_name(node.func) - Visitor->>Nodes: lookup call_name - Visitor->>Rels: append CallRelationship (resolved or unresolved) - Visitor->>Visitor: clear current_class_name diff --git a/docs/diagrams/architecture/python_analyzer-4.mmd b/docs/diagrams/architecture/python_analyzer-4.mmd deleted file mode 100644 index 8988784c..00000000 --- a/docs/diagrams/architecture/python_analyzer-4.mmd +++ /dev/null @@ -1,6 +0,0 @@ -flowchart TD - Start["analyze() called"] --> Parse["ast.parse(content)"] - Parse -->|"success"| Visit["self.visit(tree)"] - Parse -->|"SyntaxError"| WarnLog["log warning, skip file"] - Visit -->|"success"| Done["nodes + call_relationships populated"] - Visit -->|"unexpected Exception"| ErrLog["log error with traceback"] diff --git a/docs/diagrams/architecture/python_analyzer.mmd b/docs/diagrams/architecture/python_analyzer.mmd deleted file mode 100644 index c5db2c86..00000000 --- a/docs/diagrams/architecture/python_analyzer.mmd +++ /dev/null @@ -1,11 +0,0 @@ -flowchart TD - RepoAnalyzer["RepoAnalyzer"] -->|"reads .py file"| PyFile["Python Source File"] - RepoAnalyzer -->|"instantiate"| PythonASTAnalyzer["PythonASTAnalyzer"] - PyFile -->|"content"| PythonASTAnalyzer - PythonASTAnalyzer -->|"ast.parse()"| AST["Python AST"] - AST -->|"NodeVisitor traversal"| PythonASTAnalyzer - PythonASTAnalyzer -->|"produces"| Nodes["List of Node"] - PythonASTAnalyzer -->|"produces"| Relationships["List of CallRelationship"] - Nodes --> DependencyGraphBuilder["DependencyGraphBuilder"] - Relationships --> DependencyGraphBuilder - DependencyGraphBuilder --> Graph["Dependency Graph"] diff --git a/docs/diagrams/architecture/sample_fixtures-2.mmd b/docs/diagrams/architecture/sample_fixtures-2.mmd deleted file mode 100644 index 116beb42..00000000 --- a/docs/diagrams/architecture/sample_fixtures-2.mmd +++ /dev/null @@ -1,28 +0,0 @@ -classDiagram - class APIController { - +service: MainService - +request_count: int - +__init__() - +handle_request(endpoint, data) dict - } - class MainService { - +name: str - +active: bool - +__init__(name) - +start(config) bool - +process_request(data) dict - +stop() - } - class PluginInterface { - +initialize() bool - +execute(context) dict - } - class DataPlugin { - +name: str - +initialized: bool - +__init__(name) - +initialize() bool - +execute(context) dict - } - APIController --> MainService : "uses" - DataPlugin --|> PluginInterface : "implements" diff --git a/docs/diagrams/architecture/sample_fixtures-3.mmd b/docs/diagrams/architecture/sample_fixtures-3.mmd deleted file mode 100644 index d2ebf7b6..00000000 --- a/docs/diagrams/architecture/sample_fixtures-3.mmd +++ /dev/null @@ -1,11 +0,0 @@ -sequenceDiagram - participant Client - participant Controller as APIController - participant Service as MainService - - Client->>Controller: handle_request("/process", data) - Controller->>Controller: request_count += 1 - Controller->>Service: process_request(data) - Service->>Service: check active flag - Service-->>Controller: {"status": "success", "result": processed} - Controller-->>Client: response dict diff --git a/docs/diagrams/architecture/sample_fixtures-4.mmd b/docs/diagrams/architecture/sample_fixtures-4.mmd deleted file mode 100644 index 968fb7b7..00000000 --- a/docs/diagrams/architecture/sample_fixtures-4.mmd +++ /dev/null @@ -1,5 +0,0 @@ -stateDiagram-v2 - [*] --> Uninitialized: "DataPlugin(name)" - Uninitialized --> Initialized: "initialize()" - Initialized --> Initialized: "execute(context)" - Uninitialized --> Error: "execute() before initialize()" diff --git a/docs/diagrams/architecture/sample_fixtures-5.mmd b/docs/diagrams/architecture/sample_fixtures-5.mmd deleted file mode 100644 index d1775cc3..00000000 --- a/docs/diagrams/architecture/sample_fixtures-5.mmd +++ /dev/null @@ -1,4 +0,0 @@ -graph LR - Controller["main.controller.APIController"] -->|"imports"| Service["main.service.MainService"] - Service -->|"imports"| Deps["deps.helper (external)"] - Plugin["external.plugin.DataPlugin"] -->|"implements"| Iface["external.plugin.PluginInterface"] diff --git a/docs/diagrams/architecture/sample_fixtures.mmd b/docs/diagrams/architecture/sample_fixtures.mmd deleted file mode 100644 index 4a9397b3..00000000 --- a/docs/diagrams/architecture/sample_fixtures.mmd +++ /dev/null @@ -1,4 +0,0 @@ -graph TD - Parent["Test Multi Path"] --> SampleFixtures["Sample Fixtures"] - Parent --> TestSuites["Test Suites"] - SampleFixtures -.->|"exercised by"| TestSuites diff --git a/docs/diagrams/architecture/test-clustering-2.mmd b/docs/diagrams/architecture/test-clustering-2.mmd deleted file mode 100644 index 667f93aa..00000000 --- a/docs/diagrams/architecture/test-clustering-2.mmd +++ /dev/null @@ -1,23 +0,0 @@ -classDiagram - class TestResultsBase { - +add_test(name, passed, details) - +print_summary() bool - } - class DebugResults - class ForcedResults - class LocalResults - class IntegrationResults { - +passed int - +failed int - } - class ValidationResults { - +success bool - } - class IdBasedResults - - TestResultsBase <|-- DebugResults - TestResultsBase <|-- ForcedResults - TestResultsBase <|-- LocalResults - TestResultsBase <|-- IntegrationResults - TestResultsBase <|-- ValidationResults - TestResultsBase <|-- IdBasedResults diff --git a/docs/diagrams/architecture/test-clustering-3.mmd b/docs/diagrams/architecture/test-clustering-3.mmd deleted file mode 100644 index bc44214d..00000000 --- a/docs/diagrams/architecture/test-clustering-3.mmd +++ /dev/null @@ -1,15 +0,0 @@ -sequenceDiagram - participant Dev as Developer - participant Debug as "test_clustering_debug.py" - participant Patch as "Monkey-patched LLM Client" - participant Cluster as "cluster_modules()" - - Dev->>Debug: python3 test_clustering_debug.py - Debug->>Patch: patch create_llm_client() - Debug->>Cluster: cluster_modules(leaf_nodes, components, config) - Cluster->>Patch: client.call(prompt) - Patch-->>Cluster: raw LLM response text - Patch-->>Debug: capture response (global variable) - Cluster-->>Debug: module_tree dict - Debug->>Dev: print captured response + module_tree summary - Debug->>Dev: TestResults.print_summary() diff --git a/docs/diagrams/architecture/test-clustering.mmd b/docs/diagrams/architecture/test-clustering.mmd deleted file mode 100644 index baeef457..00000000 --- a/docs/diagrams/architecture/test-clustering.mmd +++ /dev/null @@ -1,37 +0,0 @@ -flowchart TD - subgraph scripts["Test Clustering Scripts"] - Debug["test_clustering_debug.py"] - Forced["test_clustering_forced.py"] - Local["test_clustering_local.py"] - Integration["test_clustering_integration.py"] - Validation["test_clustering_validation.py"] - IdBased["test_id_based_clustering.py"] - end - - subgraph backend["Backend Clustering Pipeline"] - ClusterFn["cluster_modules()"] - IdMap["create_component_id_map()"] - Normalize["normalize_component_ids_by_lookup()"] - LLMClient["LLM Client"] - end - - ConfigMod["Config"] - NodeMod["Node"] - - Debug -->|"invokes"| ClusterFn - Forced -->|"invokes"| ClusterFn - Local -->|"invokes"| ClusterFn - ClusterFn -->|"calls"| LLMClient - - Integration -->|"invokes directly"| IdMap - Integration -->|"invokes directly"| Normalize - - Validation -->|"simulates validation logic of"| ClusterFn - IdBased -->|"simulates parsing/validation logic of"| ClusterFn - - Debug -->|"constructs"| NodeMod - Forced -->|"constructs"| NodeMod - Local -->|"constructs"| NodeMod - Debug -->|"constructs"| ConfigMod - Forced -->|"constructs"| ConfigMod - Local -->|"constructs"| ConfigMod diff --git a/docs/diagrams/architecture/test-multi-path-2.mmd b/docs/diagrams/architecture/test-multi-path-2.mmd deleted file mode 100644 index 5b0d9b95..00000000 --- a/docs/diagrams/architecture/test-multi-path-2.mmd +++ /dev/null @@ -1,12 +0,0 @@ -sequenceDiagram - participant Suite as "Test Script" - participant Cfg as "Config" - participant Builder as "DependencyGraphBuilder" - - Suite->>Cfg: Create Config with repo_path and additional_source_paths - Suite->>Cfg: validate_source_paths() (optional, for invalid-path test) - Suite->>Builder: DependencyGraphBuilder(config) - Suite->>Builder: build_dependency_graph() - Builder-->>Suite: Returns (components, leaf_nodes) - Suite->>Suite: Assert namespaces, counts, and cross-path edges - Suite->>Suite: Print summary and exit code diff --git a/docs/diagrams/architecture/test-multi-path.mmd b/docs/diagrams/architecture/test-multi-path.mmd deleted file mode 100644 index 993d8fc5..00000000 --- a/docs/diagrams/architecture/test-multi-path.mmd +++ /dev/null @@ -1,25 +0,0 @@ -flowchart TD - subgraph fixtures["Sample Application Fixtures"] - Service["MainService"] - Controller["APIController"] - Plugin["DataPlugin / PluginInterface"] - end - - subgraph suites["Test Suites and Runners"] - MultiPathSuite["test_multi_path.py suite"] - IntegrationRunner["IntegrationTestRunner"] - end - - subgraph external_deps["External Dependencies"] - ConfigCls["Config"] - Builder["DependencyGraphBuilder"] - end - - MultiPathSuite -->|"constructs"| ConfigCls - IntegrationRunner -->|"constructs"| ConfigCls - ConfigCls -->|"declares additional_source_paths"| Builder - Builder -->|"parses"| Service - Builder -->|"parses"| Controller - Builder -->|"parses"| Plugin - MultiPathSuite -->|"asserts on"| Builder - IntegrationRunner -->|"asserts on"| Builder diff --git a/docs/diagrams/architecture/test_suites-2.mmd b/docs/diagrams/architecture/test_suites-2.mmd deleted file mode 100644 index da3a803e..00000000 --- a/docs/diagrams/architecture/test_suites-2.mmd +++ /dev/null @@ -1,19 +0,0 @@ -flowchart TD - Start["run_all_tests()"] --> T1["test_single_path()"] - Start --> T2["test_multiple_paths()"] - Start --> T3["test_component_namespacing()"] - Start --> T4["test_cross_path_dependencies()"] - Start --> T5["test_invalid_path_handling()"] - Start --> T6["test_empty_additional_paths()"] - Start --> T7["test_relative_vs_absolute_paths()"] - T1 --> Builder["DependencyGraphBuilder.build_dependency_graph()"] - T2 --> Builder - T3 --> Builder - T4 --> Builder - T6 --> Builder - T7 --> Builder - T5 --> Validate["Config.validate_source_paths()"] - Builder --> Results["TestResults.add_test()"] - Validate --> Results - Results --> Summary["TestResults.print_summary()"] - Summary --> ExitCode["Process exit code (0 or 1)"] diff --git a/docs/diagrams/architecture/test_suites-3.mmd b/docs/diagrams/architecture/test_suites-3.mmd deleted file mode 100644 index f559e4ad..00000000 --- a/docs/diagrams/architecture/test_suites-3.mmd +++ /dev/null @@ -1,25 +0,0 @@ -sequenceDiagram - participant Main as "main()" - participant Runner as "IntegrationTestRunner" - participant Cfg as "Config" - participant Builder as "DependencyGraphBuilder" - participant Results as "IntegrationTestResults" - - Main->>Runner: run() - Runner->>Runner: setup_test_environment() - Note over Runner: Creates main/, deps/, vendor/ with cross-importing fixture files - Runner->>Cfg: create_config() - Cfg-->>Runner: Config(repo_path, additional_source_paths) - Runner->>Runner: validate_paths() - Runner->>Builder: execute_dependency_parser() - Builder->>Builder: build_dependency_graph() - Builder-->>Runner: components, leaf_nodes - Runner->>Results: verify_namespaces() - Runner->>Results: verify_cross_path_dependencies() - Runner->>Results: verify_no_warnings() - Runner->>Results: verify_file_counts() - Runner->>Runner: print_detailed_output() - Runner->>Results: print_summary() - Results-->>Runner: all_passed() - Runner->>Runner: cleanup() - Runner-->>Main: exit code 0, 1, or 2 diff --git a/docs/diagrams/architecture/test_suites-4.mmd b/docs/diagrams/architecture/test_suites-4.mmd deleted file mode 100644 index df7aea08..00000000 --- a/docs/diagrams/architecture/test_suites-4.mmd +++ /dev/null @@ -1,15 +0,0 @@ -flowchart LR - subgraph TS["Test Suites"] - SM["Scenario Test Suite (test_multi_path.py)"] - IT["Integration Test Runner (integration_test.py)"] - end - subgraph CC["Config Core"] - Config["Config"] - end - subgraph DA["Dependency Analyzer Core"] - DGB["DependencyGraphBuilder"] - end - SM --> Config - IT --> Config - SM --> DGB - IT --> DGB diff --git a/docs/diagrams/architecture/test_suites.mmd b/docs/diagrams/architecture/test_suites.mmd deleted file mode 100644 index 3b7dc665..00000000 --- a/docs/diagrams/architecture/test_suites.mmd +++ /dev/null @@ -1,46 +0,0 @@ -classDiagram - class Colors { - +GREEN - +RED - +YELLOW - +BLUE - +BOLD - +END - } - class ScenarioTestResults { - -results dict - +add_test(name, passed) - +print_summary() int - } - class IntegrationTestRunner { - -test_dir - -main_path - -deps_path - -vendor_path - -config - -builder - -components - -leaf_nodes - -results - +setup_test_environment() - +create_config() - +validate_paths() - +execute_dependency_parser() - +verify_namespaces() - +verify_cross_path_dependencies() - +verify_no_warnings() - +verify_file_counts() - +print_detailed_output() - +cleanup() - +run() int - } - class IntegrationTestResults { - -tests list - +add_test(name, passed, details) - +print_summary() - +all_passed() bool - } - IntegrationTestRunner --> IntegrationTestResults : records into - IntegrationTestRunner --> Config : creates - IntegrationTestRunner --> DependencyGraphBuilder : invokes - ScenarioTestResults ..> Colors : uses for output diff --git a/docs/diagrams/architecture/tree-sitter-analyzers.mmd b/docs/diagrams/architecture/tree-sitter-analyzers.mmd deleted file mode 100644 index 4fd82f40..00000000 --- a/docs/diagrams/architecture/tree-sitter-analyzers.mmd +++ /dev/null @@ -1,22 +0,0 @@ -flowchart TD - Caller["Dependency Parser"] -->|"dispatches by file extension"| Router["Language Router"] - Router -->|".c / .h"| CAnalyzer["TreeSitterCAnalyzer"] - Router -->|".cpp / .hpp / .cc"| CppAnalyzer["TreeSitterCppAnalyzer"] - Router -->|".cs"| CSharpAnalyzer["TreeSitterCSharpAnalyzer"] - Router -->|".java"| JavaAnalyzer["TreeSitterJavaAnalyzer"] - Router -->|".js / .jsx / .mjs"| JSAnalyzer["TreeSitterJSAnalyzer"] - Router -->|".ts / .tsx"| TSAnalyzer["TreeSitterTSAnalyzer"] - Router -->|".php"| PHPAnalyzer["TreeSitterPHPAnalyzer"] - Router -->|".py"| PyAnalyzer["PythonASTAnalyzer"] - - CAnalyzer --> Output["Node and CallRelationship lists"] - CppAnalyzer --> Output - CSharpAnalyzer --> Output - JavaAnalyzer --> Output - JSAnalyzer --> Output - TSAnalyzer --> Output - PHPAnalyzer --> Output - PyAnalyzer --> Output - - Output --> Builder["Dependency Graph Builder"] - Output --> CallGraph["Call Graph Analyzer"] diff --git a/docs/diagrams/architecture/utils-2.mmd b/docs/diagrams/architecture/utils-2.mmd deleted file mode 100644 index 064dda9f..00000000 --- a/docs/diagrams/architecture/utils-2.mmd +++ /dev/null @@ -1,22 +0,0 @@ -sequenceDiagram - participant Pipeline as "CLI Pipeline" - participant Logger as "CLILogger" - participant Tracker as "ProgressTracker" - participant ModBar as "ModuleProgressBar" - - Pipeline->>Logger: create_logger(verbose) - Pipeline->>Tracker: start_stage(1, "Dependency Analysis") - Pipeline->>Logger: info("Analyzing repository...") - Tracker-->>Pipeline: update_stage(progress) - Pipeline->>Tracker: complete_stage() - - Pipeline->>Tracker: start_stage(3, "Documentation Generation") - Pipeline->>ModBar: new ModuleProgressBar(total_modules) - loop "for each module" - Pipeline->>ModBar: update(module_name, cached) - Pipeline->>Logger: debug("module details") - end - Pipeline->>ModBar: finish() - Pipeline->>Tracker: complete_stage("Generation finished") - - Pipeline->>Logger: success("Documentation generated") diff --git a/docs/diagrams/architecture/utils-3.mmd b/docs/diagrams/architecture/utils-3.mmd deleted file mode 100644 index 9ef7a30b..00000000 --- a/docs/diagrams/architecture/utils-3.mmd +++ /dev/null @@ -1,11 +0,0 @@ -flowchart TD - A["create_logger(verbose)"] --> B["ProgressTracker(total_stages=5)"] - B --> C["tracker.start_stage(1, 'Dependency Analysis')"] - C --> D["perform analysis; call tracker.update_stage()"] - D --> E["tracker.complete_stage()"] - E --> F["tracker.start_stage(3, 'Documentation Generation')"] - F --> G["ModuleProgressBar(total_modules)"] - G --> H["for each module: bar.update(name, cached)"] - H --> I["bar.finish()"] - I --> J["tracker.complete_stage()"] - J --> K["logger.success('Done')"] diff --git a/docs/diagrams/architecture/utils.mmd b/docs/diagrams/architecture/utils.mmd deleted file mode 100644 index 69a2014a..00000000 --- a/docs/diagrams/architecture/utils.mmd +++ /dev/null @@ -1,36 +0,0 @@ -classDiagram - class CLILogger { - +bool verbose - +datetime start_time - +debug(message) void - +info(message) void - +success(message) void - +warning(message) void - +error(message) void - +step(message, step, total) void - +elapsed_time() str - } - - class ProgressTracker { - +int total_stages - +int current_stage - +float stage_progress - +float start_time - +bool verbose - +STAGE_WEIGHTS dict - +STAGE_NAMES dict - +start_stage(stage, description) void - +update_stage(progress, message) void - +complete_stage(message) void - +get_overall_progress() float - +get_eta() str - } - - class ModuleProgressBar { - +int total_modules - +int current_module - +bool verbose - +bar - +update(module_name, cached) void - +finish() void - } diff --git a/docs/getting-started/.gitignore b/docs/getting-started/.gitignore deleted file mode 100644 index a5d01be5..00000000 --- a/docs/getting-started/.gitignore +++ /dev/null @@ -1,7 +0,0 @@ -# VoltAgent temp files -temp/ - -# JSON intermediate files (except schema/config) -*.json -!*-schema.json -!*-config.json diff --git a/docs/getting-started/first-steps.md b/docs/getting-started/first-steps.md deleted file mode 100644 index 192c1604..00000000 --- a/docs/getting-started/first-steps.md +++ /dev/null @@ -1,95 +0,0 @@ -# First Steps - -You've installed CodeWiki, configured your LLM credentials, and generated your first documentation set. Here's what to explore next. - -## 1. Inspect and Confirm Your Configuration - -Use `codewiki config show` to review everything CodeWiki currently knows about your setup — models, base URLs, token limits, temperature settings, and agent instructions. API keys are always masked (only the first/last few characters shown). - -```bash -codewiki config show -codewiki config show --json -``` - -Use `codewiki config validate` any time you change providers or suspect something is misconfigured. It checks the config file, verifies all three API keys are present, validates base URL formats, confirms models are set, and (unless `--quick` is passed) performs a live connectivity test against each configured provider. - -```bash -codewiki config validate -codewiki config validate --quick # Skip live API connectivity test -codewiki config validate --verbose # Step-by-step diagnostic output -``` - -## 2. Customize What Gets Documented - -The `generate` command accepts several options that narrow or reshape the analysis without touching your saved configuration: - -```bash -# Only analyze C# files, skip test projects -codewiki generate --include "*.cs" --exclude "*Tests*,*Specs*,test_*" - -# Focus documentation on specific modules/paths -codewiki generate --focus "src/core,src/api" --doc-type architecture - -# Add free-form custom instructions for the documentation agent -codewiki generate --instructions "Focus on public APIs and include usage examples" - -# Include additional source directories (e.g., vendored dependencies) -codewiki generate --additional-paths "vendor/packages,external/deps" -``` - -If you want these choices to become your **default** behavior for every future run (rather than one-off flags), persist them with: - -```bash -codewiki config agent --include "*.cs" --exclude "*Tests*,*Specs*" -codewiki config agent --doc-type architecture -codewiki config agent --instructions "Focus on public APIs and include usage examples" - -# Clear all saved agent instructions -codewiki config agent --clear -``` - -`--doc-type` accepts one of: `api`, `architecture`, `user-guide`, or `developer`. - -## 3. Tune Token Budgets and Depth for Large Repositories - -If your repository is very large or the LLM response is being truncated, adjust token and depth limits either per-run or persistently: - -```bash -# Per-run override -codewiki generate --max-tokens 32768 --max-token-per-module 40000 --max-token-per-leaf-module 20000 --max-depth 3 - -# Persist as defaults -codewiki config set --cluster-max-tokens 128000 --main-max-tokens 128000 \ - --max-token-per-module 40000 --max-token-per-leaf-module 20000 --max-depth 3 -``` - -`--max-depth` controls how many levels of hierarchical module decomposition are produced (default: 2). - -## 4. Explore the Git and GitHub Pages Workflow - -If you're documenting a Git-tracked project, CodeWiki can create a dedicated branch for the generated docs and prepare a GitHub Pages-ready static site: - -```bash -codewiki generate --create-branch --github-pages -``` - -`--create-branch` requires a clean working tree and creates a timestamped branch. `--github-pages` renders a self-contained `index.html` from the generated `module_tree.json` and `metadata.json`, suitable for publishing directly. - -For CI/CD pipelines where you don't want interactive prompts, add `--force` to overwrite existing documentation without prompting, and `--no-cache` to force a full regeneration. - -## 5. Explore the Generated Output Structure - -After a run, look inside your output directory (default `./docs`) for: - -- Individual Markdown files per analyzed module (leaves generated first, then parent overview pages) -- `module_tree.json` — the hierarchical module structure used for navigation -- `metadata.json` — job status and generation statistics -- `index.html` (only if `--github-pages` was used) - -If you used `--diagrams-output`, Mermaid diagrams extracted from the generated Markdown are also saved separately as `.mmd` files. - -## Where to Get Help - -- Run `codewiki --help`, `codewiki generate --help`, or `codewiki config --help` / `codewiki config set --help` for full flag references directly in your terminal — these are the most up-to-date source of truth for available options. -- For questions, feedback, or community discussion, join the OpenMSP Slack community: [https://www.openmsp.ai/](https://www.openmsp.ai/) ([join link](https://join.slack.com/t/openmsp/shared_invite/zt-36bl7mx0h-3~U2nFH6nqHqoTPXMaHEHA)). -- If you plan to contribute code or documentation improvements back to CodeWiki itself, continue on to the development section of this documentation. diff --git a/docs/getting-started/introduction.md b/docs/getting-started/introduction.md deleted file mode 100644 index 81f5512a..00000000 --- a/docs/getting-started/introduction.md +++ /dev/null @@ -1,85 +0,0 @@ -# Introduction to CodeWiki - -## What is CodeWiki? - -**CodeWiki** is an AI-powered documentation generator for source code repositories. Point it at a codebase — locally via a command-line interface, or remotely via a web application — and it will: - -1. Analyze the repository's file structure and cross-file call relationships -2. Build a dependency graph of functions, classes, and modules -3. Cluster related components into meaningful, hierarchical modules -4. Use LLM-backed agents to write structured Markdown documentation for each module -5. Produce navigable output, including an optional static HTML viewer suitable for GitHub Pages - -CodeWiki is built as a Python package (`codewiki`, requires Python `>=3.12`) and supports analysis of **Python, Java, JavaScript, TypeScript, C, C++, C#, and PHP** source files through dedicated tree-sitter based language analyzers. - -## Key Features - -- **Two entry points, one engine** — a `codewiki` CLI for local repositories and a FastAPI web application for submitting GitHub repository URLs. Both drive the same backend documentation pipeline. -- **Multi-language dependency analysis** — tree-sitter powered analyzers extract call graphs and structural relationships across eight languages. -- **LLM-driven module clustering** — instead of a flat file listing, components are grouped into meaningful modules (e.g., "Auth Module", "API Module") by an LLM, then documented leaf-first before parent overviews are generated. -- **Per-provider LLM configuration** — separate model, API key, base URL, token limit, and temperature settings for the *cluster*, *main*, and *fallback* LLM roles, so you can mix providers (e.g., a cheaper model for clustering, a stronger model for generation). -- **Secure credential storage** — API keys are stored in the OS keyring (macOS Keychain, Windows Credential Manager, Linux Secret Service), never in plaintext configuration files. -- **Git-aware workflow** — the CLI can validate a clean working tree, create a timestamped documentation branch, and prepare it for a pull request. -- **Static HTML output** — an optional, self-contained `index.html` viewer can be generated for GitHub Pages deployment. -- **Caching for the web app** — the FastAPI frontend caches generated documentation by repository URL (with configurable expiry) to avoid redundant regeneration. -- **Multi-path analysis** — repositories whose source is split across multiple root directories (e.g., a monorepo with `main/`, `deps/`, `vendor/`) can be analyzed as a single unified documentation set via `additional_source_paths`. - -## Target Audience - -CodeWiki is for: - -- **Individual developers** who want up-to-date architecture documentation for a repository without manually writing it. -- **Teams and maintainers** who want a repeatable, CI-friendly way to regenerate documentation as code evolves (`--force` flag for non-interactive use). -- **Open-source project maintainers** who want a GitHub Pages-ready documentation site generated directly from their codebase. -- **Platform operators** who want to offer documentation-as-a-service through a hosted web application backed by GitHub repository submissions. - -## High-Level Architecture - -```mermaid -flowchart TD - User["Developer or Web User"] --> Entry{{"Choose entry point?"}} - Entry -->|CLI| CLI["CLI Core"] - Entry -->|Web| Frontend["Frontend Core"] - - CLI --> RuntimeConfig["Config Core"] - Frontend --> RuntimeConfig - - CLI --> GitOps["Git Integration"] - Frontend --> RepoProcessor["GitHub Repository Processor"] - - GitOps --> Source["Source Repository"] - RepoProcessor --> Source - - RuntimeConfig --> Generator["DocumentationGenerator"] - Source --> Generator - - Generator --> Analysis["Dependency Analysis"] - Analysis --> Parsers["Language Parsers"] - Parsers --> Graph["Dependency Graph"] - - Graph --> Clustering["Module Clustering"] - Clustering --> Agents["LLM Agent Orchestration"] - Agents --> Docs["Markdown Documentation"] - - Docs --> Metadata["Module Tree and Metadata"] - Metadata --> HTML["Optional HTML Viewer"] - Docs --> Output["Generated Documentation Output"] - HTML --> Output -``` - -## Core Modules at a Glance - -| Module | Purpose | -|---|---| -| **CLI Core** | Command-line orchestration: local config, Git integration, terminal progress, static HTML generation. | -| **Backend Core** | Repository analysis, dependency-graph construction, module clustering, LLM agent orchestration, documentation generation. | -| **Frontend Core** | FastAPI web application: repository submission, background job processing, caching, doc serving. | -| **Config Core** | Shared runtime `Config` model for source paths, output locations, provider settings, token limits, and agent instructions. | - -> **Note:** CodeWiki itself is a documentation-generation *tool* — the "modules" above describe CodeWiki's own internal architecture, not application features you configure for an end-user product. When you run CodeWiki against your own repository, it analyzes *your* code and produces documentation about *your* project. - -## Where to Go Next - -- [Prerequisites](prerequisites.md) — required software, versions, and environment setup before installing CodeWiki. -- [Quick Start](quick-start.md) — a five-minute walkthrough of installing CodeWiki and generating your first documentation set. -- [First Steps](first-steps.md) — what to explore after your first successful run. diff --git a/docs/getting-started/prerequisites.md b/docs/getting-started/prerequisites.md deleted file mode 100644 index 4869d69e..00000000 --- a/docs/getting-started/prerequisites.md +++ /dev/null @@ -1,71 +0,0 @@ -# Prerequisites - -Before installing CodeWiki, make sure your environment meets the following requirements. - -## Required Software - -| Software | Minimum Version | Why It's Needed | -|---|---|---| -| Python | `>=3.12` | CodeWiki is a Python package (`pyproject.toml` requires `requires-python = ">=3.12"`). | -| pip | Bundled with Python 3.12 | Used to install the `codewiki` package and its dependencies from `requirements.txt` / `pyproject.toml`. | -| Git | Any recent version | Required for repository cloning (web app), git-branch workflows (`--create-branch`), and commit-hash/branch detection. | -| Node.js | `>=14.0.0` | Declared as a build requirement in `pyproject.toml` (`[external] build-requires`) — used by `mermaid-py` to validate Mermaid diagrams embedded in generated documentation. | -| Docker & Docker Compose | Recent version | Optional — only needed if you want to run the web application via `docker/docker-compose.yml` instead of running it directly with Python. | - -## System Requirements - -- **OS**: Linux, macOS, or Windows. Keyring-backed credential storage uses the OS-native secret store: macOS Keychain, Windows Credential Manager, or a Linux Secret Service implementation (e.g., GNOME Keyring). If no keyring backend is available, CodeWiki degrades gracefully (API keys must then be supplied another way each run). -- **Disk space**: Sufficient space to clone the target repository plus generated output (`output/cache`, `output/temp`, `output/docs`, `output/dependency_graphs`). -- **Network access**: Outbound HTTPS access to your configured LLM provider endpoint(s) (e.g., OpenAI, Anthropic, or an OpenAI-compatible base URL) is required at generation time. - -## Account / Access Requirements - -CodeWiki does not include a bundled LLM — you must bring your own provider credentials for **each** of the three configurable model roles: - -| Role | Purpose | -|---|---| -| **Cluster model** | Groups discovered code components into a hierarchical module tree. Recommended: a strong/top-tier model, since clustering quality drives overall documentation structure. | -| **Main model** | Writes the actual Markdown documentation for each module. | -| **Fallback model** | Used automatically if the main model call fails or is rate-limited. | - -Each role has its own API key, base URL, API version (for Anthropic-style APIs), max-token setting, and temperature setting — they do not need to be the same provider. - -## Environment Variables - -CodeWiki's CLI does **not** rely on ad-hoc environment variables for normal operation — persistent configuration is stored in `~/.codewiki/config.json` (non-secret settings) and the OS keyring (API keys), managed via `codewiki config set`. - -For the **test/diagnostic scripts** included in the repository (e.g., `test_clustering_*.py`), the following environment variables (or a local `.env.local` file) may be read directly: - -| Variable | Purpose | -|---|---| -| `OPENAI_API_KEY` | API key for OpenAI-compatible providers used by ad-hoc clustering test scripts. | -| `ANTHROPIC_API_KEY` | API key for Anthropic providers used by ad-hoc clustering test scripts. | -| `MAIN_API_KEY` / `CLUSTER_API_KEY` / `FALLBACK_API_KEY` | Per-role overrides used by the same test scripts. | - -For the **Docker Compose** deployment of the web application, environment values are loaded from an `.env` file referenced in `docker/docker-compose.yml` (`env_file: ../.env`), and `APP_PORT` controls the host port mapping (defaults to `8000`). - -## Verification Commands - -Run these commands to confirm your environment is ready before installing CodeWiki: - -```bash -# Check Python version (must be 3.12 or higher) -python3 --version - -# Check pip is available -pip3 --version - -# Check Git is installed -git --version - -# Check Node.js is installed (required by mermaid-py for diagram validation) -node --version - -# Optional: check Docker and Docker Compose (only needed for the web app container) -docker --version -docker compose version -``` - -> **Note:** If `python3 --version` reports an older version than 3.12, install a compatible Python before proceeding — CodeWiki's `pyproject.toml` will refuse to install otherwise. - -Once these checks pass, continue to the [Quick Start](quick-start.md) guide to install and run CodeWiki for the first time. diff --git a/docs/getting-started/quick-start.md b/docs/getting-started/quick-start.md deleted file mode 100644 index 330b15fc..00000000 --- a/docs/getting-started/quick-start.md +++ /dev/null @@ -1,114 +0,0 @@ -# Quick Start - -This guide gets you from zero to a generated documentation set in about five minutes, using the `codewiki` CLI against a local repository. - -> If you'd rather run CodeWiki as a hosted web service (submit a GitHub URL, poll for job status, view cached results), see the Docker-based setup mentioned at the end of this guide instead. - -## Step 1: Install CodeWiki - -Clone the repository and install the package (editable install is convenient for exploring the source): - -```bash -git clone https://github.com/flamingo-stack/CodeWiki.git -cd CodeWiki -pip install -e . -``` - -This registers the `codewiki` console command, defined in `pyproject.toml` as: - -```text -[project.scripts] -codewiki = "codewiki.cli.main:cli" -``` - -Verify the install: - -```bash -codewiki --version -codewiki version -``` - -## Step 2: Configure Your LLM Credentials - -CodeWiki needs API credentials for at least a **main model** and a **cluster model** (a fallback model is optional but recommended). Credentials are stored securely in your OS keyring; non-secret settings go to `~/.codewiki/config.json`. - -```bash -codewiki config set \ - --cluster-api-key "sk-your-cluster-provider-key" \ - --main-api-key "sk-your-main-provider-key" \ - --cluster-model "your-cluster-model-name" \ - --main-model "your-main-model-name" \ - --cluster-base-url "https://api.your-provider.com/v1" \ - --main-base-url "https://api.your-provider.com/v1" -``` - -> **Note:** Replace the model names, base URLs, and API keys with values for your actual LLM provider. CodeWiki does not ship with default credentials — you must supply your own. - -Confirm the configuration was saved and is complete: - -```bash -codewiki config validate -``` - -## Step 3: Generate Documentation for a Repository - -Navigate to any Git repository you want to document, then run: - -```bash -cd /path/to/your/project -codewiki generate -``` - -By default, output is written to `./docs`. The CLI runs through four staged checks and then the documentation pipeline itself: - -```text -Validating configuration... -Validating repository... -Analyzing dependencies... -Generating documentation... -``` - -## Expected Output - -After a successful run, you should see a `docs/` directory in your project containing: - -- Markdown files for each analyzed module (leaf modules first, then parent overview pages) -- A `module_tree.json` describing the hierarchical module structure -- A `metadata.json` describing job statistics and status - -## Example: Verbose Run with Custom Output - -```bash -codewiki generate --output ./generated-docs --verbose -``` - -Add `--verbose` any time you want detailed stage-by-stage progress and debug information printed to your terminal. - -## Example: Generate a GitHub Pages Site - -```bash -codewiki generate --github-pages --create-branch -``` - -This additionally renders a self-contained `index.html` viewer (from the generated `module_tree.json` and `metadata.json`) and creates a timestamped Git branch for the documentation changes, ready to push and open a pull request. - -## Running the Web Application Instead - -If you prefer the hosted web workflow (submit a GitHub repo URL through a browser, track job status, and view cached results), you can run the FastAPI app directly: - -```bash -python codewiki/run_web_app.py -``` - -Or via Docker Compose: - -```bash -cd docker -docker compose up --build -``` - -The web app listens on port `8000` by default (configurable via the `APP_PORT` environment variable read by `docker/docker-compose.yml`). - -## Next Steps - -Once you've generated your first documentation set, continue to [First Steps](first-steps.md) to learn about customizing what gets documented, exploring the CLI's other options, and where to find help. diff --git a/docs/reference/architecture/.gitignore b/docs/reference/architecture/.gitignore deleted file mode 100644 index b4f4c7a6..00000000 --- a/docs/reference/architecture/.gitignore +++ /dev/null @@ -1,8 +0,0 @@ -# CodeWiki temp files (dependency graphs can be 7GB+) -temp/ -dependency_graphs/ - -# JSON intermediate files (except schema/config) -*.json -!*-schema.json -!*-config.json diff --git a/docs/reference/architecture/README.md b/docs/reference/architecture/README.md deleted file mode 100644 index 6d9394e0..00000000 --- a/docs/reference/architecture/README.md +++ /dev/null @@ -1,123 +0,0 @@ -# CodeWiki Overview - -CodeWiki is an AI-powered documentation generator for source repositories. It analyzes a codebase, builds dependency relationships, groups components into meaningful modules, and uses LLM-backed agents to produce structured Markdown documentation with navigation metadata and optional static HTML output. - -The repository supports two primary entry points: - -- **CLI workflow** for local repository documentation, configuration management, Git workflows, progress reporting, and static-site generation. -- **Web workflow** for submitting GitHub repositories, queueing asynchronous generation jobs, caching results, and serving generated documentation. - -## End-to-End Architecture - -```mermaid -flowchart TD - User["Developer or Web User"] --> Entry{{"Choose entry point?"}} - Entry -->|CLI| CLI["CLI Core"] - Entry -->|Web| Frontend["Frontend Core"] - - CLI --> RuntimeConfig["Config Core"] - Frontend --> RuntimeConfig - - CLI --> GitOps["Git Integration"] - Frontend --> RepoProcessor["GitHub Repository Processor"] - - GitOps --> Source["Source Repository"] - RepoProcessor --> Source - - RuntimeConfig --> Generator["DocumentationGenerator"] - Source --> Generator - - Generator --> Analysis["Dependency Analysis"] - Analysis --> Parsers["Language Parsers"] - Parsers --> Graph["Dependency Graph"] - - Graph --> Clustering["Module Clustering"] - Clustering --> Agents["LLM Agent Orchestration"] - Agents --> Docs["Markdown Documentation"] - - Docs --> Metadata["Module Tree and Metadata"] - Metadata --> HTML["Optional HTML Viewer"] - Docs --> Output["Generated Documentation Output"] - HTML --> Output -``` - -## Documentation Generation Flow - -```mermaid -sequenceDiagram - participant Caller - participant Config as "Config Core" - participant Generator as "DocumentationGenerator" - participant Analyzer as "Dependency Analyzer" - participant Agents as "AgentOrchestrator" - participant Output as "Documentation Output" - - Caller->>Config: Build validated runtime configuration - Caller->>Generator: Start documentation run - Generator->>Analyzer: Analyze repository files and calls - Analyzer-->>Generator: Dependency graph and leaf components - Generator->>Generator: Cluster components into modules - Generator->>Agents: Generate leaf module documentation - Agents-->>Generator: Generated module content - Generator->>Output: Write Markdown, module tree, and metadata - Generator-->>Caller: Generation complete -``` - -## Core Modules - -| Module | Purpose | Source | -|---|---|---| -| [CLI Core](cli-core.md) | Command-line orchestration, local configuration, Git integration, terminal progress, and static HTML generation. | [`codewiki/cli`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/cli) | -| [Backend Core](backend-core.md) | Repository analysis, dependency-graph construction, module clustering, LLM agent orchestration, and documentation generation. | [`codewiki/src/be`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/src/be) | -| [Frontend Core](frontend-core.md) | FastAPI web application, repository submission, background job processing, caching, and documentation serving. | [`codewiki/src/fe`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/src/fe) | -| [Config Core](config-core.md) | Shared runtime `Config` model for source paths, output locations, provider settings, token limits, and agent instructions. | [`codewiki/src`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/src) | -| [Test Multi Path](test-multi-path.md) | Fixtures and executable checks for analysis across multiple source roots. | [`test-multi-path`](https://github.com/flamingo-stack/CodeWiki/tree/main/test-multi-path) | -| [Test Clustering](test-clustering.md) | Diagnostic and validation scripts for LLM-based module clustering and component-ID normalization. | [`test_clustering`](https://github.com/flamingo-stack/CodeWiki/tree/main/test_clustering) | - -## Module Relationships - -```mermaid -flowchart LR - Config["Config Core"] --> CLI["CLI Core"] - Config --> Frontend["Frontend Core"] - Config --> Backend["Backend Core"] - - CLI --> Backend - Frontend --> Backend - - Backend --> AgentTools["Agent Tools Core"] - Backend --> Analyzer["Dependency Analyzer Core"] - Backend --> Language["Tree-sitter Analyzers"] - Backend --> LLM["LLM Services"] - - TestPaths["Test Multi Path"] --> Config - TestPaths --> Analyzer - - TestCluster["Test Clustering"] --> Config - TestCluster --> Backend -``` - -## Backend Documentation References - -The Backend Core module is decomposed into focused subsystems: - -- [Agent Tools Core](backend-core/agent-tools-core/agent-tools-core.md) — controlled repository inspection and documentation editing tools. -- [Dependency Analyzer Core](backend-core/dependency-analyzer-core/dependency-analyzer-core.md) — file discovery, AST parsing, call analysis, and graph construction. -- [Tree Sitter Analyzers](backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md) — language support for C, C++, C#, Java, JavaScript, TypeScript, PHP, and Python. -- [Dependency Analyzer Models](backend-core/dependency-analyzer-models/dependency-analyzer-models.md) — repository, node, relationship, and analysis-result contracts. -- [Documentation Generator](backend-core/documentation-generator/documentation-generator.md) — top-level pipeline coordination and documentation output. -- [LLM Services](backend-core/llm-services/llm-services.md) — LLM model selection, request counting, and fallback handling. -- [Logging Config](backend-core/logging-config/logging-config.md) — shared colorized logging support. - -## CLI Documentation References - -- [Generation](cli-core/generation/generation.md) — CLI adapter for staged documentation generation. -- [Configuration](cli-core/configuration/configuration.md) — persisted settings, keyring-backed credentials, and agent instructions. -- [Job Models](cli-core/job_models/job_models.md) — generation job status, statistics, and LLM configuration models. -- [Git Integration](cli-core/git_integration/git_integration.md) — clean-tree checks, branch creation, commits, and remote URL handling. -- [HTML Generation](cli-core/html_generation/html_generation.md) — static documentation viewer generation. -- [Utils](cli-core/utils/utils.md) — terminal logging and progress tracking. - -## Summary - -CodeWiki separates user-facing workflows from its reusable generation engine. CLI and web layers construct a shared runtime configuration and delegate to Backend Core, which transforms source code into dependency-aware, module-oriented documentation. Test modules provide targeted coverage for multi-root analysis and the LLM-driven clustering stage. \ No newline at end of file diff --git a/docs/reference/architecture/backend-core/agent-tools-core/agent-tools-core.md b/docs/reference/architecture/backend-core/agent-tools-core/agent-tools-core.md deleted file mode 100644 index 0301cf95..00000000 --- a/docs/reference/architecture/backend-core/agent-tools-core/agent-tools-core.md +++ /dev/null @@ -1,123 +0,0 @@ -# Agent Tools Core - -## Purpose - -The Agent Tools Core module provides the foundational toolset used by the CodeWiki documentation-generation agent to interact with the filesystem during automated documentation authoring. It defines: - -- **`CodeWikiDeps`** — the shared dependency container passed to every agent tool invocation, carrying paths, the component registry, module-tree position, and LLM configuration. -- **`EditTool`** (and its helpers `Filemap`, `WindowExpander`) — a filesystem editor that lets the documentation agent view source/repository files and view, create, and edit documentation files, mirroring the SWE-agent `str_replace_editor` tool contract used by Anthropic-compatible agent frameworks. - -This module is a direct child of [Backend Core](../backend-core.md) and is consumed by the [Agent Orchestrator](../backend-core.md) during the documentation generation pipeline orchestrated for each module in the [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md) output tree. - -## Architecture Overview - -```mermaid -flowchart TD - Orchestrator["AgentOrchestrator"] -->|"builds"| Deps["CodeWikiDeps"] - Orchestrator -->|"invokes tool with"| ToolCall["str_replace_editor Tool Call"] - ToolCall -->|"uses ctx.deps"| Deps - ToolCall -->|"instantiates"| EditTool["EditTool"] - EditTool -->|"large .py view"| Filemap["Filemap"] - EditTool -->|"view/edit windows"| WindowExpander["WindowExpander"] - EditTool -->|"reads"| RepoFS[("Repository Files")] - EditTool -->|"reads/writes"| DocsFS[("Generated Docs Files")] - ToolCall -->|"post-edit .md check"| MermaidValidator["validate_mermaid_diagrams"] - Deps -->|"references"| Config["Config"] - Deps -->|"references"| Node["Node (component registry)"] -``` - -The module bridges the AI agent's tool-calling layer (built on `pydantic_ai`) and the local filesystem, enforcing safety constraints (absolute paths only, read-only access to the source repository, write access limited to the docs output directory) so that the agent can safely author markdown documentation while inspecting source code. - -## Core Components - -### `CodeWikiDeps` — Agent Dependency Container - -`CodeWikiDeps` (in `codewiki/src/be/agent_tools/deps.py`) is a `dataclass` that bundles everything a tool call needs to operate correctly within the current documentation generation run: - -| Field | Purpose | -|---|---| -| `absolute_docs_path` | Root directory where generated documentation is written | -| `absolute_repo_path` | Root directory of the analyzed source repository (read-only) | -| `registry` | Shared mutable dict used for cross-call state (e.g., file edit history, mermaid validation retry counters) | -| `components` | Mapping of component identifiers to `Node` objects from the dependency graph | -| `path_to_current_module` / `current_module_name` | Position of the module currently being documented within the module tree | -| `module_tree` | The full hierarchical module tree being documented | -| `max_depth` / `current_depth` | Recursion bounds for nested sub-module documentation generation | -| `config` | LLM configuration (`Config`) used to drive model behavior | -| `custom_instructions` | Optional user-supplied instructions injected into the agent's prompt | - -This container is instantiated once per documentation job/module by the orchestrator and passed as `ctx.deps` into every `pydantic_ai.RunContext` during a tool call, giving each tool a consistent view of the run's state without global variables. - -### `EditTool`, `Filemap`, and `WindowExpander` — Filesystem Editor - -The `str_replace_editor` tool (in `codewiki/src/be/agent_tools/str_replace_editor.py`) is adapted from the SWE-agent reference implementation and exposes five commands to the agent: - -- **`view`** — Displays a file (with `cat -n`-style line numbers) or lists a directory up to two levels deep. -- **`create`** — Creates a new file, auto-creating any missing parent directories (used for hierarchical docs output). -- **`str_replace`** — Replaces a unique occurrence of `old_str` with `new_str` in a file; rejects ambiguous or missing matches. -- **`insert`** — Inserts text at a specific line number. -- **`undo_edit`** — Reverts the most recent edit, using a per-file history stack persisted in `CodeWikiDeps.registry`. - -**Safety model:** The tool enforces `working_dir` semantics — when `working_dir="repo"`, only `view` is permitted (the source repository is never mutated); when `working_dir="docs"`, all commands are available against the documentation output tree. All paths must resolve to absolute paths under the appropriate root, and a leading-slash stripping fix prevents `Path` composition bugs where an absolute-looking relative path could escape the intended docs root. - -```mermaid -sequenceDiagram - participant Agent - participant Tool as "str_replace_editor" - participant Edit as "EditTool" - participant FS as "Filesystem" - participant Validator as "Mermaid Validator" - - Agent->>Tool: command="create", working_dir="docs", path="module.md" - Tool->>Tool: resolve absolute_path under docs root - Tool->>Edit: EditTool(registry, docs_path) - Edit->>Edit: validate_path(command, path) - Edit->>FS: create parent dirs + write_file - FS-->>Edit: written - Edit-->>Tool: success log - Tool->>Validator: validate_mermaid_diagrams(path) - Validator-->>Tool: validation result - Tool-->>Agent: combined result string -``` - -#### Supporting Helper Classes - -- **`Filemap`**: Uses `tree-sitter` to parse Python source and elide long function bodies (`>= 5` lines) when a `.py` file exceeds the response length limit, producing a condensed "filemap" view so the agent can navigate large files without exhausting context. This is currently gated behind the `USE_FILEMAP` flag. -- **`WindowExpander`**: Expands a requested `view_range` or edit snippet window outward to natural code boundaries (blank lines, `def`/`class`/decorator lines for Python) so that partial views don't cut a function or class definition in half. Expansion size is controlled by `MAX_WINDOW_EXPANSION_VIEW` / `MAX_WINDOW_EXPANSION_EDIT_CONFIRM` (both set to `0` by default, effectively disabling automatic expansion in the current configuration). - -### Mermaid Validation Hook - -After any non-`view` command that touches a `.md` file under `working_dir="docs"`, the tool asynchronously calls `validate_mermaid_diagrams` to catch malformed Mermaid diagrams as soon as the agent writes them. A per-file retry counter is stored in `ctx.deps.registry` (`mermaid_attempts:{path}`) to cap validation retries at `MAX_MERMAID_ATTEMPTS = 3`, preventing infinite fix-and-retry loops when the agent cannot resolve a diagram syntax error. - -```mermaid -flowchart LR - Edit["Edit .md file"] --> Check{{"path ends with .md?"}} - Check -->|"no"| Done["Return result"] - Check -->|"yes"| Attempts["Read mermaid_attempts counter"] - Attempts --> Limit{{"attempts >= MAX?"}} - Limit -->|"yes"| Skip["Skip validation, warn agent"] - Limit -->|"no"| Validate["validate_mermaid_diagrams"] - Validate --> HasError{{"errors found?"}} - HasError -->|"yes"| Increment["Increment counter in registry"] - HasError -->|"no"| Reset["Reset counter to 0"] - Increment --> Done - Reset --> Done - Skip --> Done -``` - -### Optional Linting Integration - -The module includes `flake8`-based linting utilities (`Flake8Error`, `format_flake8_output`, `flake8`) that can compare pre- and post-edit lint output for Python files and surface newly-introduced errors to the agent after a `str_replace` edit. This is gated by the `USE_LINTER` flag and is primarily relevant when the agent edits Python source rather than markdown documentation. - -## Integration with the Wider System - -- **[Backend Core](../backend-core.md)**: The parent module's `AgentOrchestrator` constructs `CodeWikiDeps` for each documentation job and registers the `str_replace_editor_tool` (a `pydantic_ai.Tool`) with the agent, along with the component registry produced by the [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md) and its supporting [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) (`Node`, `Repository`, etc.). -- **Configuration**: `CodeWikiDeps.config` is populated from the top-level `Config` object (see [Config Core](../../config-core.md)), which supplies LLM provider settings used throughout the generation pipeline. -- **Documentation Output**: Every markdown file produced by the [Documentation Generator](../documentation-generator/documentation-generator.md) subsystem is ultimately written to disk via `EditTool.create_file` / `EditTool.str_replace`, making this module the final write path for all agent-authored documentation. - -## Key Design Considerations - -1. **Read-only source, writable docs**: The `working_dir` parameter is the sole gate that separates safe read-only inspection of the analyzed repository from the mutable documentation workspace, preventing the agent from accidentally modifying source code. -2. **Absolute path enforcement**: All paths passed to `EditTool` must be absolute, with an explicit fix to prevent `pathlib.Path`'s "absolute path discards base" behavior from silently escaping the intended root directory. -3. **Bounded retries for generated content**: Both the mermaid-validation retry counter and the lint-comparison logic are designed to give the agent actionable feedback without entering unbounded correction loops. -4. **Stateful history via registry**: Rather than relying on in-process instance state (which would not survive across tool calls within the same agent run), file edit history and validation counters are persisted in the shared `CodeWikiDeps.registry` dict, keyed by file path. diff --git a/docs/reference/architecture/backend-core/backend-core.md b/docs/reference/architecture/backend-core/backend-core.md deleted file mode 100644 index 3d844e7f..00000000 --- a/docs/reference/architecture/backend-core/backend-core.md +++ /dev/null @@ -1,98 +0,0 @@ -# Backend Core - -Backend Core is CodeWiki’s server-side documentation pipeline. Located in `codewiki/src/be`, it analyzes source repositories, constructs dependency graphs, organizes components into modules, invokes LLM-backed agents to write documentation, and provides the filesystem tools and logging needed to run those workflows safely. - -Its primary entry point is `DocumentationGenerator`, which coordinates repository analysis and hierarchical documentation output. `AgentOrchestrator` manages agent-driven generation for individual modules, while `CountingFallbackModel` provides resilient LLM execution. - -## Architecture - -```mermaid -flowchart TD - Input["Source Repository"] --> Generator["DocumentationGenerator"] - Generator --> GraphBuilder["DependencyGraphBuilder"] - GraphBuilder --> Parser["DependencyParser"] - Parser --> Analysis["AnalysisService"] - Analysis --> RepoAnalyzer["RepoAnalyzer"] - Analysis --> CallAnalyzer["CallGraphAnalyzer"] - CallAnalyzer --> LanguageAnalyzers["Tree-sitter and Python Analyzers"] - LanguageAnalyzers --> Models["Dependency Analyzer Models"] - - GraphBuilder --> Components["Component Dependency Graph"] - Components --> Generator - - Generator --> Orchestrator["AgentOrchestrator"] - Orchestrator --> AgentTools["Agent Tools Core"] - Orchestrator --> LLM["CountingFallbackModel"] - LLM --> Provider["Configured LLM Provider"] - - AgentTools --> Docs["Generated Markdown Documentation"] - Generator --> Docs -``` - -## Documentation Generation Flow - -```mermaid -sequenceDiagram - participant Caller - participant Generator as "DocumentationGenerator" - participant Builder as "DependencyGraphBuilder" - participant Agent as "AgentOrchestrator" - participant Tools as "Agent Tools" - participant Output as "Documentation Files" - - Caller->>Generator: run() - Generator->>Builder: build_dependency_graph() - Builder-->>Generator: components and leaf nodes - Generator->>Generator: cluster modules and order leaves first - Generator->>Agent: process leaf module - Agent->>Tools: inspect repository and write docs - Tools-->>Agent: tool results - Agent-->>Generator: module documentation complete - Generator->>Output: write parent overviews and metadata - Generator-->>Caller: documentation complete -``` - -## Core Components - -| Component | Responsibility | -|---|---| -| `AgentOrchestrator` | Creates and runs documentation-writing agents for module-level generation. | -| `DocumentationGenerator` | Coordinates graph building, module clustering, leaf-first generation, parent overviews, and metadata output. | -| `AnalysisService` | Orchestrates repository structure analysis and call-graph extraction. | -| `RepoAnalyzer` | Discovers repository files and builds filtered file-tree representations. | -| `CallGraphAnalyzer` | Routes source files to language analyzers and aggregates call relationships. | -| `DependencyParser` | Converts analysis results into namespaced dependency-graph components. | -| `DependencyGraphBuilder` | Builds, validates, filters, and persists the repository dependency graph. | -| `CountingFallbackModel` | Wraps LLM models with request counting and automatic fallback behavior. | -| `CodeWikiDeps` | Carries shared run context and configuration into agent tools. | -| `EditTool` | Provides controlled repository viewing and documentation-file editing for agents. | -| `ColoredFormatter` | Produces readable, colorized backend console logs. | - -## Backend Subsystems - -- [Agent Tools Core](agent-tools-core/agent-tools-core.md) — Safe filesystem interaction through `CodeWikiDeps`, `EditTool`, `Filemap`, and `WindowExpander`. -- [Dependency Analyzer Core](dependency-analyzer-core/dependency-analyzer-core.md) — End-to-end repository analysis and dependency-graph construction. -- [Tree Sitter Analyzers](tree-sitter-analyzers/tree-sitter-analyzers.md) — Language-specific parsing for C, C++, C#, Java, JavaScript, TypeScript, PHP, and Python. -- [Dependency Analyzer Models](dependency-analyzer-models/dependency-analyzer-models.md) — Shared `Node`, `CallRelationship`, `Repository`, and analysis-result contracts. -- [Documentation Generator](documentation-generator/documentation-generator.md) — The top-level documentation generation workflow. -- [LLM Services](llm-services/llm-services.md) — Model factories, fallback handling, token configuration, and direct LLM calls. -- [Logging Config](logging-config/logging-config.md) — Shared colorized logging utilities. - -## Component Relationships - -```mermaid -flowchart LR - Models["Dependency Analyzer Models"] --> Analyzer["Dependency Analyzer Core"] - Language["Tree-sitter Analyzers"] --> Analyzer - Logging["Logging Config"] -.-> Analyzer - - Analyzer --> Generator["Documentation Generator"] - LLM["LLM Services"] --> Generator - Generator --> Orchestrator["AgentOrchestrator"] - Orchestrator --> Tools["Agent Tools Core"] - Tools --> Output["Markdown Documentation"] -``` - -## Source Location - -Backend Core source is maintained under [`codewiki/src/be`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/src/be). The module is consumed by the CLI and frontend layers to transform a repository into structured, navigable documentation. \ No newline at end of file diff --git a/docs/reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md b/docs/reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md deleted file mode 100644 index ff7af08d..00000000 --- a/docs/reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md +++ /dev/null @@ -1,274 +0,0 @@ -# Analysis Pipeline - -The Analysis Pipeline module is the orchestration layer responsible for turning a raw repository (either a cloned GitHub repository or a local folder) into a structured, language-aware call graph. It coordinates repository structure discovery, multi-language source parsing, and call-relationship resolution, producing the `AnalysisResult` data that downstream documentation and visualization components consume. - -This module sits inside the dependency analyzer subsystem of the backend and provides the primary entry point for "what does this codebase look like and how do its functions call each other?" - -## Purpose and Scope - -The Analysis Pipeline answers three questions for any supported repository: - -1. **What files exist?** — via structure analysis and pattern-based filtering. -2. **What functions/methods/classes exist in each file?** — via per-language AST parsing. -3. **How do those functions call each other?** — via cross-file/cross-language call relationship resolution. - -The pipeline does **not** implement language-specific parsing logic itself; it delegates that to the [Tree-Sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) module. It also does not define the data model shapes it operates on — those come from [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md). Instead, this module focuses purely on **orchestration**: cloning, filtering, dispatching, aggregating, and cleaning up. - -## Core Components - -| Component | Responsibility | -|---|---| -| `AnalysisService` | Top-level façade that orchestrates the full analysis workflow: clone → structure analysis → call graph analysis → README extraction → result assembly → cleanup | -| `RepoAnalyzer` | Builds a filtered file tree for a repository (or multiple repositories), applying include/exclude glob patterns and computing summary statistics | -| `CallGraphAnalyzer` | Central multi-language orchestrator that extracts code files from a file tree, dispatches each file to the correct language analyzer, aggregates functions/relationships, resolves call targets, deduplicates edges, and generates visualization data | - -## Architecture - -`AnalysisService` is the public API of this module. It composes a `RepoAnalyzer` (created per-call, since include/exclude patterns can vary) and a long-lived `CallGraphAnalyzer` instance. `CallGraphAnalyzer` never parses source code itself — it lazily imports the language-specific analyzer function (e.g. `analyze_python_file`, `analyze_javascript_file_treesitter`) from the [Tree-Sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) module at call time, keeping this module decoupled from heavy per-language parser dependencies. - -```mermaid -flowchart TD - Client["Caller (e.g. Documentation Generator)"] --> AS["AnalysisService"] - AS -->|"clone_repository()"| Clone["Repository Cloning (analysis.cloning)"] - AS -->|"analyze_repository_structure()"| RA["RepoAnalyzer"] - AS -->|"analyze_code_files()"| CGA["CallGraphAnalyzer"] - RA -->|"file_tree"| CGA - CGA -->|"extract_code_files()"| Extract["Code File Extraction"] - CGA -->|"dispatch by language"| Analyzers["Language Analyzers"] - Analyzers --> TS["Tree-Sitter Analyzers module"] - CGA -->|"Node / CallRelationship"| Models["Dependency Analyzer Models module"] - AS -->|"AnalysisResult"| Models - AS -->|"Repository"| Models -``` - -## Component Relationships - -```mermaid -classDiagram - class AnalysisService { - +call_graph_analyzer: CallGraphAnalyzer - +analyze_local_repository(repo_path, max_files, languages) dict - +analyze_repository_full(github_url, include_patterns, exclude_patterns) AnalysisResult - +analyze_repository_structure_only(github_url, include_patterns, exclude_patterns) dict - +cleanup_all() - -_clone_repository(github_url) str - -_analyze_structure(repo_dir, include, exclude) dict - -_analyze_call_graph(file_tree, repo_dir) dict - -_read_readme_file(repo_dir) str - -_cleanup_repository(temp_dir) - } - class RepoAnalyzer { - +include_patterns: list - +exclude_patterns: list - +analyze_repository_structure(repo_dir) dict - -_build_file_tree(repo_dir) dict - -_should_exclude_path(path, filename) bool - -_should_include_file(path, filename) bool - -_count_files(tree) int - -_calculate_size(tree) float - } - class CallGraphAnalyzer { - +functions: dict~str, Node~ - +call_relationships: list~CallRelationship~ - +analyze_code_files(code_files, base_dir) dict - +extract_code_files(file_tree) list - -_analyze_code_file(repo_dir, file_info) - -_resolve_call_relationships() - -_deduplicate_relationships() - -_generate_visualization_data() dict - +generate_llm_format() dict - -_select_most_connected_nodes(target_count) - } - class AnalysisResult { - +repository: Repository - +functions: list - +relationships: list - +file_tree: dict - +summary: dict - +visualization: dict - +readme_content: str - } - class Node { - +id: str - +name: str - +node_type: str - +file_path: str - +component_id: str - +docstring: str - +parameters: list - } - class CallRelationship { - +caller: str - +callee: str - +call_line: int - +is_resolved: bool - } - - AnalysisService --> RepoAnalyzer : uses - AnalysisService --> CallGraphAnalyzer : uses - AnalysisService --> AnalysisResult : produces - CallGraphAnalyzer --> Node : produces - CallGraphAnalyzer --> CallRelationship : produces - AnalysisResult --> Node : contains - AnalysisResult --> CallRelationship : contains -``` - -`AnalysisResult`, `Node`, `CallRelationship`, and `Repository` are defined in the [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) module and are reused here as the canonical output contract. - -## Workflow: Full Repository Analysis - -`analyze_repository_full` is the primary entry point used when a complete call graph (functions + relationships + visualization) is required for a GitHub repository. It is a sequential pipeline with cleanup guaranteed on both success and failure paths. - -```mermaid -sequenceDiagram - participant Caller - participant AS as AnalysisService - participant Clone as "Cloning Utility" - participant RA as RepoAnalyzer - participant CGA as CallGraphAnalyzer - participant FS as Filesystem - - Caller->>AS: analyze_repository_full(github_url) - AS->>Clone: clone_repository(github_url) - Clone-->>AS: temp_dir - AS->>AS: parse_github_url(github_url) - AS->>RA: analyze_repository_structure(temp_dir) - RA->>FS: walk directory tree - RA-->>AS: file_tree + summary - AS->>CGA: extract_code_files(file_tree) - CGA-->>AS: code_files - AS->>CGA: analyze_code_files(code_files, temp_dir) - CGA->>CGA: dispatch per language, parse each file - CGA->>CGA: resolve_call_relationships() - CGA->>CGA: deduplicate_relationships() - CGA->>CGA: generate_visualization_data() - CGA-->>AS: functions, relationships, call_graph, visualization - AS->>FS: read README file - AS->>AS: build AnalysisResult - AS->>Clone: cleanup_repository(temp_dir) - AS-->>Caller: AnalysisResult -``` - -If any step raises an exception, `AnalysisService` still cleans up the temporary clone directory before re-raising a `RuntimeError`, preventing orphaned checkouts from accumulating on disk. - -## Workflow: Structure-Only and Local Analysis - -Two lighter-weight workflows exist alongside the full analysis path: - -- **`analyze_repository_structure_only`** clones a GitHub repository and runs only `RepoAnalyzer`, skipping call graph generation entirely. This is used when only the file tree and summary statistics are needed (e.g. for quick previews). -- **`analyze_local_repository`** skips cloning altogether and operates directly on a local folder path, optionally filtering by `languages` and capping the number of files via `max_files`. It returns a simplified dict (`nodes`, `relationships`, `summary`) rather than a full `AnalysisResult`. - -```mermaid -flowchart LR - subgraph Full["analyze_repository_full"] - F1["Clone"] --> F2["Structure Analysis"] --> F3["Call Graph Analysis"] --> F4["README Read"] --> F5["AnalysisResult"] - end - subgraph StructureOnly["analyze_repository_structure_only"] - S1["Clone"] --> S2["Structure Analysis"] --> S3["Structure Dict"] - end - subgraph Local["analyze_local_repository"] - L1["Local Path"] --> L2["Structure Analysis"] --> L3["Filter by language / max_files"] --> L4["Call Graph Analysis"] --> L5["Simplified Dict"] - end -``` - -## Repository Structure Analysis (`RepoAnalyzer`) - -`RepoAnalyzer` walks a repository directory recursively and builds a nested `file_tree` dictionary, applying two categories of patterns: - -- **Include patterns**: if explicitly provided, they *replace* the defaults and restrict results to matching files only. -- **Exclude patterns**: if provided, they are *merged* with a built-in default ignore list (e.g. `.git`, `node_modules`). - -It defends against unsafe filesystem traversal by rejecting symlinks and any resolved path that escapes the repository root. - -`RepoAnalyzer` also supports **multi-repository analysis**: when given a list of paths instead of a single path, it builds a namespaced, merged tree — each repository's subtree is wrapped with a namespace derived from its folder name, and summary statistics (`total_files`, `total_size_kb`, `repositories`, `namespaces`) are aggregated across all inputs. - -```mermaid -flowchart TD - Input["repo_dir: str or list[str]"] --> Check{{"Single path?"}} - Check -->|"Yes"| Single["_build_file_tree(repo_dir)"] - Check -->|"No"| Multi["_analyze_multiple_repositories(repo_dirs)"] - Multi --> Loop["For each repo_dir: compute namespace, build tree"] - Loop --> Wrap["Wrap tree with namespace prefix"] - Wrap --> Merge["Merge into single root tree"] - Single --> Summary1["Compute total_files, total_size_kb"] - Merge --> Summary2["Aggregate totals across repositories"] - Summary1 --> Out["file_tree + summary"] - Summary2 --> Out -``` - -Path filtering combines glob matching (`fnmatch`) on both the full relative path and the bare filename, along with segment-level and prefix checks, so patterns like `node_modules`, `*.test.js`, or `build/` are all honored. - -## Call Graph Analysis (`CallGraphAnalyzer`) - -`CallGraphAnalyzer` is the multi-language orchestrator. Its responsibilities, in order: - -1. **Extraction** — `extract_code_files` walks the file tree and filters files by known code extensions (mapped to language names via a shared extension table). -2. **Dispatch** — `_analyze_code_file` routes each file to a private per-language method (`_analyze_python_file`, `_analyze_javascript_file`, `_analyze_typescript_file`, `_analyze_java_file`, `_analyze_csharp_file`, `_analyze_c_file`, `_analyze_cpp_file`, `_analyze_php_file`). Each of these lazily imports the corresponding analyzer function from the [Tree-Sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) module and normalizes the returned functions into the shared `Node` map keyed by function ID. -3. **Resolution** — `_resolve_call_relationships` builds a lookup table from function ID, bare name, `component_id`, and trailing method name, then attempts to match every call relationship's `callee` string against it, marking matches as `is_resolved`. -4. **Deduplication** — `_deduplicate_relationships` removes duplicate `(caller, callee)` edges, keeping only the first occurrence. -5. **Visualization** — `_generate_visualization_data` emits Cytoscape.js-compatible `elements` (nodes classified by type/language, edges limited to resolved relationships) plus a summary of node/edge counts. - -```mermaid -flowchart TD - FT["file_tree"] --> Extract["extract_code_files()"] - Extract --> CodeFiles["code_files list"] - CodeFiles --> Dispatch{{"Dispatch by language"}} - Dispatch -->|"python"| PyA["Python AST Analyzer"] - Dispatch -->|"javascript"| JsA["JS Tree-Sitter Analyzer"] - Dispatch -->|"typescript"| TsA["TS Tree-Sitter Analyzer"] - Dispatch -->|"java"| JavaA["Java Tree-Sitter Analyzer"] - Dispatch -->|"csharp"| CsA["C# Tree-Sitter Analyzer"] - Dispatch -->|"c"| CA["C Tree-Sitter Analyzer"] - Dispatch -->|"cpp"| CppA["C++ Tree-Sitter Analyzer"] - Dispatch -->|"php"| PhpA["PHP Tree-Sitter Analyzer"] - PyA --> Agg["functions dict, call_relationships list"] - JsA --> Agg - TsA --> Agg - JavaA --> Agg - CsA --> Agg - CA --> Agg - CppA --> Agg - PhpA --> Agg - Agg --> Resolve["_resolve_call_relationships()"] - Resolve --> Dedup["_deduplicate_relationships()"] - Dedup --> Viz["_generate_visualization_data()"] - Viz --> Result["functions, relationships, call_graph, visualization"] -``` - -Each per-language method (e.g. `_analyze_python_file`) follows the same normalization pattern: on success, returned `functions`/`relationships` are merged into `self.functions` (keyed by `func.id` or a synthesized `path:name` fallback) and appended to `self.call_relationships`; on failure, the exception is logged and analysis continues for remaining files, ensuring a single malformed file cannot abort the whole run. - -`_filter_supported_languages` (in `AnalysisService`) additionally narrows the extracted code files to a fixed set of supported languages (`python`, `javascript`, `typescript`, `java`, `csharp`, `c`, `cpp`, `php`, `go`, `rust`) before invoking the call graph analyzer, and reports how many files were skipped as unsupported. - -## Call Relationship Resolution Logic - -Resolving a raw `callee` string (as extracted from source, e.g. a bare function name, a dotted method reference, or a fully-qualified component ID) into an actual function node is the most delicate part of the pipeline, since different language analyzers may extract call targets in different formats. - -```mermaid -flowchart TD - Start["callee: str"] --> Direct{{"callee in func_lookup?"}} - Direct -->|"Yes"| Resolved["Mark resolved, rewrite callee to func_id"] - Direct -->|"No"| HasDot{{"Contains '.'?"}} - HasDot -->|"No"| Unresolved["Leave unresolved"] - HasDot -->|"Yes"| MethodName["Extract trailing segment after last '.'"] - MethodName --> MethodLookup{{"method_name in func_lookup?"}} - MethodLookup -->|"Yes"| Resolved - MethodLookup -->|"No"| Unresolved -``` - -The lookup table (`func_lookup`) is populated with multiple keys per function — its ID, bare name, `component_id`, and the last dotted segment of `component_id` — to maximize the chance of matching call sites extracted with varying levels of qualification across languages. - -## Cleanup and Resource Management - -`AnalysisService` tracks every temporary clone directory it creates in `_temp_directories`. Both `analyze_repository_full` and `analyze_repository_structure_only` clean up their temp directory in the `except` branch as well as on the success path, and `cleanup_all()` (also invoked from `__del__`) provides a final safety net to remove any directories that were never explicitly cleaned up — guarding against disk space leaks from crashed or aborted analysis runs. - -## Backward-Compatible Function API - -Two module-level functions, `analyze_repository` and `analyze_repository_structure_only`, wrap `AnalysisService` for callers that expect the older tuple-returning function signature (`(result, temp_dir)`), returning `None` in place of `temp_dir` since cleanup is now handled internally by the service. - -## Relationship to Other Modules - -- **[Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md)** — supplies the `AnalysisResult`, `Node`, `CallRelationship`, and `Repository` data structures produced by this pipeline. -- **[Tree-Sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md)** — supplies the actual per-language parsing functions (`analyze_python_file`, `analyze_javascript_file_treesitter`, etc.) that `CallGraphAnalyzer` dispatches to. -- **[Graph Construction](../graph_construction/graph_construction.md)** — a sibling module handling AST parsing and dependency graph assembly (`DependencyParser`, `DependencyGraphBuilder`); the Analysis Pipeline focuses on repository-level orchestration and call graph resolution, while graph construction focuses on assembling the broader dependency graph from parsed nodes. -- **[Dependency Analyzer Core](../dependency-analyzer-core.md)** — the parent module that groups the Analysis Pipeline together with Graph Construction as the two core analysis capabilities of the dependency analyzer subsystem. -- **[Documentation Generator](../../documentation-generator/documentation-generator.md)** — a consumer of `AnalysisResult` data produced by this pipeline, using it as input for generating documentation content. diff --git a/docs/reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md b/docs/reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md deleted file mode 100644 index 3e4335de..00000000 --- a/docs/reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md +++ /dev/null @@ -1,124 +0,0 @@ -# Dependency Analyzer Core - -## Overview - -The Dependency Analyzer Core module is the central orchestration engine of the CodeWiki backend's static analysis pipeline. It coordinates the full lifecycle of turning a raw source repository into a structured, queryable **dependency graph** of code components (functions, classes, methods, interfaces, etc.) connected by call/usage relationships. - -This module answers three fundamental questions for any supported repository: - -1. **What files exist and which are relevant?** — repository structure discovery and filtering. -2. **What code components exist, and who calls whom?** — multi-language AST/call-graph extraction. -3. **How do these components form a navigable, deduplicated dependency graph?** — component construction, namespacing, and graph persistence. - -The resulting dependency graph is the primary input consumed by downstream stages of the CodeWiki backend, including documentation generation and LLM-driven summarization (see the [Documentation Generator](../documentation-generator/documentation-generator.md) module in `backend-core`). - -## Position in the System - -Dependency Analyzer Core lives inside the `backend-core` module, alongside several closely related sibling modules: - -- [Tree-sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) — language-specific parsers (Python, JavaScript, TypeScript, Java, C#, C, C++, PHP) invoked by this module's call graph orchestrator. -- [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) — the shared `Node`, `CallRelationship`, `Repository`, `AnalysisResult`, and `NodeSelection` data models produced and consumed by this module. -- [Logging Config](../logging-config/logging-config.md) — shared logging formatting utilities used across the analysis pipeline. -- [Documentation Generator](../documentation-generator/documentation-generator.md) — consumes the dependency graph produced here to generate documentation. -- [Agent Tools Core](../agent-tools-core/agent-tools-core.md) — provides file-editing and repository interaction tools used by the broader backend agent workflow. -- [LLM Services](../llm-services/llm-services.md) — provides model routing/fallback used elsewhere in `backend-core`. - -For the full picture of how these modules interrelate, see the parent [Backend Core](../backend-core.md) documentation. - -## Architecture - -The module is organized into two cooperating sub-systems: - -1. **Analysis Pipeline** — discovers repository structure and extracts functions/classes and their call relationships for a single repository path. -2. **Graph Construction** — wraps the analysis pipeline to build fully-qualified, namespaced `Node` components (supporting both single-repository and multi-repository/dependency scenarios), resolves cross-namespace edges, validates graph completeness, and persists the final dependency graph to disk. - -```mermaid -flowchart TD - subgraph AnalysisPipeline["Analysis Pipeline"] - direction TB - RepoAnalyzer["RepoAnalyzer"] -->|"file_tree"| AnalysisService["AnalysisService"] - CallGraphAnalyzer["CallGraphAnalyzer"] -->|"functions + relationships"| AnalysisService - end - - subgraph GraphConstruction["Graph Construction"] - direction TB - DependencyParser["DependencyParser"] -->|"components (Node)"| DependencyGraphBuilder["DependencyGraphBuilder"] - end - - Config["Config"] -.->|"repo paths, patterns"| DependencyGraphBuilder - DependencyGraphBuilder -->|"drives"| DependencyParser - DependencyParser -->|"uses"| AnalysisService - CallGraphAnalyzer -->|"delegates per-language"| TreeSitterAnalyzers["Tree-sitter Analyzers"] - AnalysisService -->|"produces"| Models["Node / CallRelationship / AnalysisResult"] - DependencyGraphBuilder -->|"writes"| GraphFile[("dependency_graph.json")] - - click TreeSitterAnalyzers "../tree-sitter-analyzers/tree-sitter-analyzers.md" - click Models "../dependency-analyzer-models/dependency-analyzer-models.md" -``` - -### High-Level Data Flow - -```mermaid -sequenceDiagram - participant Builder as "DependencyGraphBuilder" - participant Parser as "DependencyParser" - participant Service as "AnalysisService" - participant RepoA as "RepoAnalyzer" - participant CallA as "CallGraphAnalyzer" - participant TS as "Tree-sitter Analyzers" - - Builder->>Parser: parse_repository() - Parser->>Service: _analyze_structure(repo_path) - Service->>RepoA: analyze_repository_structure() - RepoA-->>Service: file_tree + summary - Parser->>Service: _analyze_call_graph(file_tree, repo_path) - Service->>CallA: extract_code_files() / analyze_code_files() - CallA->>TS: analyze_python_file / analyze_javascript_file_treesitter / ... - TS-->>CallA: functions, relationships - CallA-->>Service: functions, relationships, visualization - Service-->>Parser: call_graph_result - Parser->>Parser: build namespaced Node components - Parser-->>Builder: components (Dict[str, Node]) - Builder->>Builder: build_graph_from_components / validate / filter leaves - Builder-->>Builder: (components, leaf_nodes) -``` - -## Sub-modules - -### Analysis Pipeline - -Handles repository cloning support, file-tree discovery with include/exclude filtering, and multi-language call graph extraction by delegating to per-language analyzers. Composed of `AnalysisService`, `RepoAnalyzer`, and `CallGraphAnalyzer`. - -See [Analysis Pipeline](analysis_pipeline.md) for full details. - -### Graph Construction - -Turns raw analysis results into fully-qualified, namespaced `Node` components, resolves single- and multi-repository (dependency-aware) call relationships, and drives graph validation, leaf-node filtering, and persistence to JSON. Composed of `DependencyParser` and `DependencyGraphBuilder`. - -See [Graph Construction](graph_construction.md) for full details. - -## Key Concepts - -### Fully-Qualified Domain Names (FQDNs) - -Every extracted code component is assigned a unique identifier of the form `namespace.module.path::ComponentName`. The `namespace` prefix is derived from the source repository's directory name and prevents ID collisions when analyzing multiple repositories together (e.g., a primary repo plus its dependencies). See [Graph Construction](graph_construction.md) for details on how `DependencyParser` builds and resolves these identifiers. - -### Multi-Language Support - -The pipeline supports Python, JavaScript, TypeScript, Java, C#, C, C++, PHP, Go, and Rust. Python uses a native AST analyzer; all other languages are handled via tree-sitter based analyzers described in [Tree-sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md). - -### Single-Path vs. Multi-Path Analysis - -`DependencyGraphBuilder` and `DependencyParser` both support analyzing either a single repository or multiple repository paths simultaneously (for example, a main repository plus one or more external dependency sources). In multi-path mode, each source directory is assigned its own namespace, and cross-namespace dependency edges are resolved after all components have been extracted. See [Graph Construction](graph_construction.md). - -## Summary - -| Component | Responsibility | -|---|---| -| `AnalysisService` | Top-level orchestration: repository cloning, structure + call-graph analysis, README extraction, cleanup | -| `RepoAnalyzer` | Builds filtered file trees for one or more repository paths | -| `CallGraphAnalyzer` | Routes files to per-language analyzers, resolves and deduplicates call relationships, builds visualization data | -| `DependencyParser` | Converts analysis results into namespaced `Node` components; resolves cross-namespace dependencies | -| `DependencyGraphBuilder` | Drives parsing end-to-end, validates graph completeness, filters leaf nodes, persists the dependency graph | - -For details on each area, see [Analysis Pipeline](analysis_pipeline.md) and [Graph Construction](graph_construction.md). diff --git a/docs/reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md b/docs/reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md deleted file mode 100644 index 12ef0668..00000000 --- a/docs/reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md +++ /dev/null @@ -1,218 +0,0 @@ -# Graph Construction - -The Graph Construction module is the stage-two engine of CodeWiki's dependency analysis pipeline. It transforms raw structural and call-graph data produced by the [Analysis Pipeline](../analysis_pipeline/analysis_pipeline.md) into a fully-resolved, persistable dependency graph made of typed `Node` objects. It is responsible for assigning fully-qualified domain names (FQDNs) to every code component, namespacing components when multiple repositories are analyzed together, resolving both intra- and cross-namespace dependency edges, and filtering the resulting graph down to a set of "leaf" components suitable for downstream clustering and documentation generation. - -This module contains two core components: - -- **`DependencyParser`** — converts raw analysis output (functions, classes, relationships) into namespaced `Node` objects and resolves dependency edges between them, supporting both single-repository and multi-repository ("multi-path") modes. -- **`DependencyGraphBuilder`** — orchestrates the end-to-end graph-building workflow: invoking the parser, persisting the graph to disk, building an in-memory traversal graph, validating its completeness, and filtering leaf nodes for downstream consumption. - -## Purpose and Scope - -The Graph Construction module sits between raw source analysis and the higher-level clustering/documentation stages of CodeWiki. Its responsibilities are strictly bounded to: - -1. **Component identity assignment** — converting analyzer-produced identifiers into canonical FQDNs of the form `{namespace}.{module.path}::{ComponentName}`. -2. **Namespace management** — when analyzing multiple source directories (multi-path mode), prefixing components by their originating repository/directory to avoid ID collisions. -3. **Dependency edge resolution** — mapping raw caller/callee identifiers extracted by language analyzers into resolved FQDN-to-FQDN edges, including matching across namespace boundaries by component name when a direct ID match is unavailable. -4. **Graph persistence** — serializing the final component graph to a deterministic (sorted) JSON representation on disk. -5. **Leaf node filtering** — identifying and validating the "leaf" components (typically classes, interfaces, structs, or functions for C-style codebases) that anchor the hierarchical documentation generation process. - -It does **not** perform language-specific parsing itself (that responsibility belongs to the [Tree-Sitter Analyzers](../../tree-sitter-analyzers/tree-sitter-analyzers.md)) nor does it run the structural/call-graph analysis (handled by `AnalysisService` in the [Analysis Pipeline](../analysis_pipeline/analysis_pipeline.md)). It also delegates the `Node` data model itself to [Dependency Analyzer Models](../../dependency-analyzer-models/dependency-analyzer-models.md). - -## Architecture - -```mermaid -flowchart TD - Config["Config"] --> Builder["DependencyGraphBuilder"] - Builder -->|"instantiates"| Parser["DependencyParser"] - Parser -->|"uses"| AnalysisSvc["AnalysisService"] - AnalysisSvc -->|"structure + call graph"| Parser - Parser -->|"produces"| Nodes["Node objects (Dict[str, Node])"] - Builder -->|"build_graph_from_components()"| TraversalGraph["In-memory traversal graph"] - Builder -->|"validate_graph_completeness()"| Validation["Graph validation"] - Builder -->|"get_leaf_nodes()"| LeafFilter["Leaf node filtering"] - Parser -->|"save_dependency_graph()"| JSONFile["dependency_graph.json"] - LeafFilter --> Output["(components, leaf_nodes)"] -``` - -`DependencyGraphBuilder` is constructed with a [Config](../../../config-core.md) instance that supplies the repository path(s), output directories, and include/exclude filter patterns. It then instantiates a `DependencyParser` scoped to either a single repository path (`str`) or a list of paths (multi-path mode), delegating all structural/call-graph analysis to `AnalysisService` from the [Analysis Pipeline](../analysis_pipeline/analysis_pipeline.md). - -## Core Components - -### DependencyParser - -`DependencyParser` is the component-extraction engine. It wraps an `AnalysisService` instance and converts its raw output into the canonical `Node` representation used throughout the rest of CodeWiki. - -**Initialization** accepts either a single repository path or a list of paths, along with optional `include_patterns` / `exclude_patterns` file filters: - -```python -parser = DependencyParser( - repo_path=["/path/to/main-repo", "/path/to/dependency-repo"], - include_patterns=["*.py", "*.ts"], - exclude_patterns=["*Tests*"] -) -components = parser.parse_repository() -``` - -**Key responsibilities:** - -| Method | Purpose | -|---|---| -| `parse_repository()` | Entry point; dispatches to single- or multi-path parsing based on how many repo paths were configured. | -| `_parse_single_repository()` | Backward-compatible path: analyzes structure and call graph for one repository, then builds components via `_build_components_from_analysis()`. | -| `_parse_multiple_repositories()` | Analyzes each configured path independently, namespaces its components, merges all namespaces together, then resolves cross-namespace dependencies. | -| `_build_namespaced_components()` | First pass creates `Node` objects with FQDN keys (`{namespace}.{original_id}`); second pass wires up `depends_on` edges within the same namespace. | -| `_resolve_cross_namespace_dependencies()` | For any dependency edge that didn't resolve within its own namespace, attempts a name-based match against components in *other* namespaces. | -| `_build_components_from_analysis()` | Single-path equivalent of namespaced component construction; also tracks legacy `file_path:name` IDs for backward-compatible dependency resolution. | -| `save_dependency_graph()` | Serializes all components (sorted by ID, with `depends_on` sets converted to sorted lists) to a JSON file for deterministic, diffable output. | - -#### FQDN Construction - -Every component is identified by a Fully Qualified Domain Name in the form: - -```text -{namespace}.{module.path}::{ComponentName} -``` - -- `namespace` is derived from the last path segment of the source directory being analyzed (e.g., `openframe-frontend`, `ui-kit`), computed by `_get_namespace_from_path()`. -- `module.path::ComponentName` is the raw identifier produced by the language-specific analyzer (see [Tree-Sitter Analyzers](../../tree-sitter-analyzers/tree-sitter-analyzers.md)). -- In single-path mode, `is_from_deps` is always `False`. In multi-path mode, the first configured path (`repo_index == 0`) is treated as the primary repository, while subsequent paths are marked `is_from_deps=True`. - -Each `Node` also retains `short_id` (the original, un-namespaced identifier) purely for display purposes, alongside the `namespace` string itself — see the [Dependency Analyzer Models](../../dependency-analyzer-models/dependency-analyzer-models.md) documentation for the full `Node` schema. - -#### Single-Path vs. Multi-Path Parsing - -```mermaid -flowchart TD - Start["parse_repository()"] --> Check{{"len(repo_paths) == 1?"}} - Check -->|"Yes"| Single["_parse_single_repository()"] - Check -->|"No"| Multi["_parse_multiple_repositories()"] - - Single --> S1["analysis_service._analyze_structure()"] - S1 --> S2["analysis_service._analyze_call_graph()"] - S2 --> S3["_build_components_from_analysis()"] - S3 --> SResult["self.components"] - - Multi --> M1["For each repo path: compute namespace"] - M1 --> M2["analysis_service._analyze_structure()"] - M2 --> M3["analysis_service._analyze_call_graph()"] - M3 --> M4["_build_namespaced_components()"] - M4 --> M5["Merge into all_components"] - M5 --> M6{{"More paths?"}} - M6 -->|"Yes"| M1 - M6 -->|"No"| M7["_resolve_cross_namespace_dependencies()"] - M7 --> MResult["self.components"] -``` - -In multi-path mode, each configured source directory is analyzed independently by `AnalysisService`, then merged into a single component dictionary keyed by namespaced FQDN. A `namespace_mapping` dict (original ID → FQDN) is built incrementally as each repository is processed, enabling within-namespace dependency resolution during `_build_namespaced_components()`. - -#### Cross-Namespace Dependency Resolution - -After all repositories are parsed and merged, `_resolve_cross_namespace_dependencies()` performs a second reconciliation pass. For every component's `depends_on` set: - -1. If the dependency ID already matches an entry in `all_components`, it is kept as-is. -2. Otherwise, the parser extracts the trailing component name (`dep_id.split(".")[-1]`) and searches all components for a name match belonging to a *different* namespace, logging the resolution as a cross-namespace dependency. -3. If no match is found at all, the original (unresolved) dependency ID is preserved rather than dropped, ensuring no data loss even when resolution fails. - -This name-based fallback allows dependency edges to survive even when different analyzers or namespaces produce slightly different identifier formats for the same logical target. - -### DependencyGraphBuilder - -`DependencyGraphBuilder` is the orchestration layer that wraps `DependencyParser` with graph persistence, traversal-graph construction, validation, and leaf-node filtering. It is the primary entry point used by the rest of the backend pipeline to obtain a ready-to-cluster dependency graph. - -```python -builder = DependencyGraphBuilder(config) -components, leaf_nodes = builder.build_dependency_graph() -``` - -**Workflow performed by `build_dependency_graph()`:** - -```mermaid -flowchart TD - A["build_dependency_graph()"] --> B["Ensure dependency_graph_dir exists"] - B --> C["Compute sanitized dependency_graph_path"] - C --> D["Resolve include/exclude patterns from Config"] - D --> E["Build repo_paths from config.all_source_paths"] - E --> F["Instantiate DependencyParser"] - F --> G["parser.parse_repository()"] - G --> H["Log component type breakdown"] - H --> I["parser.save_dependency_graph(path)"] - I --> J["build_graph_from_components(components)"] - J --> K["validate_graph_completeness(components, graph)"] - K --> L["get_leaf_nodes(graph, components)"] - L --> M["Determine valid leaf types from available component types"] - M --> N["Filter leaf_nodes: skip invalid / wrong-type / not-found"] - N --> O["Return (components, keep_leaf_nodes)"] -``` - -**Key behaviors:** - -- **Source path resolution**: Uses `config.all_source_paths` (primary `repo_path` plus any `additional_source_paths`) to decide whether to invoke single-path or multi-path parsing on `DependencyParser`. See [Config](../../../config-core.md) for how these paths are validated and exposed. -- **Deterministic output naming**: The dependency graph JSON file is named `{sanitized_repo_name}_dependency_graph.json`, where the repository's base directory name is sanitized to alphanumeric characters and underscores. -- **Graph traversal preparation**: After parsing, `build_graph_from_components()` and `get_leaf_nodes()` (internal graph/topology utilities) convert the flat `Node` dictionary into a traversable graph structure and extract nodes with no outgoing dependencies as candidate "leaves." -- **Post-build validation**: `validate_graph_completeness()` is invoked immediately after graph construction to catch structural inconsistencies before leaf filtering proceeds. -- **Leaf node type filtering**: Leaf nodes are only retained if their `component_type` is one of `class`, `interface`, or `struct` — unless *none* of the parsed components have those types (e.g., a purely procedural/C-style codebase), in which case `function` is also accepted. Nodes that are empty, contain error-like keywords (`error`, `exception`, `failed`, `invalid`), or are absent from the parsed `components` dictionary are skipped and logged with a specific reason. - -#### Leaf Node Filtering Logic - -```mermaid -flowchart TD - Start["For each leaf_node in leaf_nodes"] --> V1{{"Is leaf_node a valid non-empty string without error keywords?"}} - V1 -->|"No"| SkipInvalid["skipped_invalid += 1"] - V1 -->|"Yes"| V2{{"leaf_node in components?"}} - V2 -->|"No"| SkipNotFound["skipped_not_found += 1"] - V2 -->|"Yes"| V3{{"component_type in valid_types?"}} - V3 -->|"No"| SkipType["skipped_type += 1"] - V3 -->|"Yes"| Keep["keep_leaf_nodes.append(leaf_node)"] -``` - -The resulting `keep_leaf_nodes` list, together with the full `components` dictionary, is returned to the caller and forms the input for the clustering and documentation-generation stages that follow in the broader backend pipeline. - -## Data Model - -Both `DependencyParser` and `DependencyGraphBuilder` operate on the `Node` model defined in [Dependency Analyzer Models](../../dependency-analyzer-models/dependency-analyzer-models.md). The fields most relevant to graph construction are: - -| Field | Description | -|---|---| -| `id` | Primary key; the FQDN (`{namespace}.{original_id}`). | -| `short_id` | Original, un-namespaced identifier from the language analyzer. | -| `namespace` | Source directory namespace (e.g., repository directory name). | -| `is_from_deps` | `True` if the component came from a non-primary (dependency) source path in multi-path mode. | -| `component_type` | Classifies the node (`class`, `interface`, `struct`, `function`, `method`, etc.); drives leaf-node filtering. | -| `depends_on` | Set of FQDNs this component depends on; populated and reconciled by `DependencyParser`. | - -## Component Interactions - -```mermaid -sequenceDiagram - participant Caller as "Pipeline Caller" - participant Builder as "DependencyGraphBuilder" - participant Parser as "DependencyParser" - participant Svc as "AnalysisService" - participant FS as "File System" - - Caller->>Builder: build_dependency_graph() - Builder->>Parser: new DependencyParser(repo_paths, patterns) - Builder->>Parser: parse_repository() - Parser->>Svc: _analyze_structure(repo_path) - Svc-->>Parser: structure_result - Parser->>Svc: _analyze_call_graph(file_tree, repo_path) - Svc-->>Parser: call_graph_result - Parser->>Parser: _build_components_from_analysis() / _build_namespaced_components() - Parser-->>Builder: components (Dict[str, Node]) - Builder->>Parser: save_dependency_graph(path) - Parser->>FS: write JSON - Builder->>Builder: build_graph_from_components(components) - Builder->>Builder: validate_graph_completeness(components, graph) - Builder->>Builder: get_leaf_nodes(graph, components) - Builder-->>Caller: (components, leaf_nodes) -``` - -## Relationship to the Broader Pipeline - -- **Upstream**: Relies on `AnalysisService` from the [Analysis Pipeline](../analysis_pipeline/analysis_pipeline.md) for structural file-tree scanning and call-graph extraction, which in turn delegates language-specific parsing to the [Tree-Sitter Analyzers](../../tree-sitter-analyzers/tree-sitter-analyzers.md). -- **Configuration**: Reads repository paths, filter patterns, and output directories from the [Config](../../../config-core.md) object. -- **Data Model**: Produces and manipulates `Node` instances defined in [Dependency Analyzer Models](../../dependency-analyzer-models/dependency-analyzer-models.md). -- **Downstream**: The `(components, leaf_nodes)` tuple returned by `DependencyGraphBuilder.build_dependency_graph()` feeds into the clustering and documentation-generation stages that consume the persisted dependency graph and the filtered leaf set for hierarchical documentation planning. - -For the broader dependency-analysis subsystem this module belongs to, see the [Dependency Analyzer Core](../dependency-analyzer-core.md) overview. diff --git a/docs/reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md b/docs/reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md deleted file mode 100644 index a65c76c6..00000000 --- a/docs/reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md +++ /dev/null @@ -1,194 +0,0 @@ -# Dependency Analyzer Models - -## Introduction - -The Dependency Analyzer Models module defines the core Pydantic data contracts used throughout the dependency analysis pipeline of the backend. It contains no business logic of its own — instead, it establishes the shared vocabulary (data shapes) that every other backend component relies on to represent source code entities, their relationships, the repositories being analyzed, and the final analysis output. - -Because these models are pure data definitions with no external side effects, this module acts as the **foundation layer** of the backend dependency-analysis subsystem. Every producer (parsers, analyzers, graph builders) and every consumer (documentation generator, agent orchestrator, CLI) of analysis data depends on these types either directly or transitively. - -The module is composed of two files: - -- `models/core.py` — the fundamental building blocks: `Node`, `CallRelationship`, `Repository` -- `models/analysis.py` — the higher-level, aggregate result types: `AnalysisResult`, `NodeSelection` - -Given its small size and purely declarative nature, this module is documented as a single cohesive unit rather than split into further sub-modules. - -## Purpose in the Overall System - -The dependency-analysis pipeline works roughly as follows: - -1. Language-specific tree-sitter/AST analyzers parse source files and extract code entities and call edges. -2. The `DependencyParser` and `DependencyGraphBuilder` (part of the Dependency Analyzer Core module) assemble these extracted entities into `Node` and `CallRelationship` instances. -3. The `AnalysisService`, `RepoAnalyzer`, and `CallGraphAnalyzer` (also part of Dependency Analyzer Core) orchestrate the end-to-end analysis of a repository, producing an `AnalysisResult`. -4. Downstream consumers — the Documentation Generator, the Agent Orchestrator, and the CLI — read the `AnalysisResult` (and its nested `Node`/`CallRelationship`/`Repository` data) to generate documentation, drive AI agents, or render reports. - -This module's models are therefore the "wire format" passed between all of those stages. - -```mermaid -flowchart TD - Analyzers["Tree-sitter Analyzers"] -->|"produce"| Node["Node"] - Analyzers -->|"produce"| CallRel["CallRelationship"] - Parser["Dependency Parser / Graph Builder"] -->|"assembles"| Node - Parser -->|"assembles"| CallRel - RepoModel["Repository"] -->|"describes source of"| Node - AnalysisSvc["Analysis Service / Repo Analyzer / Call Graph Analyzer"] -->|"aggregates"| Node - AnalysisSvc -->|"aggregates"| CallRel - AnalysisSvc -->|"aggregates"| RepoModel - AnalysisSvc -->|"produces"| Result["AnalysisResult"] - Selection["NodeSelection"] -->|"filters"| Result - Result -->|"consumed by"| DocGen["Documentation Generator"] - Result -->|"consumed by"| Agent["Agent Orchestrator"] - Result -->|"consumed by"| CLIGen["CLI Documentation Generator"] -``` - -## Core Components - -### `Node` - -`Node` (defined in `models/core.py`) is the canonical representation of a single source-code entity discovered during analysis — a function, class, method, or other addressable code unit. - -Key fields: - -- `id`: Fully-qualified domain name (FQDN) in the format `{namespace}.{original_id}`, used as the unique identifier across the entire system. -- `name`, `display_name`, `class_name`: Human-readable naming information; `get_display_name()` falls back to `name` when `display_name` is not set. -- `component_type`, `node_type`: Classify the kind of entity (e.g., function, class, method). -- `file_path`, `relative_path`: Location of the entity within the analyzed repository. -- `depends_on`: A `Set[str]` of other node IDs this node depends on — the backbone of the dependency graph. -- `source_code`, `start_line`, `end_line`: Raw source snippet and its location for documentation/reference purposes. -- `has_docstring`, `docstring`: Extracted documentation comments. -- `parameters`, `base_classes`: Signature and inheritance metadata (when applicable). -- `component_id`: Optional external identifier correlating a node with a documentation component. -- FQDN metadata: `short_id` (original ID without namespace), `namespace` (e.g., `main`, `deps`, `ui-kit`), and `is_from_deps` (whether the node originates from a dependency rather than the main repository). - -`Node` is intentionally permissive — most fields beyond `id`, `name`, `component_type`, `file_path`, and `relative_path` are optional or default-valued, allowing different language analyzers (Python, Java, C/C++/C#, JavaScript/TypeScript, PHP) to populate only the metadata that is available for their language. - -### `CallRelationship` - -`CallRelationship` models a directed edge between two `Node` instances, representing a call or reference from one code entity to another. - -Fields: - -- `caller`: ID of the node initiating the call. -- `callee`: ID of the node being called. -- `call_line`: Optional line number where the call occurs, useful for navigation and documentation cross-referencing. -- `is_resolved`: Whether the callee could be definitively resolved to a known `Node` (as opposed to remaining an unresolved/external reference). - -Collections of `CallRelationship` objects, combined with `Node.depends_on` sets, form the call graph that the Call Graph Analyzer and Dependency Graph Builder operate on. - -### `Repository` - -`Repository` captures metadata about the source repository under analysis: - -- `url`: Origin location of the repository (e.g., a Git remote URL). -- `name`: Human-readable repository name. -- `clone_path`: Local filesystem path where the repository was cloned/checked out for analysis. -- `analysis_id`: Identifier correlating this repository record with a specific analysis run. - -This model is embedded directly inside `AnalysisResult` to associate analysis output with its source. - -### `AnalysisResult` - -`AnalysisResult` (defined in `models/analysis.py`) is the top-level aggregate produced at the end of a full repository analysis. It bundles together everything downstream consumers need: - -- `repository`: The `Repository` that was analyzed. -- `functions`: A `List[Node]` of all extracted code entities. -- `relationships`: A `List[CallRelationship]` describing the call graph. -- `file_tree`: A `Dict[str, Any]` representing the hierarchical file/directory structure of the repository. -- `summary`: A `Dict[str, Any]` with aggregate statistics or high-level findings. -- `visualization`: An optional `Dict[str, Any]` holding precomputed visualization data (defaults to an empty dict). -- `readme_content`: An optional string containing the repository's README content, when available. - -This model is the primary data structure passed from the analysis pipeline (Analysis Service, Repo Analyzer, Call Graph Analyzer) to consumers such as the Documentation Generator and the Agent Orchestrator. - -### `NodeSelection` - -`NodeSelection` supports partial/filtered exports of an analysis result — for example, when a user wants documentation generated for only a subset of discovered nodes. - -Fields: - -- `selected_nodes`: A `List[str]` of node IDs to include (defaults to an empty list). -- `include_relationships`: Whether relationships between selected nodes should also be included (defaults to `True`). -- `custom_names`: A `Dict[str, str]` mapping node IDs to user-supplied display names, allowing renaming without mutating the underlying `Node` data. - -## Data Model Relationships - -```mermaid -classDiagram - class Node { - +str id - +str name - +str component_type - +str file_path - +str relative_path - +Set~str~ depends_on - +Optional~str~ source_code - +int start_line - +int end_line - +bool has_docstring - +str docstring - +Optional~List~str~~ parameters - +Optional~str~ node_type - +Optional~List~str~~ base_classes - +Optional~str~ class_name - +Optional~str~ display_name - +Optional~str~ component_id - +str short_id - +str namespace - +bool is_from_deps - +get_display_name() str - } - - class CallRelationship { - +str caller - +str callee - +Optional~int~ call_line - +bool is_resolved - } - - class Repository { - +str url - +str name - +str clone_path - +str analysis_id - } - - class AnalysisResult { - +Repository repository - +List~Node~ functions - +List~CallRelationship~ relationships - +Dict file_tree - +Dict summary - +Dict visualization - +Optional~str~ readme_content - } - - class NodeSelection { - +List~str~ selected_nodes - +bool include_relationships - +Dict~str,str~ custom_names - } - - AnalysisResult "1" --> "1" Repository : describes - AnalysisResult "1" --> "*" Node : contains - AnalysisResult "1" --> "*" CallRelationship : contains - CallRelationship "*" --> "1" Node : caller/callee reference (by id) - NodeSelection "1" --> "*" Node : references (by id) -``` - -## Usage Across the System - -These models form the shared contract consumed by several other backend modules: - -- The tree-sitter language analyzers and the AST/graph construction components (`DependencyParser`, `DependencyGraphBuilder`) construct `Node` and `CallRelationship` instances as they walk source files. -- The analysis pipeline components (`AnalysisService`, `RepoAnalyzer`, `CallGraphAnalyzer`) assemble these into a single `Repository`-scoped `AnalysisResult`. -- The Documentation Generator consumes `AnalysisResult` to produce human-readable documentation, optionally filtering via `NodeSelection` for partial exports. -- The Agent Orchestrator and its tooling read `Node` and `CallRelationship` data to answer questions about code structure and to drive AI-assisted documentation generation. -- The CLI's documentation generation adapter ultimately surfaces `AnalysisResult` data to end users. - -Because these are plain Pydantic `BaseModel` classes, they also provide built-in serialization/validation, making it straightforward to persist analysis results as JSON, pass them between processes, or validate data received from external sources. - -## Design Notes - -- **FQDN-based identity**: The `id` field on `Node` always follows the `{namespace}.{original_id}` convention, ensuring uniqueness even when analyzing a main repository alongside its dependencies. The `namespace` and `is_from_deps` fields make it possible to distinguish first-party code from vendored/dependency code without losing the original identifier (`short_id`). -- **Permissive defaults**: Nearly every field beyond the minimal identity fields on `Node` has a sensible default (empty set, `None`, empty string), which allows the model to accommodate the varying levels of metadata different language analyzers can extract. -- **Separation of raw graph data from aggregate results**: `Node`/`CallRelationship`/`Repository` represent the atomic units of the dependency graph, while `AnalysisResult` represents the fully-assembled output of a completed analysis run. `NodeSelection` sits alongside `AnalysisResult` as a lightweight filter/view specification rather than a graph primitive. diff --git a/docs/reference/architecture/backend-core/documentation-generator/documentation-generator.md b/docs/reference/architecture/backend-core/documentation-generator/documentation-generator.md deleted file mode 100644 index 488c05a1..00000000 --- a/docs/reference/architecture/backend-core/documentation-generator/documentation-generator.md +++ /dev/null @@ -1,228 +0,0 @@ -# Documentation Generator - -The Documentation Generator module contains the `DocumentationGenerator` class — the top-level orchestrator that drives the entire end-to-end documentation generation pipeline for a target repository. It coordinates dependency graph construction, module clustering, hierarchical (dynamic-programming style) documentation generation via AI agents, and final metadata assembly. - -This module is the primary entry point invoked by both the [CLI Core](../../cli-core.md) (`CLIDocumentationGenerator`) and the [Frontend Core](../../frontend-core.md) (`BackgroundWorker`) layers whenever a documentation job needs to be executed against a cloned or provided repository. - -## Purpose and Responsibilities - -`DocumentationGenerator` ties together several backend subsystems to transform raw source code into a hierarchical set of Markdown documentation files: - -- **Dependency graph construction**: Delegates to `DependencyGraphBuilder` (see [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md)) to parse the repository and produce a component graph plus a filtered list of "leaf nodes" (classes, interfaces, structs, or functions worth documenting). -- **Module clustering**: Groups leaf-level components into a hierarchical module tree using `cluster_modules`, with a **synthetic module fallback** that prevents context-overflow failures when clustering yields no groups. -- **Documentation generation ordering**: Computes a topological "leaf-modules-first" processing order so that child module documentation is always available before its parent module summary is generated. -- **Agent-driven content generation**: Delegates the actual LLM-backed documentation writing to `AgentOrchestrator` (see [Backend Core](../../backend-core.md)) for leaf modules, and to its own `generate_parent_module_docs` routine (using `REPO_OVERVIEW_PROMPT` / `MODULE_OVERVIEW_PROMPT`) for parent/overview modules. -- **Hierarchical file layout**: Computes nested working directories (`module_path` → `docs_dir/A/B/C/C.md`) so that every module's documentation lives in its own directory alongside its children. -- **Metadata generation**: After generation completes, writes a `metadata.json` describing generation statistics and the list of all generated Markdown files. - -## Position in the System - -```mermaid -flowchart TD - CLI["CLI Core: CLIDocumentationGenerator"] --> DG["Documentation Generator: DocumentationGenerator"] - Web["Frontend Core: BackgroundWorker"] --> DG - DG --> GraphBuilder["Dependency Analyzer Core: DependencyGraphBuilder"] - DG --> Cluster["cluster_modules"] - DG --> Orchestrator["Backend Core: AgentOrchestrator"] - DG --> LLM["LLM Services: call_llm / CountingFallbackModel"] - GraphBuilder --> Analyzers["Tree-sitter Analyzers"] - GraphBuilder --> Models["Dependency Analyzer Models"] - Orchestrator --> AgentTools["Agent Tools Core"] - DG --> ConfigCore["Config Core: Config"] -``` - -- **Upstream callers**: [CLI Core](../../cli-core.md) and [Frontend Core](../../frontend-core.md) both construct a `Config` (see [Config Core](../../../config-core.md)) and instantiate `DocumentationGenerator(config, commit_id)` before calling `run()`. -- **Downstream collaborators**: - - `DependencyGraphBuilder` from [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md) parses the repository via `DependencyParser` and the [Tree-sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md). - - `AgentOrchestrator` (in [Backend Core](../../backend-core.md)) creates and runs `pydantic-ai` agents backed by `CountingFallbackModel` from [LLM Services](../llm-services/llm-services.md), using tools defined in [Agent Tools Core](../agent-tools-core/agent-tools-core.md). - - `Node`, `AnalysisResult` and related types from [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) represent the parsed components passed throughout the pipeline. - -## Core Component - -### `DocumentationGenerator` - -```python -class DocumentationGenerator: - def __init__(self, config: Config, commit_id: str = None): - self.config = config - self.commit_id = commit_id - self.graph_builder = DependencyGraphBuilder(config) - self.agent_orchestrator = AgentOrchestrator(config) -``` - -On construction, `DocumentationGenerator` immediately wires up its two main collaborators: - -| Attribute | Type | Role | -|---|---|---| -| `config` | `Config` | Repository path, model configuration, output directories, agent instructions | -| `commit_id` | `str \| None` | Optional git commit SHA recorded in generation metadata | -| `graph_builder` | `DependencyGraphBuilder` | Parses the repository and produces components + leaf nodes | -| `agent_orchestrator` | `AgentOrchestrator` | Runs LLM agents to write documentation for leaf modules | - -#### Public / Key Methods - -| Method | Purpose | -|---|---| -| `run()` | Entry point — runs the complete pipeline: graph build → clustering (with synthetic fallback) → module documentation generation → metadata creation | -| `generate_module_documentation(components, leaf_nodes)` | Iterates the module tree in leaf-first order, invoking either the agent orchestrator (leaf modules) or `generate_parent_module_docs` (parent modules), and finally the repository overview | -| `generate_parent_module_docs(module_path, working_dir)` | Builds a summarized child-doc structure and calls the LLM (`REPO_OVERVIEW_PROMPT` / `MODULE_OVERVIEW_PROMPT`) to synthesize a parent/overview document | -| `get_processing_order(module_tree, parent_path)` | Computes a topological (post-order/leaf-first) traversal of the module tree | -| `is_leaf_module(module_info)` | Determines whether a module has no children (i.e., should be documented directly from components) | -| `build_overview_structure(module_tree, module_path, working_dir)` | Produces a JSON-serializable subtree with one level of children's already-generated docs embedded, marking the current target module | -| `_get_nested_working_dir(base_dir, module_path)` | Computes the hierarchical output directory for a given module path (e.g. `docs_dir/Backend/Auth/JWT/`) | -| `create_documentation_metadata(working_dir, components, num_leaf_nodes)` | Walks the output directory for generated `.md` files and writes `metadata.json` | - -## End-to-End Generation Flow - -```mermaid -sequenceDiagram - participant Caller as "CLI/Web Caller" - participant DG as "DocumentationGenerator" - participant GB as "DependencyGraphBuilder" - participant Cluster as "cluster_modules" - participant AO as "AgentOrchestrator" - participant LLM as "call_llm" - participant FS as "file_manager" - - Caller->>DG: run() - DG->>GB: build_dependency_graph() - GB-->>DG: components, leaf_nodes - alt "module tree cache exists" - DG->>FS: load_json(first_module_tree.json) - else "no cache" - DG->>Cluster: cluster_modules(leaf_nodes, components, config) - Cluster-->>DG: module_tree - DG->>FS: save_json(first_module_tree.json) - end - alt "module_tree empty but leaf_nodes exist" - DG->>DG: build synthetic modules by top-level directory - DG->>FS: save_json(synthetic module_tree) - end - DG->>FS: save_json(module_tree.json) - DG->>DG: generate_module_documentation(components, leaf_nodes) - loop "for each module in leaf-first order" - alt "is_leaf_module" - DG->>AO: process_module(name, components, ids, path, dir) - AO->>LLM: agent.run(user_prompt) - LLM-->>AO: markdown docs - AO->>FS: save module_name.md - else "parent module" - DG->>DG: build_overview_structure(...) - DG->>LLM: call_llm(MODULE_OVERVIEW_PROMPT) - LLM-->>DG: "..." - DG->>FS: save parent module_name.md - end - end - DG->>DG: generate_parent_module_docs([], working_dir) - DG->>LLM: call_llm(REPO_OVERVIEW_PROMPT) - DG->>FS: save overview.md - DG->>DG: create_documentation_metadata(working_dir, components, len(leaf_nodes)) - DG-->>Caller: "documentation complete" -``` - -### 1. Dependency Graph Construction - -`run()` first delegates to `self.graph_builder.build_dependency_graph()`, returning: -- `components`: a dict mapping component IDs to parsed `Node` objects (see [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md)) -- `leaf_nodes`: a filtered list of component IDs that are candidates for direct documentation (classes/interfaces/structs, or functions for C-style codebases) - -### 2. Module Clustering (with Synthetic Fallback) - -If a cached `first_module_tree.json` exists, it is reused. Otherwise `cluster_modules(leaf_nodes, components, config)` groups leaf nodes into a semantic module hierarchy. - -A **synthetic module patch** guards against an important failure mode: if clustering returns an empty tree despite having leaf nodes (which would otherwise trigger an expensive "whole repository in one pass" fallback that can exceed LLM context limits), `DocumentationGenerator` instead groups leaf nodes by their top-level source directory and constructs flat synthetic modules directly. This synthetic tree is persisted back to the cache file so subsequent runs skip re-clustering. - -### 3. Leaf-First Processing Order - -`get_processing_order` performs a recursive collection over the module tree: for any module with children, it recurses into the children **first** and appends the parent path afterward. This guarantees a true dynamic-programming order — every child module's Markdown file already exists on disk by the time its parent module is processed, since `build_overview_structure` reads children's generated docs directly from `working_dir/child_name/child_name.md`. - -```mermaid -flowchart LR - subgraph Tree["Module Tree Example"] - Root["Backend"] --> Auth["Authentication"] - Root --> API["API Layer"] - Auth --> JWT["JWT"] - Auth --> OAuth["OAuth"] - end - subgraph Order["Processing Order"] - O1["1: JWT (leaf)"] --> O2["2: OAuth (leaf)"] - O2 --> O3["3: Authentication (parent)"] - O3 --> O4["4: API Layer (leaf/parent)"] - O4 --> O5["5: Backend (parent)"] - end -``` - -### 4. Per-Module Documentation Generation - -For each module in the processing order, `generate_module_documentation`: -1. Resolves `module_info` by walking the module tree along `module_path`. -2. Skips modules already processed (idempotency across retries). -3. Computes the nested working directory via `_get_nested_working_dir` (e.g. `docs_dir/Backend/Authentication/JWT/`) and ensures it exists. -4. **Leaf modules** — delegates to `self.agent_orchestrator.process_module(...)`, which builds a `pydantic-ai` agent (simple or "complex" with sub-module tools) and runs it against the module's component list. See [Backend Core](../../backend-core.md) for orchestrator internals. -5. **Parent modules** — delegates to `self.generate_parent_module_docs(...)`, which synthesizes an overview from already-generated child documentation. -6. Errors for any single module are logged with a full traceback but do **not** halt the overall run — processing continues with the next module (graceful degradation). - -After all modules are processed, a final call to `generate_parent_module_docs([], working_dir)` produces the top-level repository `overview.md`. - -### 5. Parent / Overview Document Synthesis - -`generate_parent_module_docs`: -1. Loads the canonical `module_tree.json` from the **base** docs directory (not the nested `working_dir`, since `module_tree.json` is always written at the top level). -2. Returns early if `overview.md` or the module's own `.md` already exists (idempotent caching). -3. Calls `build_overview_structure` to construct a subtree containing the target module marked with `is_target_for_overview_generation` and each direct child's already-generated Markdown embedded under a `docs` key. -4. Serializes this structure to JSON and formats it into `MODULE_OVERVIEW_PROMPT` (for non-root modules) or `REPO_OVERVIEW_PROMPT` (for the repository root, when `module_path` is empty). -5. Invokes `call_llm(prompt, self.config)` (see [LLM Services](../llm-services/llm-services.md)), extracts the content between `` / `` tags (falling back to the raw response if tags are missing), and writes it to disk via `file_manager.save_text`. - -### 6. Small-Repository Fast Path - -If clustering (including the synthetic fallback) still yields **zero** modules, `run()` falls back to treating the entire repository as a single module: it calls `agent_orchestrator.process_module` directly with all `leaf_nodes`, then renames the resulting `.md` to `overview.md`. - -### 7. Metadata Generation - -Once all documentation is written, `create_documentation_metadata` walks the output directory tree, records every generated `.md` file (relative path), and writes a `metadata.json` capturing: -- Generation timestamp, model used, generator version, repo path, and commit ID -- Component/leaf-node/max-depth statistics -- The full list of generated files - -## Directory Layout Convention - -`DocumentationGenerator` enforces a strict hierarchical output convention via `_get_nested_working_dir`: every module — leaf or parent, at any depth — gets its own subdirectory named after itself, containing a same-named Markdown file: - -```text -docs_dir/ -├── overview.md # repository-level overview -├── module_tree.json # final clustered/processed module tree -├── first_module_tree.json # initial clustering result (cache) -├── metadata.json # generation metadata -├── Backend/ -│ ├── Backend.md # parent module overview -│ ├── Authentication/ -│ │ ├── Authentication.md -│ │ ├── JWT/ -│ │ │ └── JWT.md # leaf module docs -│ │ └── OAuth/ -│ │ └── OAuth.md -│ └── API/ -│ └── API.md -``` - -This matches the same hierarchical linking convention used across all generated module documentation in this system. - -## Error Handling and Resilience - -- **Per-module isolation**: Exceptions raised while processing an individual module in `generate_module_documentation` are caught, logged with a traceback, and the loop continues — a single failing module does not abort the entire documentation run. -- **Idempotent caching**: Both `process_module` (in `AgentOrchestrator`) and `generate_parent_module_docs` check for existing output files before invoking the LLM, allowing interrupted runs to be resumed cheaply. -- **Context-overflow prevention**: The synthetic module fallback in `run()` avoids the case where an empty clustering result would force the entire repository through a single oversized LLM call. -- **Tag-tolerant parsing**: `generate_parent_module_docs` gracefully handles LLM responses that omit the expected `` wrapper tags by falling back to the full raw response instead of failing. - -## Related Modules - -- [Backend Core](../../backend-core.md) — parent module; hosts `AgentOrchestrator`, agent tools, and the dependency analysis subsystem this module depends on -- [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md) — provides `DependencyGraphBuilder` used to parse the repository -- [Tree-sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) — language-specific parsers invoked during dependency graph construction -- [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) — `Node` and related data models flowing through this pipeline -- [Agent Tools Core](../agent-tools-core/agent-tools-core.md) — tools (`EditTool`, `Filemap`, `CodeWikiDeps`) used by agents during leaf-module generation -- [LLM Services](../llm-services/llm-services.md) — `CountingFallbackModel` and `call_llm`, used both directly (for parent/overview synthesis) and indirectly (via `AgentOrchestrator`) for leaf modules -- [Config Core](../../../config-core.md) — the `Config` object supplied to `DocumentationGenerator` on construction -- [CLI Core](../../../cli-core.md) — command-line entry point that constructs `Config` and invokes `DocumentationGenerator.run()` -- [Frontend Core](../../../frontend-core.md) — web application entry point (`BackgroundWorker`) that runs the same generation pipeline for submitted repositories diff --git a/docs/reference/architecture/backend-core/llm-services/llm-services.md b/docs/reference/architecture/backend-core/llm-services/llm-services.md deleted file mode 100644 index 965ca114..00000000 --- a/docs/reference/architecture/backend-core/llm-services/llm-services.md +++ /dev/null @@ -1,224 +0,0 @@ -# Llm Services - -The Llm Services module is the central factory and runtime layer for all Large Language Model (LLM) interactions in CodeWiki's backend. It builds configured `pydantic-ai` model instances (main, fallback, and cluster models) from a [Config](../../config-core.md) object, wraps them with automatic failover and request-counting behavior, and exposes a low-level synchronous `call_llm()` helper for direct OpenAI-compatible API calls. Every LLM-driven component in the backend — the [Documentation Generator](../documentation-generator/documentation-generator.md) and the agent orchestration layer that powers it — depends on this module to obtain ready-to-use model clients without needing to know provider-specific details (base URLs, API keys, temperature support, or token-limit parameter naming). - -## Purpose and Scope - -CodeWiki supports multiple LLM providers (OpenAI, Anthropic, and OpenAI-compatible endpoints) across three distinct operational stages: - -- **Main/generation** — the primary model used to write module documentation. -- **Fallback** — a secondary model automatically invoked if the main model fails. -- **Cluster** — a (usually cheaper/faster) model used for the module-clustering stage of documentation generation. - -Because each stage can use a different provider with different base URLs, API keys, temperature support, and token-limit parameter names (`max_tokens` vs. `max_completion_tokens` for reasoning models like o3/o3-mini), the Llm Services module centralizes this provider-resolution logic in one place so the rest of the backend can remain provider-agnostic. - -## Core Responsibilities - -1. **Model factory functions** — build `pydantic-ai` `OpenAIModel` instances per stage (`create_main_model`, `create_fallback_model`, `create_cluster_model`), each validating that the required `base_url` and `api_key` are present in [Config](../../config-core.md) and raising descriptive `ValueError`s otherwise. -2. **Fallback chain assembly** — `create_fallback_models()` combines the main and fallback models into a single `CountingFallbackModel`, the core component of this module, which `pydantic-ai`'s `Agent` uses transparently: if the main model errors, the fallback model is tried automatically. -3. **Request counting/telemetry** — a module-level counter (`_request_counter`) tracks how many LLM requests have been made for the current module being documented, logging progress every 100 requests via `increment_request_counter()` / `reset_request_counter()`. -4. **Direct/synchronous LLM calls** — `call_llm()` provides a simpler, non-agentic path for callers (such as the parent-module overview generation flow in [Documentation Generator](../documentation-generator/documentation-generator.md)) that just need a single prompt/response round-trip using the raw `openai` Python SDK client built by `create_openai_client()`. -5. **Token-limit and temperature normalization** — `get_max_output_tokens()` and `get_model_max_token_field()` resolve environment-variable overrides (`MAX_OUTPUT_TOKENS`, `CODEWIKI_*_MAX_TOKEN_FIELD`) so reasoning models that require `max_completion_tokens` instead of `max_tokens` work without code changes. - -## Core Component - -### CountingFallbackModel - -`CountingFallbackModel` extends `pydantic-ai`'s `FallbackModel` and overrides `request()` to call `increment_request_counter()` before delegating to the parent implementation. This gives CodeWiki visibility into LLM usage volume per module without modifying `pydantic-ai` internals or scattering counting logic throughout the agent code. It is the object returned by `create_fallback_models()` and passed directly to `pydantic-ai.Agent` by the [Agent Orchestrator](../documentation-generator/documentation-generator.md) (part of `AgentOrchestrator.create_agent`), meaning **every** agent-driven LLM call in CodeWiki flows through this wrapper. - -```python -class CountingFallbackModel(FallbackModel): - """FallbackModel wrapper that counts and logs requests every 100 calls.""" - - async def request(self, *args, **kwargs): - """Wrap request to count calls.""" - increment_request_counter() - return await super().request(*args, **kwargs) -``` - -## Architecture - -```mermaid -flowchart TD - Config["Config"] --> CreateMain["create_main_model()"] - Config --> CreateFallback["create_fallback_model()"] - Config --> CreateCluster["create_cluster_model()"] - Config --> CreateClient["create_openai_client()"] - - CreateMain --> MainModel["OpenAIModel (main)"] - CreateFallback --> FallbackModel_["OpenAIModel (fallback)"] - - MainModel --> CreateChain["create_fallback_models()"] - FallbackModel_ --> CreateChain - - CreateChain --> CountingModel["CountingFallbackModel"] - CountingModel --> IncCounter["increment_request_counter()"] - IncCounter --> Counter[("_request_counter")] - - CountingModel --> Agent["pydantic-ai Agent"] - - CreateCluster --> ClusterModel["OpenAIModel (cluster)"] - ClusterModel --> ClusterAgent["Clustering LLM calls"] - - CreateClient --> OpenAIClient["openai.OpenAI client"] - OpenAIClient --> CallLLM["call_llm()"] - CallLLM --> Response["LLM response text"] -``` - -## How the Module Fits Into the Backend - -```mermaid -flowchart LR - subgraph LlmServices["Llm Services"] - Factory["Model factories
(main / fallback / cluster)"] - Counting["CountingFallbackModel"] - DirectCall["call_llm()"] - end - - subgraph DocGen["Documentation Generator"] - DG["DocumentationGenerator"] - Orchestrator["AgentOrchestrator"] - end - - ConfigCore["Config"] --> Factory - Factory --> Counting - Orchestrator -->|"create_fallback_models(config)"| Factory - Orchestrator -->|"used by pydantic-ai Agent"| Counting - DG -->|"call_llm(prompt, config)"| DirectCall - - DirectCall --> OpenAISDK["openai.OpenAI client"] - Counting --> PydanticAI["pydantic-ai FallbackModel / OpenAIProvider"] -``` - -*[Documentation Generator](../documentation-generator/documentation-generator.md) owns the `AgentOrchestrator`, which calls `create_fallback_models(config)` once per module to build the agent's LLM backend, and calls `call_llm()` directly when generating parent-module overview documentation (bypassing the agent framework for simpler prompt/response interactions).* - -## Model Creation Flow - -Each `create_*_model()` function follows the same validation-and-build pattern, differing only in which per-provider config fields it reads (`main_*`, `fallback_*`, or `cluster_*`): - -```mermaid -sequenceDiagram - participant Caller - participant Factory as "create_main_model() / create_fallback_model() / create_cluster_model()" - participant Config as "Config" - participant Provider as "OpenAIProvider" - participant Model as "OpenAIModel" - - Caller->>Factory: create_X_model(config) - Factory->>Config: getattr(config, "X_max_tokens", None) - Factory->>Config: getattr(config, "X_temperature", 0.0) - Factory->>Config: getattr(config, "X_temperature_supported", True) - Factory->>Config: getattr(config, "X_base_url", None) - alt base_url missing - Factory-->>Caller: raise ValueError("X_base_url is required...") - end - Factory->>Config: getattr(config, "X_api_key", None) - alt api_key missing - Factory-->>Caller: raise ValueError("X_api_key is required...") - end - Factory->>Provider: OpenAIProvider(base_url, api_key) - Factory->>Model: OpenAIModel(model_name, provider, settings) - Model-->>Caller: configured OpenAIModel -``` - -## Fallback Chain Construction - -```mermaid -sequenceDiagram - participant Orchestrator as "AgentOrchestrator" - participant LlmServices as "llm_services" - participant MainModel as "OpenAIModel (main)" - participant FallbackModelObj as "OpenAIModel (fallback)" - participant CFM as "CountingFallbackModel" - - Orchestrator->>LlmServices: create_fallback_models(config) - LlmServices->>LlmServices: create_main_model(config) - LlmServices-->>MainModel: main - LlmServices->>LlmServices: create_fallback_model(config) - LlmServices-->>FallbackModelObj: fallback - LlmServices->>CFM: CountingFallbackModel(main, fallback) - CFM-->>Orchestrator: fallback_models - - Note over Orchestrator,CFM: Passed as the model backend
for pydantic-ai Agent -``` - -At runtime, when `pydantic-ai`'s `Agent.run()` invokes `CountingFallbackModel.request()`: - -1. `increment_request_counter()` runs first, bumping the module-scoped counter and logging a progress line every 100 requests. -2. The call is delegated to `FallbackModel.request()` (the parent class), which attempts the **main** model first and automatically retries with the **fallback** model if the main model raises an error. - -## Direct LLM Invocation (`call_llm`) - -For flows that do not need the full `pydantic-ai` agent/tool-calling machinery — such as generating a parent-module overview from already-generated child documentation — the module exposes a synchronous `call_llm()` helper built directly on the `openai` SDK. - -```mermaid -sequenceDiagram - participant Caller - participant CallLLM as "call_llm()" - participant ClientFactory as "create_openai_client()" - participant Config as "Config" - participant OpenAISDK as "openai.OpenAI" - - Caller->>CallLLM: call_llm(prompt, config, model, temperature) - CallLLM->>CallLLM: resolve stage (main / cluster / fallback) - CallLLM->>ClientFactory: create_openai_client(config, model) - ClientFactory->>Config: resolve base_url / api_key / api_version for stage - alt base_url or api_key missing - ClientFactory-->>CallLLM: raise ValueError(...) - end - ClientFactory->>OpenAISDK: OpenAI(base_url, api_key, default_headers) - OpenAISDK-->>CallLLM: client - CallLLM->>CallLLM: get_model_max_token_field(stage) - CallLLM->>OpenAISDK: client.chat.completions.create(**kwargs) - alt success - OpenAISDK-->>CallLLM: response - CallLLM-->>Caller: response_content - else OpenAIError - CallLLM-->>Caller: raise RuntimeError(context-wrapped) - end -``` - -Key behaviors of `call_llm()`: - -- **Stage detection** — determines whether `model` matches `config.cluster_model`, `config.fallback_model`, or defaults to main/generation, so it can select the correct `base_url`, `api_key`, `api_version`, `max_tokens` value, temperature support flag, and token-field name for that specific provider. -- **Dynamic token-field naming** — uses `get_model_max_token_field(stage)` to decide whether to send `max_tokens` or `max_completion_tokens` in the request payload, accommodating reasoning models (o3, o3-mini) that reject the standard `max_tokens` parameter. -- **Conditional temperature** — only includes `temperature` in the request if the resolved `*_temperature_supported` config flag is `True`, since some reasoning models reject custom temperature values entirely. -- **Rich error context** — wraps both `OpenAIError` and generic exceptions in a `RuntimeError` that includes the stage, model, base URL, temperature, and max-token settings to aid debugging misconfigured providers. -- **Structured logging** — emits detailed, tree-formatted log lines (stage, model, base URL, prompt length/preview, temperature, token settings) both before the request and after a successful response, easing observability during large documentation runs. - -## Environment-Driven Configuration Helpers - -Two helper functions decouple token-limit behavior from hardcoded values, reading overrides from environment variables at call time: - -| Function | Environment Variable(s) | Purpose | -|---|---|---| -| `get_max_output_tokens()` | `MAX_OUTPUT_TOKENS` | Returns the default max output tokens (16384) unless overridden; used as a fallback when a stage-specific `*_max_tokens` field is not set on [Config](../../config-core.md). | -| `get_model_max_token_field(stage)` | `CODEWIKI_CLUSTER_MAX_TOKEN_FIELD`, `CODEWIKI_GENERATION_MAX_TOKEN_FIELD`, `CODEWIKI_FALLBACK_MAX_TOKEN_FIELD` | Returns `max_tokens` (default) or `max_completion_tokens` for the given stage, letting operators switch reasoning-model support without code changes. | - -## Relationship to Configuration - -All factory functions and `call_llm()` read their provider settings exclusively from the [Config](../../config-core.md) object — specifically its per-provider fields (`main_base_url`, `main_api_key`, `main_temperature`, `cluster_base_url`, `cluster_api_key`, `fallback_base_url`, `fallback_api_key`, etc.). This module performs no environment-variable parsing of its own for credentials; all credential/URL resolution responsibility lives in `Config.from_args()`, `Config.from_cli()`, and `Config.from_web_job()`. Llm Services only reads the resulting per-provider attributes via `getattr()`, defaulting gracefully where sensible (e.g., temperature defaults to `0.0`, token limits default to `get_max_output_tokens()`). - -## Error Handling Philosophy - -Every factory function fails fast with a descriptive `ValueError` when required configuration (`base_url`, `api_key`) is missing, naming the exact CLI flag or config-file field the caller should set: - -```python -raise ValueError( - "main_base_url is required in configuration for main/generation model.\n" - f"Model: {config.main_model}\n" - "Please set via CLI: --main-base-url \n" - "Or in config file: main_base_url = ''" -) -``` - -This "actionable error message" pattern is applied consistently across `create_main_model`, `create_fallback_model`, `create_cluster_model`, and `create_openai_client`, minimizing debugging time when a user misconfigures a provider for one of the three stages. - -## Summary - -The Llm Services module is a thin but critical abstraction layer that: - -- Converts a single [Config](../../config-core.md) object into fully-configured, provider-agnostic `pydantic-ai` models for three distinct LLM stages (main, fallback, cluster). -- Provides `CountingFallbackModel` as the automatic-failover, telemetry-instrumented model backend used by every `pydantic-ai` `Agent` created by the [Documentation Generator](../documentation-generator/documentation-generator.md)'s agent orchestration layer. -- Offers a simpler synchronous `call_llm()` path for non-agentic prompt/response use cases, with the same provider-resolution and error-handling guarantees. -- Normalizes cross-provider quirks (reasoning-model token-field naming, optional temperature support) so callers never need provider-specific branching logic. diff --git a/docs/reference/architecture/backend-core/logging-config/logging-config.md b/docs/reference/architecture/backend-core/logging-config/logging-config.md deleted file mode 100644 index d6a398f1..00000000 --- a/docs/reference/architecture/backend-core/logging-config/logging-config.md +++ /dev/null @@ -1,181 +0,0 @@ -# Logging Config - -The Logging Config module provides a colorized console logging formatter and helper functions used across the CodeWiki backend to produce readable, severity-differentiated log output. It centers on the `ColoredFormatter` class, a custom subclass of Python's standard `logging.Formatter` that decorates log records with ANSI colors based on log level, and two convenience setup functions (`setup_logging` and `setup_module_logging`) that wire the formatter into console handlers for the root logger or a specific named logger. - -This module is a small, focused utility that other backend components depend on for consistent, human-friendly terminal logging during dependency analysis, documentation generation, and related long-running backend operations. - -## Purpose and Scope - -Backend processes such as repository analysis, call graph construction, and documentation generation emit a large volume of log messages while processing potentially large codebases. Plain, uncolored log output makes it difficult to visually scan for warnings and errors in a busy terminal. The Logging Config module solves this by: - -- Coloring log messages according to severity (DEBUG, INFO, WARNING, ERROR, CRITICAL) -- Coloring timestamps and other structural elements distinctly from the message body -- Providing simple, one-call setup functions so any part of the backend can enable colored logging without repeating formatter/handler boilerplate -- Ensuring cross-platform compatibility (including Windows terminals) via the `colorama` library - -## Core Component - -### ColoredFormatter - -`ColoredFormatter` extends `logging.Formatter` and overrides the `format(record)` method to inject ANSI color codes into the rendered log line. - -**Color scheme:** - -| Log Level / Element | Color | -|---|---| -| DEBUG | Blue | -| INFO | Cyan | -| WARNING | Yellow | -| ERROR | Red | -| CRITICAL | Red + Bright | -| Timestamp | Blue | -| Reset | Style reset (no color) | - -Two internal class-level dictionaries drive this behavior: - -- `COLORS`: maps `record.levelname` (e.g. `"INFO"`, `"ERROR"`) to a `colorama.Fore` color code -- `COMPONENT_COLORS`: maps structural elements (`timestamp`, `module`, `reset`) to their respective colors - -The `format()` method builds the final log line by: -1. Looking up the color for the record's level (falling back to no color if the level is unrecognized) -2. Formatting the timestamp (`HH:MM:SS`) and wrapping it in the timestamp color -3. Formatting the level name (left-padded to 8 characters) in the level's color -4. Formatting the message text in the same color as the level, for visual consistency -5. Concatenating timestamp, level, and message into a single colored line -6. Appending formatted exception traceback text (uncolored) if the record carries exception info - -```mermaid -flowchart TD - A["logging.LogRecord"] --> B["ColoredFormatter.format(record)"] - B --> C["Look up level color in COLORS"] - B --> D["Format timestamp HH:MM:SS"] - D --> E["Wrap timestamp in blue"] - C --> F["Wrap levelname in level color"] - C --> G["Wrap message in level color"] - E --> H["Concatenate: timestamp + level + message"] - F --> H - G --> H - H --> I{"record.exc_info present?"} - I -->|"Yes"| J["Append formatException() output"] - I -->|"No"| K["Return colored log line"] - J --> K -``` - -## Setup Functions - -Alongside `ColoredFormatter`, the module exposes two helper functions that configure logging handlers using the formatter. - -### setup_logging(level=logging.INFO) - -Configures the **root logger** for the entire application: -1. Creates a `logging.StreamHandler` writing to `sys.stdout` -2. Attaches a `ColoredFormatter` instance to the handler -3. Clears any existing handlers on the root logger (to avoid duplicate output when called more than once) -4. Sets the root logger's level and attaches the new handler - -This is intended to be called once, early in a process's lifecycle (e.g., at the start of a CLI or backend service run), to enable colored output application-wide. - -### setup_module_logging(module_name, level=logging.INFO) - -Configures a **named logger** for a specific module rather than the root logger: -1. Retrieves (or creates) a logger via `logging.getLogger(module_name)` -2. Creates a `StreamHandler` to `stdout` with a `ColoredFormatter` -3. Clears existing handlers on that logger -4. Sets `logger.propagate = False` to prevent messages from bubbling up to the root logger (which would otherwise cause duplicate log lines) -5. Returns the configured logger for direct use - -This allows individual backend components — for example an analyzer or the analysis service — to have isolated, independently configured colored logging without interfering with (or being interfered by) the root logger's configuration. - -```mermaid -sequenceDiagram - participant Caller as "Backend Component" - participant Setup as "setup_module_logging()" - participant Logger as "logging.Logger" - participant Handler as "StreamHandler(stdout)" - participant Fmt as "ColoredFormatter" - - Caller->>Setup: setup_module_logging("my_module", level) - Setup->>Logger: logging.getLogger("my_module") - Setup->>Handler: create StreamHandler(sys.stdout) - Setup->>Fmt: create ColoredFormatter() - Setup->>Handler: setFormatter(Fmt) - Setup->>Logger: handlers.clear() - Setup->>Logger: addHandler(Handler) - Setup->>Logger: propagate = False - Setup-->>Caller: return configured logger - Caller->>Logger: logger.info("message") - Logger->>Handler: emit(record) - Handler->>Fmt: format(record) - Fmt-->>Handler: colored log line - Handler-->>Caller: printed to stdout -``` - -## Class Structure - -```mermaid -classDiagram - class Formatter { - <> - +format(record) - +formatTime(record, datefmt) - +formatException(exc_info) - } - class ColoredFormatter { - +COLORS : dict - +COMPONENT_COLORS : dict - +format(record) str - } - Formatter <|-- ColoredFormatter -``` - -## Integration with the Backend - -The Logging Config module is a leaf utility within the broader backend codebase. It is imported wherever colored console output is desired, most notably by components that perform long-running, verbose operations such as dependency graph analysis and repository scanning. Because it only depends on the Python standard library `logging` module and `colorama`, it can be adopted independently by any backend component without introducing coupling to other subsystems. - -Within the backend hierarchy, this module sits alongside sibling utility and domain modules such as [Dependency Analyzer Core](dependency-analyzer-core/dependency-analyzer-core.md), [Tree-Sitter Analyzers](tree-sitter-analyzers/tree-sitter-analyzers.md), [Dependency Analyzer Models](dependency-analyzer-models/dependency-analyzer-models.md), [Documentation Generator](documentation-generator/documentation-generator.md), and [LLM Services](llm-services/llm-services.md), all of which are children of the top-level backend module. - -Note that the CLI package maintains its own separate logging utility (`CLILogger`) for command-line output; the Logging Config module described here is specific to the backend's internal logging needs and is not shared code with the CLI's logging utilities. - -```mermaid -graph TD - Backend["Backend Core"] --> LoggingConfig["Logging Config"] - Backend --> DepAnalyzer["Dependency Analyzer Core"] - Backend --> TreeSitter["Tree-Sitter Analyzers"] - Backend --> DepModels["Dependency Analyzer Models"] - Backend --> DocGen["Documentation Generator"] - Backend --> LLMServices["LLM Services"] - Backend --> AgentTools["Agent Tools Core"] - - DepAnalyzer -.->|"colored console output"| LoggingConfig - TreeSitter -.->|"colored console output"| LoggingConfig - DocGen -.->|"colored console output"| LoggingConfig -``` - -## Usage Pattern - -Typical usage within a backend component follows one of two patterns: - -**Application-wide setup** (once, at process start): - -```python -from codewiki.src.be.dependency_analyzer.utils.logging_config import setup_logging -import logging - -setup_logging(level=logging.INFO) -``` - -**Per-module isolated setup:** - -```python -from codewiki.src.be.dependency_analyzer.utils.logging_config import setup_module_logging -import logging - -logger = setup_module_logging(__name__, level=logging.DEBUG) -logger.debug("Starting analysis...") -``` - -Both patterns rely on the same underlying `ColoredFormatter` to ensure consistent color-coded output regardless of which setup function is used. - -## Relationship to the Parent Module - -This module is part of the backend codebase and is documented as a child of [Backend Core](../backend-core.md). It has no further child modules of its own, as its single component (`ColoredFormatter`) and the accompanying setup functions form a complete, self-contained unit of functionality. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md deleted file mode 100644 index 290b37c3..00000000 --- a/docs/reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md +++ /dev/null @@ -1,260 +0,0 @@ -# C Family Analyzers - -## Introduction - -The C Family Analyzers module provides static-analysis front-ends for three closely related C-style languages: **C**, **C++**, and **C#**. Each analyzer parses a single source file with a language-specific [tree-sitter](https://tree-sitter.github.io/tree-sitter/) grammar, walks the resulting concrete syntax tree, and produces two normalized outputs: - -- **Nodes** — structural components such as functions, classes, structs, methods, namespaces, and global variables -- **Call Relationships** — edges describing how those components interact (calls, inheritance, instantiation, field/property usage) - -These outputs conform to the shared `Node` and `CallRelationship` data models used across the entire dependency-analysis pipeline, allowing the C-family analyzers to be interchangeable with analyzers for other languages such as Java, Python, PHP, and JavaScript/TypeScript. - -The module is a leaf-level component in the analyzer layer: it has no knowledge of cross-file resolution, graph construction, or clustering — that responsibility belongs to higher-level orchestration components described later in this document. - -## Purpose and Scope - -| Analyzer | Component | Source Extensions | Grammar Package | -|---|---|---|---| -| C | `TreeSitterCAnalyzer` | `.c`, `.h` | `tree_sitter_c` | -| C++ | `TreeSitterCppAnalyzer` | `.cpp`, `.cc`, `.cxx`, `.hpp`, `.h` | `tree_sitter_cpp` | -| C# | `TreeSitterCSharpAnalyzer` | `.cs` | `tree_sitter_c_sharp` | - -Each analyzer is instantiated per-file with the file path, raw source content, and (optionally) the repository root path. Construction is eager: the constructor immediately runs the full analysis pipeline (`_analyze()`), after which the caller reads the populated `nodes` and `call_relationships` lists. - -```python -analyzer = TreeSitterCAnalyzer(file_path, content, repo_path) -functions_and_structs = analyzer.nodes -relationships = analyzer.call_relationships -``` - -Each analyzer file also exposes a thin module-level convenience function (`analyze_c_file`, `analyze_cpp_file`, `analyze_csharp_file`) that wraps this construction pattern and returns a `(nodes, call_relationships)` tuple — this is the entry point used by upstream orchestration. - -## Architecture - -All three analyzers share an identical structural pattern despite differences in the underlying grammars: a two-pass traversal (node extraction, then relationship extraction) built on top of tree-sitter's parse tree. - -```mermaid -flowchart TD - Source["Source File Content"] --> Parser["tree-sitter Parser"] - Parser --> Tree["Concrete Syntax Tree"] - Tree --> Extract["Node Extraction Pass"] - Extract --> TopLevel["Top-Level Node Registry"] - TopLevel --> Rel["Relationship Extraction Pass"] - Rel --> Nodes["List of Node objects"] - Rel --> Calls["List of CallRelationship objects"] -``` - -### Common Construction Flow - -1. **Language binding** — Each analyzer loads its grammar via the corresponding `tree_sitter_*` package and wraps it in a `Language`/`Parser` pair. -2. **Parsing** — The raw UTF-8 file content is parsed once into a `tree_sitter` syntax tree. -3. **Node extraction (`_extract_nodes`)** — A recursive depth-first traversal identifies structurally significant syntax node types (e.g. `function_definition`, `class_specifier`, `struct_specifier`) and converts them into a shared `Node` model, while also registering them in a local `top_level_nodes` dictionary keyed by name (or qualified name for C++ methods). -4. **Relationship extraction (`_extract_relationships`)** — A second recursive traversal inspects call expressions, inheritance clauses, instantiation expressions, and identifier usages, cross-referencing them against `top_level_nodes` to emit `CallRelationship` edges. - -### Component ID Convention - -Every analyzer derives a fully-qualified component identifier from the file's path relative to the repository root, converting path separators to dots and stripping the language-specific file extension: - -```text -path/to/file.cpp -> path.to.file -component_id = "path.to.file::FunctionName" -``` - -For C++ methods, the identifier additionally embeds the containing class using dot notation: `path.to.file::ClassName.methodName`. This mirrors the identifier scheme used by the other language analyzers so that identifiers remain comparable across the whole codebase graph. - -## Component Details - -### TreeSitterCAnalyzer - -Parses C source using the `tree_sitter_c` grammar. It recognizes: - -- **Functions** (`function_definition`) — extracted via the `function_declarator` → `identifier` chain -- **Structs** (`struct_specifier`, and `typedef struct { ... } Name;` via `type_definition`) -- **Global variables** (`declaration` nodes not nested inside a function body) - -Relationship extraction covers two cases: -- **Function calls** (`call_expression`) — the callee is recorded by its *simple name only* (`is_resolved=False`), deferring cross-file resolution to the call-graph analyzer described below. A hardcoded set of common C-standard-library and SDL functions (`printf`, `malloc`, `SDL_Init`, etc.) is filtered out to avoid noise. -- **Global variable usage** — identifiers inside a function body that match a known global variable are recorded as resolved, same-file relationships (`is_resolved=True`). - -Only `function` and `struct` node types are appended to the public `nodes` list; global `variable` nodes are tracked internally (in `top_level_nodes`) purely to support usage-relationship detection. - -### TreeSitterCppAnalyzer - -Parses C++ using the `tree_sitter_cpp` grammar. It extends the C model with object-oriented and namespace constructs: - -- **Classes and structs** (`class_specifier`, `struct_specifier`) -- **Functions and methods** (`function_definition`) — distinguished by walking up the parent chain to detect an enclosing `class_specifier`/`struct_specifier`; methods are keyed in `top_level_nodes` by `ClassName.methodName` -- **Namespaces** (`namespace_definition`) -- **Global variables** (`declaration` nodes outside any function/class/struct body) - -Relationship extraction is the richest of the three analyzers, detecting: - -| Relationship Type | Triggering Syntax | Notes | -|---|---|---| -| `calls` | `call_expression` | Resolves plain function calls and, for method calls via `field_expression`, attempts to locate the owning class through `_find_class_containing_method` | -| `inherits` | `base_class_clause` | Extracts base `type_identifier` names | -| `creates` | `new_expression` | Detects object instantiation (`new ClassName(...)`) | -| `uses` | bare `identifier` | Detects references to global variables from within functions/methods | - -A hardcoded system-function filter (`printf`, `cout`, `new`, `delete`, etc.) suppresses standard-library noise, mirroring the C analyzer. - -### TreeSitterCSharpAnalyzer - -Parses C# using the `tree_sitter_c_sharp` grammar. Its node vocabulary is the broadest of the three, reflecting C#'s richer type-declaration surface: - -- **Classes** — further classified as `class`, `abstract class`, or `static class` based on detected `modifier` nodes -- **Interfaces** (`interface_declaration`) -- **Structs** (`struct_declaration`) -- **Enums** (`enum_declaration`) -- **Records** (`record_declaration`) -- **Delegates** (`delegate_declaration`) - -Unlike the C and C++ analyzers, the C# analyzer does **not** attempt call-expression resolution. Instead, it focuses on **type-usage relationships** that reflect C#'s declarative, strongly-typed structure: - -| Relationship Type | Triggering Syntax | Notes | -|---|---|---| -| inheritance/implementation | `class_declaration` → `base_list` | Emits a resolved relationship (`is_resolved=True`) only when the base name matches another top-level node in the same file | -| property type usage | `property_declaration` | Emits an unresolved relationship (`is_resolved=False`) when the property type is non-primitive | -| field type usage | `field_declaration` | Same pattern as properties | -| parameter type usage | `method_declaration` → `parameter_list` | Emits an unresolved relationship per non-primitive parameter type | - -A `_is_primitive_type` allow-list (C# built-ins like `int`, `string`, `List`, `Dictionary`, `Task`, `DateTime`, etc.) prevents common framework types from polluting the relationship graph. - -## Data Model - -All three analyzers populate the shared `Node` and `CallRelationship` structures defined in the dependency-analyzer models layer. Key fields populated by every C-family analyzer include: - -- `Node`: `id`, `name`, `component_type`/`node_type`, `file_path`, `relative_path`, `source_code`, `start_line`, `end_line`, `display_name`, `component_id`, and (for C++ methods) `class_name` -- `CallRelationship`: `caller`, `callee`, `call_line`, `is_resolved`, and optionally `relationship_type` (C++ only: `calls`, `inherits`, `creates`, `uses`) - -Notably, the three analyzers differ in how aggressively they resolve `is_resolved`: -- The **C** analyzer defers function-call resolution entirely (`is_resolved=False`), but resolves same-file variable usage. -- The **C++** analyzer resolves relationships whenever a callee/base-class name matches a `top_level_nodes` entry in the same file. -- The **C#** analyzer resolves inheritance within the same file but leaves type-usage relationships unresolved, since types may be declared elsewhere. - -For full schema definitions, see the dependency analyzer models. - -## Process Flow: Node and Relationship Extraction - -The following sequence illustrates the two-pass traversal shared by all three analyzers, using the C++ analyzer as a representative example: - -```mermaid -sequenceDiagram - participant Caller as "analyze_cpp_file()" - participant Analyzer as "TreeSitterCppAnalyzer" - participant TS as "tree-sitter Parser" - participant Pass1 as "_extract_nodes()" - participant Pass2 as "_extract_relationships()" - - Caller->>Analyzer: "__init__(file_path, content, repo_path)" - Analyzer->>TS: "parse(content)" - TS-->>Analyzer: "syntax tree" - Analyzer->>Pass1: "traverse(root, top_level_nodes, lines)" - Pass1->>Pass1: "detect class/struct/function/namespace/variable" - Pass1-->>Analyzer: "populated top_level_nodes + nodes list" - Analyzer->>Pass2: "traverse(root, top_level_nodes)" - Pass2->>Pass2: "detect calls, inheritance, instantiation, usage" - Pass2-->>Analyzer: "call_relationships list" - Analyzer-->>Caller: "nodes, call_relationships" -``` - -## Component Relationships - -```mermaid -classDiagram - class TreeSitterCAnalyzer { - +file_path - +content - +repo_path - +nodes - +call_relationships - -_analyze() - -_extract_nodes() - -_extract_relationships() - } - class TreeSitterCppAnalyzer { - +file_path - +content - +repo_path - +nodes - +call_relationships - -_analyze() - -_extract_nodes() - -_extract_relationships() - } - class TreeSitterCSharpAnalyzer { - +file_path - +content - +repo_path - +nodes - +call_relationships - -_analyze() - -_extract_nodes() - -_extract_relationships() - } - class Node { - +id - +name - +component_type - +file_path - +source_code - } - class CallRelationship { - +caller - +callee - +call_line - +is_resolved - } - TreeSitterCAnalyzer --> Node : produces - TreeSitterCAnalyzer --> CallRelationship : produces - TreeSitterCppAnalyzer --> Node : produces - TreeSitterCppAnalyzer --> CallRelationship : produces - TreeSitterCSharpAnalyzer --> Node : produces - TreeSitterCSharpAnalyzer --> CallRelationship : produces -``` - -## Integration with the Dependency Analysis Pipeline - -The C Family Analyzers do not run in isolation. They are invoked per-file by the language-dispatching parser, and their raw (sometimes unresolved) output is later consolidated by the call-graph resolution stage: - -```mermaid -flowchart LR - Repo["Repository Source Files"] --> Parser["Dependency Parser"] - Parser -->|".c / .h"| CAnalyzer["TreeSitterCAnalyzer"] - Parser -->|".cpp / .hpp"| CppAnalyzer["TreeSitterCppAnalyzer"] - Parser -->|".cs"| CSharpAnalyzer["TreeSitterCSharpAnalyzer"] - CAnalyzer --> RawNodes["Raw Nodes + Relationships"] - CppAnalyzer --> RawNodes - CSharpAnalyzer --> RawNodes - RawNodes --> GraphBuilder["Dependency Graph Builder"] - RawNodes --> CallGraph["Call Graph Analyzer"] - CallGraph --> Resolved["Resolved Cross-File Relationships"] - GraphBuilder --> FinalGraph["Repository Dependency Graph"] - Resolved --> FinalGraph -``` - -- The **Dependency Parser** (`DependencyParser`) selects the appropriate analyzer based on file extension and orchestrates per-file invocation across the whole repository. -- The **Call Graph Analyzer** (`CallGraphAnalyzer`) consumes the unresolved `callee` names emitted by these analyzers (particularly from the C and C++ analyzers) and resolves them into fully-qualified cross-file identifiers. -- The **Dependency Graph Builder** (`DependencyGraphBuilder`) assembles the final `Node`/`CallRelationship` collections from all language analyzers — including this module — into the unified repository dependency graph. - -These orchestration components live in the backend's dependency analyzer core, which handles graph construction and the analysis pipeline for details on how per-file analyzer output is consolidated and resolved. - -The shared `Node`, `CallRelationship`, and related schema types used by all three analyzers are defined in the dependency analyzer models. - -## Relationship to Other Language Analyzers - -The C Family Analyzers module is one of several sibling analyzer groups under [Tree Sitter Analyzers](tree-sitter-analyzers.md), each following the same node/relationship extraction contract but tailored to a different language's grammar and idioms: - -- [Java Analyzer](java_analyzer.md) -- [JavaScript/TypeScript Analyzers](javascript_typescript_analyzers.md) -- [PHP Analyzer](php_analyzer.md) -- [Python Analyzer](python_analyzer.md) - -Because all analyzers emit the same `Node`/`CallRelationship` shapes, the downstream graph-construction and call-resolution logic remains language-agnostic — new language support can be added by implementing a new analyzer following this same pattern without modifying the orchestration layer. - -## Design Notes and Limitations - -- **Per-file, single-pass scope**: Each analyzer only sees one file at a time and has no visibility into other files in the repository. Cross-file symbol resolution (e.g. resolving a C function call to its definition in another translation unit) is intentionally deferred to the call-graph resolution stage. -- **Heuristic-based class/method matching**: The C++ analyzer's `_class_has_method` and `_find_class_containing_method` use lightweight source-text heuristics (substring matching on method signatures) rather than full semantic type resolution, since tree-sitter provides only syntactic — not semantic — information. -- **Standard-library filtering is hardcoded**: Both the C and C++ analyzers maintain fixed allow-lists of common standard-library/system function names to exclude from the relationship graph. This is a pragmatic simplification rather than a complete standard-library model. -- **C# favors type usage over call resolution**: Unlike C and C++, the C# analyzer does not attempt to trace method call expressions; it instead surfaces structural type dependencies (inheritance, field/property/parameter types), which are typically more informative for understanding C# codebases dominated by object-oriented composition. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md deleted file mode 100644 index 663d7f3e..00000000 --- a/docs/reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md +++ /dev/null @@ -1,178 +0,0 @@ -# Java Analyzer - -The Java Analyzer module is a language-specific static analysis component within CodeWiki's dependency analysis pipeline. It uses [tree-sitter](https://tree-sitter.github.io/tree-sitter/) to parse Java source files into an Abstract Syntax Tree (AST) and extracts both **structural components** (classes, interfaces, enums, records, annotations, and methods) and **relationships** between them (inheritance, interface implementation, field type usage, method invocations, and object instantiation). The resulting `Node` and `CallRelationship` objects feed into the broader dependency graph that CodeWiki uses to power documentation generation and repository visualization. - -This module is one of several language-specific analyzers registered under the [Tree Sitter Analyzers](tree-sitter-analyzers.md) module, sitting alongside analyzers for C/C++/C#, JavaScript/TypeScript, PHP, and Python. - -## Purpose and Scope - -The sole responsibility of the Java Analyzer is to convert raw Java source text into a set of normalized, language-agnostic data structures that downstream components (such as the dependency analyzer core and its graph construction pipeline) can consume without needing to know anything about Java syntax. - -Specifically, the analyzer: - -1. Parses a single `.java` file's source text using the `tree-sitter-java` grammar. -2. Walks the resulting AST to identify top-level and nested structural declarations. -3. Assigns each declaration a fully-qualified component ID in the `module.path::ClassName` (or `module.path::ClassName.methodName`) format expected by the rest of the system. -4. Walks the AST a second time to detect relationships between declarations (inheritance, interface implementation, field types, method calls, and object creation). -5. Exposes the extracted `Node` and `CallRelationship` lists for consumption by the parsing pipeline. - -## Core Component - -### TreeSitterJavaAnalyzer - -`TreeSitterJavaAnalyzer` (defined in `codewiki/src/be/dependency_analyzer/analyzers/java.py`) is the single class that implements the entire analysis workflow for one Java file. It is instantiated with the file's path, its content, and (optionally) the repository root path used to compute relative/module paths. - -```python -class TreeSitterJavaAnalyzer: - def __init__(self, file_path: str, content: str, repo_path: str = None): - ... - self._analyze() -``` - -On construction, the analyzer immediately runs its full pipeline (`_analyze()`), populating two public attributes: - -- `self.nodes: List[Node]` — the structural components discovered in the file. -- `self.call_relationships: List[CallRelationship]` — the relationships discovered between those components (and to external/unresolved types). - -A module-level convenience function is also provided: - -```python -def analyze_java_file(file_path: str, content: str, repo_path: str = None) -> Tuple[List[Node], List[CallRelationship]]: - analyzer = TreeSitterJavaAnalyzer(file_path, content, repo_path) - return analyzer.nodes, analyzer.call_relationships -``` - -This function is the primary entry point that callers in the dependency analysis pipeline are expected to use, mirroring the calling convention of the other tree-sitter based analyzers (C, C++, C#, JavaScript, TypeScript, PHP). - -## Architecture Overview - -The diagram below shows how `TreeSitterJavaAnalyzer` fits into the surrounding analysis pipeline and which data models it depends on. - -```mermaid -flowchart TD - subgraph Pipeline["Dependency Analysis Pipeline"] - CGA["CallGraphAnalyzer"] -->|"dispatches .java files"| JA["TreeSitterJavaAnalyzer"] - end - - subgraph JavaAnalyzerModule["Java Analyzer"] - JA -->|"1: parse source"| TS["tree-sitter-java grammar"] - TS -->|"AST"| EN["_extract_nodes()"] - TS -->|"AST"| ER["_extract_relationships()"] - EN -->|"appends"| Nodes["self.nodes: List[Node]"] - ER -->|"appends"| Rels["self.call_relationships: List[CallRelationship]"] - end - - Nodes -->|"consumed by"| DP["DependencyParser"] - Rels -->|"consumed by"| DP - DP -->|"builds"| Graph["Dependency Graph"] - - Node["Node model"] -.->|"defines schema for"| Nodes - CallRel["CallRelationship model"] -.->|"defines schema for"| Rels -``` - -- **CallGraphAnalyzer**, part of the dependency analyzer's analysis pipeline (which lives under the dependency analyzer core), dispatches individual source files to the appropriate language analyzer based on file extension. For `.java` files, it invokes `TreeSitterJavaAnalyzer`/`analyze_java_file`. -- **DependencyParser**, part of the dependency analyzer core's graph construction stage, aggregates the `Node` and `CallRelationship` objects produced by all analyzers (Java and otherwise) into a unified, namespaced dependency graph. -- **Node** and **CallRelationship** are shared Pydantic models defined in the dependency analyzer models. The Java Analyzer produces instances of these models but does not define them itself. - -## Internal Processing Flow - -`_analyze()` performs two independent, full-tree traversals over the same parsed AST: one for structural nodes, and one for relationships. This two-pass design ensures that all top-level declarations are known (via `top_level_nodes`) before relationships that reference them (e.g., method calls, field types) are resolved. - -```mermaid -sequenceDiagram - participant Caller as "analyze_java_file()" - participant Analyzer as "TreeSitterJavaAnalyzer" - participant Parser as "tree_sitter.Parser" - participant NodesPass as "_extract_nodes()" - participant RelsPass as "_extract_relationships()" - - Caller->>Analyzer: "__init__(file_path, content, repo_path)" - Analyzer->>Analyzer: "_analyze()" - Analyzer->>Parser: "parse(content)" - Parser-->>Analyzer: "AST root node" - Analyzer->>NodesPass: "_extract_nodes(root, top_level_nodes, lines)" - NodesPass-->>Analyzer: "self.nodes populated" - Analyzer->>RelsPass: "_extract_relationships(root, top_level_nodes)" - RelsPass-->>Analyzer: "self.call_relationships populated" - Analyzer-->>Caller: "nodes, call_relationships" -``` - -### Pass 1: Structural Node Extraction (`_extract_nodes`) - -This recursive method walks every node in the AST looking for the following Java declaration types: - -| AST Node Type | Resulting `component_type` | -|---|---| -| `class_declaration` (with `abstract` modifier) | `abstract class` | -| `class_declaration` (without `abstract` modifier) | `class` | -| `interface_declaration` | `interface` | -| `enum_declaration` | `enum` | -| `record_declaration` | `record` | -| `annotation_type_declaration` | `annotation` | -| `method_declaration` | `method` | - -For each match, the analyzer builds a fully-qualified `component_id` via `_get_component_id()`, which combines the file's module path (derived from its path relative to the repository root, with `/` replaced by `.`) and the declaration name using the `module.path::Name` convention (or `module.path::ClassName.methodName` for methods, where the containing class is resolved via `_find_containing_class_name`). - -A `Node` instance is constructed with the extracted source snippet, line ranges, and a human-readable `display_name` (e.g., `"class UserService"`), then appended to `self.nodes` and registered in the `top_level_nodes` dictionary for use during relationship extraction. - -### Pass 2: Relationship Extraction (`_extract_relationships`) - -This second recursive traversal identifies five categories of relationships, each producing one or more `CallRelationship` entries: - -```mermaid -flowchart LR - A["class_declaration
with superclass"] -->|"1: Inheritance"| R1["CallRelationship
class extends BaseClass"] - B["class/enum/record
with super_interfaces"] -->|"2: Interface Implementation"| R2["CallRelationship
class implements Interface"] - C["field_declaration"] -->|"3: Field Type Use"| R3["CallRelationship
class has field of Type"] - D["method_invocation"] -->|"4: Method Call"| R4["CallRelationship
caller calls object.method()"] - E["object_creation_expression"] -->|"5: Object Creation"| R5["CallRelationship
class creates new Type()"] -``` - -1. **Inheritance** — For `class_declaration` nodes with a `superclass` child, a relationship is created from the class to its base class (unless the base class is a Java primitive/built-in type). -2. **Interface Implementation** — For classes, enums, or records with a `super_interfaces` clause, a relationship is created to each implemented interface. -3. **Field Type Use** — For each `field_declaration`, if the containing class is resolvable and the field's type is a non-primitive type, a relationship is recorded from the class to the field's type. -4. **Method Calls** — For each `method_invocation`, the analyzer attempts to resolve the invoked object's declared type: first by checking if the object name matches a known top-level declaration, then by searching local variable declarations (`_search_variable_declaration`) and field declarations (`_find_variable_type`) within the enclosing method/class. If a type is resolved, a relationship is recorded from the calling method (or containing class, if outside a method) to the resolved type. -5. **Object Creation** — For each `object_creation_expression`, a relationship is recorded from the containing class to the instantiated type. - -All relationships are created with `is_resolved=False`, since the Java Analyzer only performs local, per-file resolution — full cross-file/cross-module resolution is handled later by the dependency analyzer core when building the complete dependency graph. - -### Type Filtering - -`_is_primitive_type()` filters out Java primitives (`int`, `boolean`, `char`, etc.), their boxed equivalents (`Integer`, `Boolean`, `Character`, etc.), and common JDK built-ins (`String`, `Object`, `List`, `Set`, `Map`, `Collection`, `Optional`, `void`, `Void`) so that relationships are only recorded for application-relevant types, keeping the dependency graph focused on meaningful code relationships rather than noise from standard library usage. - -## Data Model Reference - -The Java Analyzer produces instances of two shared Pydantic models defined outside this module: - -- **`Node`** — represents a structural component (class, interface, method, etc.) with fields such as `id`, `name`, `component_type`, `file_path`, `source_code`, `start_line`/`end_line`, and `display_name`. The `id` and `component_id` fields hold the same value: the analyzer's locally-computed `module.path::Name` identifier, which is later re-namespaced by the `DependencyParser` into a full FQDN. -- **`CallRelationship`** — represents a directed edge between a `caller` and `callee` component ID, with an optional `call_line` and an `is_resolved` flag. - -For full field definitions and usage across other analyzers, see the dependency analyzer models module. - -## Component ID Convention - -A key contract that `TreeSitterJavaAnalyzer` must honor is the component ID format expected by the rest of the system: `module.path::ClassName` (and `module.path::ClassName.methodName` for methods). This is implemented via two helper methods: - -- `_get_module_path()` — converts the file's path (relative to `repo_path`, if provided) into a dotted module path, stripping the `.java` extension and replacing path separators with dots. -- `_get_component_id(name, parent_class=None)` — combines the module path with the declaration name (optionally qualified by a parent class name) using the `::` separator. - -This convention ensures that IDs produced by the Java Analyzer are structurally consistent with those produced by the other analyzers under [Tree Sitter Analyzers](tree-sitter-analyzers.md) (C/C++/C#, JavaScript/TypeScript, PHP, Python), allowing the Dependency Parser to merge components from multiple languages and repositories into a single, namespaced dependency graph. - -## Integration Points - -| Consumer | Relationship | -|---|---| -| Analysis Pipeline (`CallGraphAnalyzer`) | Invokes `analyze_java_file()` for each `.java` file discovered during repository structure analysis. | -| Graph Construction (`DependencyParser`) | Consumes the raw `Node`/`CallRelationship` dictionaries (converted to `functions`/`relationships` lists) to build namespaced components and resolve dependencies, including cross-namespace resolution for multi-repository analysis. | -| Dependency Analyzer Models | Supplies the `Node` and `CallRelationship` schemas that this analyzer instantiates. | - -## Related Analyzers - -The Java Analyzer is one of several sibling language analyzers grouped under [Tree Sitter Analyzers](tree-sitter-analyzers.md): - -- [C Family Analyzers](c_family_analyzers.md) — C, C++, and C# analysis -- [JavaScript/TypeScript Analyzers](javascript_typescript_analyzers.md) -- [PHP Analyzer](php_analyzer.md) -- [Python Analyzer](python_analyzer.md) - -All of these analyzers follow the same general contract — accept a file path and content, return `Node` and `CallRelationship` lists — even though each implements language-specific AST traversal logic suited to its grammar. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md deleted file mode 100644 index da8498a6..00000000 --- a/docs/reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md +++ /dev/null @@ -1,229 +0,0 @@ -# JavaScript & TypeScript Analyzers - -The JavaScript & TypeScript Analyzers module implements the tree-sitter based source-code analyzers responsible for parsing JavaScript and TypeScript files and extracting the structural entities (functions, classes, interfaces, methods, type aliases, enums, etc.) and the relationships between them (calls, inheritance, type usage) that feed the dependency graph built by the backend analysis pipeline. - -It provides two closely related analyzer implementations: - -- `TreeSitterJSAnalyzer` — parses plain JavaScript (and JSX/mixed) source using the `tree_sitter_javascript` grammar. -- `TreeSitterTSAnalyzer` — parses TypeScript (and its embedded JavaScript constructs) using the `tree_sitter_typescript` grammar, adding support for TypeScript-only constructs such as interfaces, type aliases, enums, and type annotations. - -Both analyzers produce the same output contract — a list of `Node` objects and a list of `CallRelationship` objects — defined in the dependency analyzer models, so that downstream components can treat all language analyzers uniformly. - -## Role in the Analysis Pipeline - -This module is one of several language-specific analyzer implementations grouped under the [Tree Sitter Analyzers](tree-sitter-analyzers.md) parent module, alongside the [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), [PHP Analyzer](php_analyzer.md), and [Python Analyzer](python_analyzer.md). - -The analyzers in this module are invoked by the `DependencyParser` in the dependency analyzer core's graph construction stage, which dispatches source files to the appropriate analyzer based on file extension. The resulting `Node` and `CallRelationship` objects are then consumed by the `DependencyGraphBuilder` to assemble the full repository dependency graph, and by the analysis pipeline (`AnalysisService`, `RepoAnalyzer`, `CallGraphAnalyzer`) for higher-level repository analysis. - -```mermaid -flowchart LR - DP["DependencyParser"] -->|"dispatches .js/.jsx/.mjs/.cjs files"| JSA["TreeSitterJSAnalyzer"] - DP -->|"dispatches .ts/.tsx files"| TSA["TreeSitterTSAnalyzer"] - JSA -->|"produces"| Nodes["Node objects"] - JSA -->|"produces"| Rels["CallRelationship objects"] - TSA -->|"produces"| Nodes - TSA -->|"produces"| Rels - Nodes --> DGB["DependencyGraphBuilder"] - Rels --> DGB - DGB --> Graph["Repository Dependency Graph"] -``` - -## Component Overview - -| Component | File | Responsibility | -|---|---|---| -| `TreeSitterJSAnalyzer` | `codewiki/src/be/dependency_analyzer/analyzers/javascript.py` | Parses JavaScript source with tree-sitter, extracts functions/classes/methods and call/inheritance/JSDoc-type relationships | -| `analyze_javascript_file_treesitter` | same file | Module-level convenience function that instantiates `TreeSitterJSAnalyzer`, runs `analyze()`, and returns `(nodes, relationships)` | -| `TreeSitterTSAnalyzer` | `codewiki/src/be/dependency_analyzer/analyzers/typescript.py` | Parses TypeScript source with tree-sitter, extracts a broader set of entities (interfaces, type aliases, enums, exports) and relationships (calls, `new`, member access, type annotations, inheritance) | -| `analyze_typescript_file_treesitter` | same file | Module-level convenience function that instantiates `TreeSitterTSAnalyzer`, runs `analyze()`, and returns `(nodes, relationships)` | - -Both classes share the same overall shape: they accept `file_path`, `content`, and an optional `repo_path`, build a `tree_sitter.Parser` bound to the relevant `Language`, walk the resulting AST, and accumulate `Node` and `CallRelationship` instances (as defined in the dependency analyzer models). - -```mermaid -classDiagram - class TreeSitterJSAnalyzer { - +file_path: Path - +content: str - +repo_path: str - +nodes: List~Node~ - +call_relationships: List~CallRelationship~ - +top_level_nodes: dict - +analyze() None - -_extract_functions(node) None - -_extract_call_relationships(node) None - -_extract_jsdoc_type_dependencies(node, caller) None - } - class TreeSitterTSAnalyzer { - +file_path: Path - +content: str - +repo_path: str - +nodes: List~Node~ - +call_relationships: List~CallRelationship~ - +top_level_nodes: dict - +analyze() None - -_extract_all_entities(node, all_entities, depth) None - -_filter_top_level_declarations(all_entities) None - -_extract_all_relationships(node, all_entities) None - } - class Node { - +id: str - +name: str - +component_type: str - +source_code: str - +start_line: int - +end_line: int - +node_type: str - +base_classes: List~str~ - } - class CallRelationship { - +caller: str - +callee: str - +call_line: int - +is_resolved: bool - } - TreeSitterJSAnalyzer --> Node : produces - TreeSitterJSAnalyzer --> CallRelationship : produces - TreeSitterTSAnalyzer --> Node : produces - TreeSitterTSAnalyzer --> CallRelationship : produces -``` - -## TreeSitterJSAnalyzer - -`TreeSitterJSAnalyzer` initializes a tree-sitter `Parser` bound to the `tree_sitter_javascript` grammar. If grammar/language initialization fails, the parser is left `None` and `analyze()` becomes a no-op (logged as a warning), which keeps the pipeline resilient to environment issues. - -### Extraction flow - -1. **`analyze()`** parses the source into an AST and calls `_extract_functions()` followed by `_extract_call_relationships()`. -2. **`_extract_functions()` / `_traverse_for_functions()`** walk the AST recursively, recognizing: - - `class_declaration`, `abstract_class_declaration`, `interface_declaration` → extracted via `_extract_class_declaration()`, with their `class_body` scanned for `method_definition` and arrow-function `field_definition` members via `_extract_methods_from_class()`. - - `function_declaration` / `generator_function_declaration` (only when not nested in a class) → `_extract_function_declaration()`. - - `export_statement` → `_extract_exported_function()` (handles `export function` and `export default function`). - - `lexical_declaration` (`const`/`let`) → `_extract_arrow_function_from_declaration()` for arrow functions and function expressions assigned to a variable. -3. **`_extract_call_relationships()` / `_traverse_for_calls()`** re-walks the tree tracking the "current top-level" entity (the enclosing function or class), and records: - - `call_expression` and `await_expression`-wrapped calls → resolved via `_extract_call_from_node()`, checking `this.`/`super.` member calls against known methods to avoid duplicate self-references. - - `new_expression` → constructor instantiation relationships. - - Class `class_heritage` (`extends`) → inheritance `CallRelationship` entries. - - JSDoc comments (`@param {Type}`, `@returns {Type}`, `@type {Type}`, `@typedef`, `@interface`) → parsed by `_parse_jsdoc_types()` / `_extract_base_types_from_jsdoc()` into type-dependency relationships, filtered against a built-in type list via `_is_builtin_type_js()`. - -All relationships are de-duplicated through `_add_relationship()`, which tracks a `(caller, callee, call_line)` key in `seen_relationships`. - -```mermaid -flowchart TD - Start["analyze()"] --> Parse["Parser.parse(content)"] - Parse --> ExtractFn["_extract_functions(root_node)"] - ExtractFn --> TraverseFn["_traverse_for_functions(node)"] - TraverseFn -->|"class/abstract/interface"| ExtractClass["_extract_class_declaration()"] - ExtractClass --> ExtractMethods["_extract_methods_from_class()"] - TraverseFn -->|"function_declaration"| ExtractFunc["_extract_function_declaration()"] - TraverseFn -->|"export_statement"| ExtractExport["_extract_exported_function()"] - TraverseFn -->|"lexical_declaration"| ExtractArrow["_extract_arrow_function_from_declaration()"] - ExtractFn --> ExtractCalls["_extract_call_relationships(root_node)"] - ExtractCalls --> TraverseCalls["_traverse_for_calls(node, current_top_level)"] - TraverseCalls -->|"call_expression"| CallRel["_extract_call_from_node()"] - TraverseCalls -->|"new_expression"| NewRel["Constructor relationship"] - TraverseCalls -->|"class_heritage"| InheritRel["Inheritance relationship"] - TraverseCalls -->|"JSDoc comment"| JsdocRel["_parse_jsdoc_types()"] - CallRel --> AddRel["_add_relationship() (dedup)"] - NewRel --> AddRel - InheritRel --> AddRel - JsdocRel --> AddRel -``` - -### Component identity - -Component IDs are built by `_get_component_id()` as `::` for top-level entities and `::.` for methods, where `_get_module_path()` derives a dotted module path from the file's path relative to `repo_path`, stripping `.js`/`.ts`/`.jsx`/`.tsx`/`.mjs`/`.cjs` extensions. Note that internally, call/inheritance relationships built during traversal use a distinct dotted convention (`.`) rather than the `::` component-id separator; graph construction downstream reconciles these against the canonical `Node.id` values. - -## TreeSitterTSAnalyzer - -`TreeSitterTSAnalyzer` follows a two-pass design that is more elaborate than the JS analyzer, reflecting TypeScript's richer set of top-level declaration forms. - -### Pass 1 — Entity collection (`_extract_all_entities`) - -A single recursive traversal builds a flat `all_entities` dictionary keyed by entity name, capturing one of the following node types at any depth: -`function_declaration`, `generator_function_declaration`, `arrow_function`, `method_definition`, `class_declaration`, `abstract_class_declaration`, `interface_declaration`, `type_alias_declaration`, `enum_declaration`, `variable_declarator`, `export_statement`, `lexical_declaration`, `variable_declaration`, `ambient_declaration`. - -Each captured entity dict stores `name`, `type`/`subtype`, `code_snippet`, `display_name`, line range, and (for functions) `parameters`, plus bookkeeping fields `depth`, `node` (the raw tree-sitter node), and `parent_context` (via `_get_parent_context()`). - -### Pass 2 — Top-level filtering (`_filter_top_level_declarations`) - -For every collected entity, `_is_actually_top_level()` walks up the parent chain to decide whether the entity is genuinely a module-level declaration (directly under `program`, `export_statement`, `ambient_declaration`, `module`, or a `statement_block` that is itself inside a `module`/`ambient_declaration`) versus nested inside a function body (checked via `_is_inside_function_body()`). Only entities classified as top-level are converted into `Node` objects (via `_create_node_from_entity()`) and filtered further by `_should_include_node()`, which excludes bare `variable` entities and reserved names (`constructor`, `__proto__`, `prototype`). - -For class/abstract-class entities, `_extract_constructor_dependencies()` additionally inspects the constructor's `formal_parameters` for TypeScript type annotations and records a dependency relationship per typed parameter via `_extract_parameter_dependencies()`. - -### Pass 3 — Relationship extraction (`_extract_all_relationships`) - -A third traversal (`_traverse_for_relationships`) tracks the current enclosing top-level entity (updated whenever a "new top-level" node type is encountered per `_is_new_top_level()`/`_get_top_level_name()`) and, for each descendant node, extracts: - -| Node type | Handler | Relationship captured | -|---|---|---| -| `call_expression` | `_extract_call_relationship()` | Function/method calls, filtering `this.`/`super.` calls to nested methods of the same class | -| `new_expression` | `_extract_new_relationship()` | Constructor instantiation | -| `member_expression` | `_extract_member_relationship()` | Property/member access | -| `type_annotation` | `_extract_type_relationship()` | Parameter/return type usage, filtered by `_is_builtin_type()` | -| `type_arguments` | `_extract_type_arguments_relationship()` | Generic type parameters | -| `extends_clause` / `implements_clause` | `_extract_inheritance_relationship()` | Class/interface inheritance and interface implementation | - -All relationships are appended via `_add_relationship()`, which builds `caller`/`callee` ids as `.` and marks them `is_resolved=False` (resolution against the full repository graph happens downstream in the dependency analyzer core's graph construction stage). - -```mermaid -flowchart TD - A["analyze()"] --> B["Pass 1: _extract_all_entities()
builds all_entities dict"] - B --> C["Pass 2: _filter_top_level_declarations()
_is_actually_top_level() check"] - C --> D["Node objects appended to self.nodes
and self.top_level_nodes"] - D --> E["_extract_constructor_dependencies()
(classes only)"] - B --> F["Pass 3: _extract_all_relationships()"] - F --> G["_traverse_for_relationships()
tracks current_top_level"] - G -->|"call_expression"| H["_extract_call_relationship()"] - G -->|"new_expression"| I["_extract_new_relationship()"] - G -->|"member_expression"| J["_extract_member_relationship()"] - G -->|"type_annotation / type_arguments"| K["_extract_type_relationship()"] - G -->|"extends_clause / implements_clause"| L["_extract_inheritance_relationship()"] - H --> M["_add_relationship()"] - I --> M - J --> M - K --> M - L --> M -``` - -## JS vs TS Analyzer: Key Differences - -| Aspect | `TreeSitterJSAnalyzer` | `TreeSitterTSAnalyzer` | -|---|---|---| -| Grammar | `tree_sitter_javascript` | `tree_sitter_typescript` (`language_typescript()`) | -| Traversal design | Two combined traversals (functions, then calls) inline while walking | Three explicit passes: entity collection, top-level filtering, relationship extraction | -| TypeScript-only entities | Not supported | `interface_declaration`, `type_alias_declaration`, `enum_declaration`, `ambient_declaration` | -| Type-dependency source | JSDoc comments (`@param`, `@returns`, `@type`, `@typedef`, `@interface`) parsed with regex | Native TS syntax: `type_annotation`, `type_arguments`, `extends_clause`/`implements_clause` | -| Constructor parameter typing | Not applicable (no static types) | `_extract_constructor_dependencies()` reads `type_annotation` on constructor parameters | -| Variable/export handling | Arrow functions from `const`/`let` handled inline in traversal | Dedicated entity extractors: `_extract_lexical_declaration_entity()`, `_extract_variable_declaration_entity()`, `_extract_export_statement_entity()` | -| Built-in type filtering | `_is_builtin_type_js()` (broad DOM + JS globals list) | `_is_builtin_type()` (primitive TS types only) + `_is_builtin_function()` (currently empty) | - -## Data Flow: From Source File to Dependency Graph - -```mermaid -sequenceDiagram - participant DP as "DependencyParser" - participant JS as "TreeSitterJSAnalyzer" - participant TS as "TreeSitterTSAnalyzer" - participant DGB as "DependencyGraphBuilder" - - DP->>JS: analyze_javascript_file_treesitter(path, content, repo_path) - JS->>JS: Parser.parse(content) - JS->>JS: _extract_functions() / _extract_call_relationships() - JS-->>DP: (nodes, call_relationships) - - DP->>TS: analyze_typescript_file_treesitter(path, content, repo_path) - TS->>TS: Parser.parse(content) - TS->>TS: _extract_all_entities() / _filter_top_level_declarations() / _extract_all_relationships() - TS-->>DP: (nodes, call_relationships) - - DP->>DGB: aggregated nodes + relationships - DGB->>DGB: resolve callee ids against repository Node index - DGB-->>DP: Repository dependency graph -``` - -## Integration Points - -- **Input contract**: Both analyzers are constructed with `(file_path, content, repo_path)` and expose `analyze()` plus `nodes: List[Node]` and `call_relationships: List[CallRelationship]`, matching the shared interface used across all analyzers in [Tree Sitter Analyzers](tree-sitter-analyzers.md) (see also [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), [PHP Analyzer](php_analyzer.md), [Python Analyzer](python_analyzer.md)). -- **Output models**: `Node` and `CallRelationship` are defined in the dependency analyzer models. -- **Consumers**: The `DependencyParser` and `DependencyGraphBuilder` in the dependency analyzer core's graph construction stage invoke these analyzers per-file and merge their output into the overall repository graph; the analysis pipeline then operates on the resulting graph for call-graph and repository-level analysis. -- **Module-level entry points**: `analyze_javascript_file_treesitter()` and `analyze_typescript_file_treesitter()` are the primary functions called by upstream dispatch logic; both wrap analyzer instantiation and `analyze()` in a try/except that returns empty lists on failure, ensuring a single malformed file does not abort the overall analysis run. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md deleted file mode 100644 index 0019a4d7..00000000 --- a/docs/reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md +++ /dev/null @@ -1,225 +0,0 @@ -# PHP Analyzer - -The PHP Analyzer module is a language-specific static analysis component of CodeWiki's dependency analysis pipeline. It parses PHP source files using `tree-sitter-php` to extract structural code entities (classes, interfaces, traits, enums, functions, and methods) and the relationships between them (inheritance, interface implementation, object instantiation, static calls, constructor property promotion, and namespace imports). Its output feeds directly into the shared dependency analyzer models (`Node`, `CallRelationship`) that the rest of the CodeWiki dependency graph pipeline consumes. - -This module is one of several per-language analyzers registered under the [Tree Sitter Analyzers](tree-sitter-analyzers.md) module, alongside analyzers for [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), [JavaScript/TypeScript Analyzers](javascript_typescript_analyzers.md), and [Python Analyzer](python_analyzer.md). - -## Purpose and Scope - -PHP has language features that make dependency extraction non-trivial compared to simpler languages: - -- **Namespaces and `use` statements**, including aliasing (`use App\User as U;`) and grouped imports (`use App\{User, Post};`) -- **Fully-qualified vs. relative class names** that must be resolved relative to the current namespace and import table -- **Multiple declaration kinds** in a single file: classes (including abstract classes), interfaces, traits, enums, free functions, and methods -- **Template files** (e.g. Blade `.blade.php`, `.phtml`) that mix PHP with HTML/markup and should typically be excluded from dependency analysis -- **PHP 8+ constructor property promotion**, which implicitly creates typed dependencies on constructor parameters - -The PHP Analyzer module addresses all of these concerns through two cooperating classes: - -- **`NamespaceResolver`** — resolves short/aliased PHP type names to fully-qualified names (FQNs) -- **`TreeSitterPHPAnalyzer`** — walks the tree-sitter AST for a PHP file, using the resolver to produce `Node` and `CallRelationship` objects - -A convenience function, `analyze_php_file`, wraps analyzer construction and returns the extracted `(nodes, call_relationships)` tuple, matching the calling convention expected by the broader dependency-analysis pipeline. - -## Architecture - -```mermaid -flowchart TD - Caller["Dependency Analysis Pipeline"] -->|"analyze_php_file(path, content, repo_path)"| Analyze["analyze_php_file()"] - Analyze --> Analyzer["TreeSitterPHPAnalyzer"] - Analyzer -->|"uses"| Resolver["NamespaceResolver"] - Analyzer -->|"parses with"| TreeSitter["tree_sitter_php / Parser"] - Analyzer -->|"produces"| Nodes["List of Node"] - Analyzer -->|"produces"| Rels["List of CallRelationship"] - Nodes --> Models["Dependency Analyzer Models"] - Rels --> Models -``` - -`Node` and `CallRelationship` are defined outside this module in the shared dependency analyzer models component group, so the PHP Analyzer module only produces instances of them — it does not own their schema. - -### Class Relationships - -```mermaid -classDiagram - class NamespaceResolver { - +string current_namespace - +dict use_map - +register_namespace(ns) - +register_use(fqn, alias) - +resolve(name) string - } - class TreeSitterPHPAnalyzer { - +Path file_path - +string content - +string repo_path - +List nodes - +List call_relationships - +NamespaceResolver namespace_resolver - +_analyze() - +_extract_namespace_info(node, depth) - +_extract_nodes(node, lines, depth, parent_class) - +_extract_relationships(node, depth) - +_is_template_file() bool - +_get_component_id(name, parent_class) string - } - class Node { - +string id - +string name - +string component_type - +string file_path - +Set depends_on - +string source_code - +int start_line - +int end_line - } - class CallRelationship { - +string caller - +string callee - +int call_line - +bool is_resolved - } - TreeSitterPHPAnalyzer --> NamespaceResolver : "delegates name resolution" - TreeSitterPHPAnalyzer --> Node : "creates" - TreeSitterPHPAnalyzer --> CallRelationship : "creates" -``` - -## Core Components - -### NamespaceResolver - -`NamespaceResolver` maintains the state needed to translate PHP type names as they appear in source code into fully-qualified names (FQNs) suitable for use as dependency identifiers. - -Internal state: -- `current_namespace: str` — the namespace declared for the file (set via `register_namespace`) -- `use_map: Dict[str, str]` — maps a local alias (or the short class name if no alias) to its FQN, populated via `register_use` - -Resolution logic in `resolve(name)`: -1. If `name` already starts with `\`, it is already fully qualified — the leading backslash is stripped and it is returned as-is. -2. If `name` matches an entry in `use_map` exactly, the mapped FQN is returned. -3. If the first segment of a dotted/qualified name matches an alias in `use_map`, the base is substituted and the remainder appended. -4. Otherwise, the current namespace is prepended to `name`. -5. If there is no current namespace and no match, the name is returned unchanged. - -This mirrors PHP's own name-resolution rules for `use App\Foo as Bar; new Bar();` style code, ensuring dependency edges point at real, de-aliased classes. - -### TreeSitterPHPAnalyzer - -`TreeSitterPHPAnalyzer` is the primary entry point for analyzing a single PHP file. It is constructed with the file path, its content, and (optionally) the repository root path used to compute relative/module paths. - -**Initialization flow:** - -```mermaid -sequenceDiagram - participant Caller as "Caller" - participant Analyzer as "TreeSitterPHPAnalyzer" - participant Parser as "tree_sitter Parser" - participant Resolver as "NamespaceResolver" - - Caller->>Analyzer: "__init__(file_path, content, repo_path)" - Analyzer->>Analyzer: "_is_template_file()" - alt is template file - Analyzer-->>Caller: "return (skip analysis)" - else not a template - Analyzer->>Parser: "parse(content)" - Parser-->>Analyzer: "AST root node" - Analyzer->>Analyzer: "_extract_namespace_info(root)" - Analyzer->>Resolver: "register_namespace() / register_use()" - Analyzer->>Analyzer: "_extract_nodes(root, lines)" - Analyzer->>Analyzer: "_extract_relationships(root)" - Analyzer->>Resolver: "resolve(type_name)" - Analyzer-->>Caller: "nodes, call_relationships populated" - end -``` - -#### Template File Skipping - -Before any parsing occurs, `_is_template_file()` checks the file path against: -- Extension patterns: `.blade.php`, `.phtml`, `.twig.php` -- Directory patterns: `views`, `templates`, `resources/views` - -If matched, analysis is skipped entirely (`nodes` and `call_relationships` remain empty), avoiding noisy or irrelevant dependency data from view/template files. - -#### Three-Pass Analysis - -Once a file is confirmed to be analyzable PHP, `_analyze()` parses it with the `tree_sitter_php` grammar and runs three sequential passes over the AST: - -1. **Namespace/use extraction** (`_extract_namespace_info`) — walks the whole tree looking for `namespace_definition` and `namespace_use_declaration` nodes, registering the namespace and populating the `NamespaceResolver`'s `use_map` (including grouped `use App\{User, Post};` syntax via `_extract_use_statement`). -2. **Node extraction** (`_extract_nodes`) — walks the tree again, this time creating a `Node` for each recognized declaration type: - - `class_declaration` → `"class"` or `"abstract class"` (detected via `abstract_modifier`) - - `interface_declaration` → `"interface"` - - `trait_declaration` → `"trait"` - - `enum_declaration` → `"enum"` - - `function_definition` → `"function"` - - `method_declaration` → `"method"` (name is composed as `ContainingClass.methodName`) - - For classes/methods, extra metadata is captured: `parameters` (via `_extract_parameters`) and `base_classes` (via `_extract_base_classes`). Every node also gets a preceding PHPDoc comment attached as its `docstring`, discovered by `_get_preceding_docstring`. -3. **Relationship extraction** (`_extract_relationships`) — a further tree walk that emits `CallRelationship` objects for: - - `namespace_use_declaration` → import-based dependency from the file module to the imported FQN - - `class_declaration` with a `base_clause` → inheritance (`extends`) - - `class_declaration`/`enum_declaration` with a `class_interface_clause` → interface implementation - - `object_creation_expression` (`new Foo()`) → instantiation dependency - - `scoped_call_expression` (`Foo::bar()`) → static-call dependency - - `property_promotion_parameter` (PHP 8 constructor promotion) → typed dependency on the promoted parameter's type - -All PHP primitive/built-in types (see `PHP_PRIMITIVES`, e.g. `string`, `int`, `self`, `static`, `Exception`, `DateTime`) are filtered out via `_is_primitive` so that dependency edges reflect meaningful application code rather than language or standard-library noise. - -A `MAX_RECURSION_DEPTH` guard (100) is applied to every recursive tree walk to protect against pathologically deep ASTs causing a Python `RecursionError`. - -#### Component ID Generation - -`_get_component_id(name, parent_class)` produces the identifier used as a `Node.id` / `CallRelationship.caller`/`callee` value. If a namespace was detected for the file, the ID is built as: - -```text -Namespace.With.Dots::ParentClass.memberName -``` - -Otherwise it falls back to a module path derived from the file's location relative to the repository root (via `_get_module_path`), using `::` to separate the module path from the local declaration name — consistent with the ID convention expected by the graph construction stage's `DependencyParser`, which splits on `::` when deriving module groupings. - -### analyze_php_file - -```python -def analyze_php_file(file_path: str, content: str, repo_path: str = None) -> Tuple[List[Node], List[CallRelationship]] -``` - -This module-level function is the simple functional interface used by callers that do not need direct access to the analyzer instance: it constructs a `TreeSitterPHPAnalyzer`, runs analysis as part of construction, and returns the resulting `nodes` and `call_relationships` lists. This mirrors the calling pattern used by sibling analyzers such as those in [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), and [Python Analyzer](python_analyzer.md), allowing the orchestrating graph construction parser to treat all language analyzers uniformly. - -## Data Flow: PHP File to Dependency Graph - -```mermaid -flowchart LR - File["PHP source file"] --> Analyzer["TreeSitterPHPAnalyzer"] - Analyzer -->|"nodes"| NodeList["List of Node objects
(classes, methods, functions, etc.)"] - Analyzer -->|"relationships"| RelList["List of CallRelationship objects
(extends, implements, new, static call)"] - NodeList --> Parser["DependencyParser
(Graph Construction)"] - RelList --> Parser - Parser --> Graph["Dependency Graph
(components + depends_on edges)"] -``` - -The `Node` objects produced here carry `depends_on` as an empty set by default; it is the surrounding graph construction logic (`DependencyParser`) that consumes the emitted `CallRelationship` list to populate each `Node.depends_on` set and assign final namespaced component IDs (FQDNs) across the whole repository — potentially spanning multiple languages. - -## Relationship Types Extracted - -| PHP Construct | AST Node Type | Relationship Produced | -|---|---|---| -| `use App\Foo;` | `namespace_use_declaration` | File-to-class import dependency | -| `class A extends B` | `class_declaration` + `base_clause` | `A` → `B` (inheritance) | -| `class A implements I` | `class_declaration`/`enum_declaration` + `class_interface_clause` | `A` → `I` (implementation) | -| `new Foo()` | `object_creation_expression` | Containing class/function → `Foo` (instantiation) | -| `Foo::bar()` | `scoped_call_expression` | Containing class/function → `Foo` (static call) | -| `public function __construct(private Foo $foo)` | `property_promotion_parameter` | Containing class → `Foo` (typed dependency) | - -Each relationship is emitted as `is_resolved=False`, meaning name resolution to a final graph-wide component ID is deferred to the shared graph-construction stage rather than being finalized within this module. - -## Integration Points - -- **Input model**: raw PHP file paths and content, plus an optional repository root, supplied by the orchestrating dependency parser as part of graph construction. -- **Output model**: `Node` and `CallRelationship` instances defined in the dependency analyzer models. -- **Sibling analyzers**: registered together under [Tree Sitter Analyzers](tree-sitter-analyzers.md) for other supported languages (C/C++/C#, Java, JavaScript/TypeScript, Python). -- **Downstream consumers**: the assembled dependency graph is used by the analysis pipeline (`AnalysisService`, `RepoAnalyzer`, `CallGraphAnalyzer`) as part of generating repository-wide documentation. - -## Design Notes - -- **Template exclusion is path/extension based**, not content based — a `.blade.php` file is skipped regardless of its actual PHP content, trading recall for precision and avoiding markup-heavy files that rarely represent meaningful code dependencies. -- **Namespace resolution happens in a dedicated pre-pass** before node/relationship extraction so that `use` aliases declared anywhere in the file are available when resolving type references later in the same file, regardless of declaration order. -- **Recursion depth protection** (`MAX_RECURSION_DEPTH = 100`) is applied uniformly across all three AST walks, with graceful degradation (a warning log and early return) rather than raising unhandled exceptions. -- **Primitive/built-in type filtering** is centralized in `PHP_PRIMITIVES`, keeping dependency graphs focused on user-defined types. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md deleted file mode 100644 index 5b917ddd..00000000 --- a/docs/reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md +++ /dev/null @@ -1,196 +0,0 @@ -# Python Analyzer - -The Python Analyzer module implements an AST-based static analysis engine for Python source files. It walks the Python `ast` (Abstract Syntax Tree) produced by the standard library to extract structural information — classes, functions, and their call relationships — and converts that information into the standardized `Node` and `CallRelationship` models used throughout the dependency analysis pipeline. - -This module is one of several language-specific analyzers under the [Tree Sitter Analyzers](tree-sitter-analyzers.md) family. While its siblings (C, C++, C#, Java, JavaScript/TypeScript, PHP) rely on the `tree-sitter` parsing library, the Python analyzer uniquely leverages Python's own built-in `ast` module, since CPython ships with a fully-featured native parser for its own syntax. - -## Purpose and Role in the System - -The Python Analyzer is invoked by the dependency analyzer core during repository analysis to process individual `.py` files. Its output — lists of `Node` and `CallRelationship` objects — feeds directly into the dependency analyzer models and ultimately into the graph construction stage handled by `DependencyGraphBuilder`. - -Within the broader system, the analysis pipeline works roughly as follows: - -1. The `RepoAnalyzer` (part of the dependency analyzer core) walks the repository file tree and identifies Python source files. -2. For each Python file, `PythonASTAnalyzer` is instantiated and its `analyze()` method is invoked. -3. The analyzer parses the file content into an AST, traverses it, and produces `Node` objects (representing classes and top-level functions) plus `CallRelationship` objects (representing call-graph edges between them). -4. These results are aggregated with the output of other language analyzers and passed to `DependencyGraphBuilder` to construct the overall dependency graph. - -```mermaid -flowchart TD - RepoAnalyzer["RepoAnalyzer"] -->|"reads .py file"| PyFile["Python Source File"] - RepoAnalyzer -->|"instantiate"| PythonASTAnalyzer["PythonASTAnalyzer"] - PyFile -->|"content"| PythonASTAnalyzer - PythonASTAnalyzer -->|"ast.parse()"| AST["Python AST"] - AST -->|"NodeVisitor traversal"| PythonASTAnalyzer - PythonASTAnalyzer -->|"produces"| Nodes["List of Node"] - PythonASTAnalyzer -->|"produces"| Relationships["List of CallRelationship"] - Nodes --> DependencyGraphBuilder["DependencyGraphBuilder"] - Relationships --> DependencyGraphBuilder - DependencyGraphBuilder --> Graph["Dependency Graph"] -``` - -## Core Component - -### `PythonASTAnalyzer` - -`PythonASTAnalyzer` extends `ast.NodeVisitor` from the Python standard library, using the visitor pattern to walk the syntax tree and collect structural information as it goes. - -**Constructor parameters:** - -| Parameter | Type | Description | -|---|---|---| -| `file_path` | `str` | Absolute or relative path to the Python file being analyzed | -| `content` | `str` | Raw source code content of the file | -| `repo_path` | `Optional[str]` | Repository root path, used to compute paths relative to the repo for FQDN generation | - -**Key internal state:** - -- `nodes: List[Node]` — accumulates discovered classes and top-level functions -- `call_relationships: List[CallRelationship]` — accumulates discovered call edges -- `current_class_name` / `current_function_name` — tracks the visitor's current traversal context so calls can be attributed to the correct enclosing scope -- `top_level_nodes: dict` — maps simple names to their `Node` objects, used to resolve whether a given call target is a locally-defined class/function or an external reference - -## Component ID Generation (FQDN Format) - -A central responsibility of the analyzer is producing globally consistent, fully-qualified component IDs so that call relationships can be resolved across the whole dependency graph — not just within a single file. IDs follow the format: - -```text -:: -``` - -- The dotted module path is derived from the file's path relative to the repository root (via `_get_relative_path()` and `_get_module_path()`), stripping the `.py`/`.pyx` extension and converting path separators to dots. -- The component name is either a bare class/function name, or `ClassName.method_name` when inside a class body (tracked via `current_class_name`). - -This is implemented in `_get_module_path()` and `_get_component_id()`: - -```mermaid -flowchart LR - FilePath["file_path"] -->|"os.path.relpath"| RelPath["relative_path"] - RelPath -->|"strip .py/.pyx, replace separators with dots"| ModulePath["module.path"] - ModulePath -->|"+ '::' + ComponentName"| ComponentID["component_id"] -``` - -## AST Traversal and Extraction Logic - -### Class Extraction — `visit_ClassDef` - -When the visitor encounters a `ClassDef` node: - -1. Base classes are extracted via `_extract_base_class_name`, which handles both simple names (`class Foo(Bar)`) and dotted attribute access (`class Foo(module.Bar)`). -2. A `Node` is constructed with `component_type="class"`, capturing source code slice, docstring, line ranges, and base class names. -3. The class is registered into `top_level_nodes` for later call resolution. -4. If any base class is itself a known top-level node in the same file, an `is_resolved=True` `CallRelationship` is recorded representing the inheritance edge. -5. `current_class_name` is set before recursing into the class body (via `generic_visit`) and cleared afterward, so nested calls/methods are correctly attributed. - -### Function Extraction — `visit_FunctionDef` / `visit_AsyncFunctionDef` - -Both regular and `async def` functions are routed through `_process_function_node`: - -1. Only **top-level** functions (i.e., `current_class_name` is `None`) are added as standalone `Node` objects with `component_type="function"`. Methods defined inside classes are intentionally excluded from being top-level nodes (they are traversed for call extraction, but not separately registered as class members in this analyzer version). -2. `_should_include_function` filters out functions whose names start with `_test_` (helper filter to avoid polluting the graph with certain test scaffolding functions). -3. `current_function_name` is tracked during the recursive visit so that calls inside the function body can be attributed to it. - -### Call Relationship Extraction — `visit_Call` - -For every function/method call expression encountered during traversal: - -1. `_get_call_name` resolves the callee's name from the AST call target, handling: - - Simple name calls: `foo()` - - Attribute calls: `obj.method()` → resolved as `obj.method` - - Chained attributes: `a.b.method()` - - A hardcoded set of Python builtins (`print`, `len`, `isinstance`, etc.) is filtered out to avoid noise from standard library calls. -2. The caller ID is determined by the current traversal context — either the enclosing class or the enclosing top-level function. -3. If the resolved call name matches an entry in `top_level_nodes`, the relationship is marked `is_resolved=True` and the callee ID is fully qualified with the module path. Otherwise, the raw call name is stored as an unresolved reference (to potentially be resolved later at the cross-file graph-building stage). - -```mermaid -sequenceDiagram - participant Visitor as PythonASTAnalyzer - participant AST as ast.NodeVisitor - participant Nodes as top_level_nodes - participant Rels as call_relationships - - AST->>Visitor: visit_ClassDef(node) - Visitor->>Visitor: extract base_classes - Visitor->>Nodes: register class Node - Visitor->>Rels: append inheritance CallRelationship (if resolved) - Visitor->>Visitor: set current_class_name - Visitor->>AST: generic_visit(node) - AST->>Visitor: visit_Call(node) [inside class/function body] - Visitor->>Visitor: _get_call_name(node.func) - Visitor->>Nodes: lookup call_name - Visitor->>Rels: append CallRelationship (resolved or unresolved) - Visitor->>Visitor: clear current_class_name -``` - -## Data Model Alignment - -The `Node` and `CallRelationship` objects produced by this analyzer are defined in the dependency analyzer models. The Python Analyzer populates the following key fields: - -**`Node` fields populated:** - -| Field | Description | -|---|---| -| `id` / `component_id` | FQDN-format identifier (`module.path::Name`) | -| `name` | Simple class or function name | -| `component_type` | `"class"` or `"function"` | -| `file_path` / `relative_path` | Absolute and repo-relative file locations | -| `source_code` | Exact source lines spanning the node's definition | -| `start_line` / `end_line` | Line range from the AST node | -| `has_docstring` / `docstring` | Extracted via `ast.get_docstring()` | -| `parameters` | Argument names, for functions only | -| `node_type` | Mirrors `component_type` (`"class"` / `"function"`) | -| `base_classes` | List of resolved base class names, for classes only | -| `display_name` | Human-readable label, e.g. `"class Foo"` or `"function bar"` | - -**`CallRelationship` fields populated:** - -| Field | Description | -|---|---| -| `caller` | FQDN of the enclosing class or function | -| `callee` | FQDN (if resolved) or raw call name (if unresolved) | -| `call_line` | Line number of the call expression | -| `is_resolved` | `True` if the callee matches a locally-known top-level node | - -## Error Handling and Robustness - -The `analyze()` method wraps parsing and traversal in exception handling to ensure a single malformed file does not halt the entire repository analysis: - -- **`SyntaxError`**: Caught and logged as a warning when `ast.parse()` fails on invalid Python syntax; analysis for that file is skipped gracefully. -- **General exceptions**: Caught and logged as errors with full traceback (`exc_info=True`), keeping the pipeline resilient to unexpected AST edge cases. -- **`SyntaxWarning` suppression**: Parsing is wrapped in `warnings.catch_warnings()` to silence `SyntaxWarning`s (e.g., from invalid escape sequences in string/regex literals within analyzed source files), preventing noisy log output during large-scale repository scans. - -```mermaid -flowchart TD - Start["analyze() called"] --> Parse["ast.parse(content)"] - Parse -->|"success"| Visit["self.visit(tree)"] - Parse -->|"SyntaxError"| WarnLog["log warning, skip file"] - Visit -->|"success"| Done["nodes + call_relationships populated"] - Visit -->|"unexpected Exception"| ErrLog["log error with traceback"] -``` - -## Public Entry Point - -The module exposes a convenience function for one-shot analysis without manually managing the analyzer instance: - -```python -def analyze_python_file( - file_path: str, content: str, repo_path: Optional[str] = None -) -> Tuple[List[Node], List[CallRelationship]]: - ... -``` - -This function instantiates `PythonASTAnalyzer`, calls `.analyze()`, and returns the `(nodes, call_relationships)` tuple directly — the typical integration point used by callers such as `RepoAnalyzer` in the dependency analyzer core. - -## Design Notes and Limitations - -- **Top-level focus**: Only classes and top-level (module-scope) functions are registered as `Node` objects. Methods defined within classes are traversed for call-extraction purposes but are not independently registered as separate top-level nodes in this analyzer's current implementation — call attribution for method bodies is scoped to the enclosing class. -- **Builtin filtering**: A static allowlist of common Python builtins is excluded from call relationship extraction to reduce graph noise; this list is not exhaustive and may need periodic updates as usage patterns evolve. -- **Local-file resolution only**: Call resolution (`is_resolved`) is determined using only the nodes discovered within the same file (via `top_level_nodes`). Cross-file/cross-module call resolution is deferred to later stages of the pipeline, such as `DependencyGraphBuilder` in the dependency analyzer core. -- **Test function filtering**: Functions with names starting with `_test_` are excluded via `_should_include_function`, a convention-based filter to reduce noise from certain test helper patterns. - -## Related Modules - -- [Tree Sitter Analyzers](tree-sitter-analyzers.md) — parent grouping of all language-specific analyzers, including the [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), [JavaScript/TypeScript Analyzers](javascript_typescript_analyzers.md), and [PHP Analyzer](php_analyzer.md), which follow an analogous extraction pattern using `tree-sitter` grammars instead of Python's native `ast` module. -- Dependency Analyzer Core — hosts `RepoAnalyzer`, `AnalysisService`, `CallGraphAnalyzer`, and `DependencyGraphBuilder`, which orchestrate invocation of this analyzer and consume its output. -- Dependency Analyzer Models — defines the `Node` and `CallRelationship` Pydantic models that structure this analyzer's output. - diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md deleted file mode 100644 index 632f68f5..00000000 --- a/docs/reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md +++ /dev/null @@ -1,115 +0,0 @@ -# Tree Sitter Analyzers - -## Purpose - -The Tree Sitter Analyzers module is the multi-language source-code parsing layer of CodeWiki's dependency analysis pipeline. For every supported programming language, it provides a dedicated analyzer that reads a single source file and produces two things: - -1. **Nodes** — structural entities such as classes, functions, methods, interfaces, structs, enums, and global variables. -2. **Call Relationships** — edges describing how those entities depend on or interact with one another (function calls, inheritance, interface implementation, object creation, type usage, and more). - -These `Node` and `CallRelationship` objects are the common currency of the entire dependency analysis system. They are defined by the shared data models used across the backend and are consumed by the graph-construction stage that assembles the full repository dependency graph and the call-graph analyzer used for clustering and documentation generation. - -Because every analyzer emits data in the exact same shape, the rest of the pipeline (dependency graph building, call-graph resolution, clustering, and documentation generation) can operate on any supported language without needing to know the language-specific parsing details. - -## Supported Languages - -| Language | Analyzer | Parsing Technology | -|---|---|---| -| C | `TreeSitterCAnalyzer` | tree-sitter (`tree_sitter_c`) | -| C++ | `TreeSitterCppAnalyzer` | tree-sitter (`tree_sitter_cpp`) | -| C# | `TreeSitterCSharpAnalyzer` | tree-sitter (`tree_sitter_c_sharp`) | -| Java | `TreeSitterJavaAnalyzer` | tree-sitter (`tree_sitter_java`) | -| JavaScript | `TreeSitterJSAnalyzer` | tree-sitter (`tree_sitter_javascript`) | -| TypeScript | `TreeSitterTSAnalyzer` | tree-sitter (`tree_sitter_typescript`) | -| PHP | `TreeSitterPHPAnalyzer` | tree-sitter (`tree_sitter_php`) | -| Python | `PythonASTAnalyzer` | Python's built-in `ast` module | - -All analyzers except the Python one are built on top of [tree-sitter](https://tree-sitter.github.io/tree-sitter/) grammars. The Python analyzer instead uses Python's native `ast` module since a fully-featured, zero-dependency parser is already available in the standard library. - -## Architecture - -Every analyzer follows the same lifecycle: - -1. **Instantiate** with `(file_path, content, repo_path)`. -2. **Parse** the file content into a syntax tree (tree-sitter grammar or Python `ast`). -3. **Extract nodes** — walk the tree once (or twice, for languages that need entity pre-collection) to identify top-level declarations and record them as `Node` objects. -4. **Extract relationships** — walk the tree again to detect calls, inheritance, instantiation, and type usage, producing `CallRelationship` objects that reference the nodes discovered in step 3. -5. **Expose results** via `self.nodes` and `self.call_relationships`, typically also through a module-level `analyze__file(...)` convenience function that returns a `(nodes, relationships)` tuple. - -```mermaid -flowchart TD - Caller["Dependency Parser"] -->|"dispatches by file extension"| Router["Language Router"] - Router -->|".c / .h"| CAnalyzer["TreeSitterCAnalyzer"] - Router -->|".cpp / .hpp / .cc"| CppAnalyzer["TreeSitterCppAnalyzer"] - Router -->|".cs"| CSharpAnalyzer["TreeSitterCSharpAnalyzer"] - Router -->|".java"| JavaAnalyzer["TreeSitterJavaAnalyzer"] - Router -->|".js / .jsx / .mjs"| JSAnalyzer["TreeSitterJSAnalyzer"] - Router -->|".ts / .tsx"| TSAnalyzer["TreeSitterTSAnalyzer"] - Router -->|".php"| PHPAnalyzer["TreeSitterPHPAnalyzer"] - Router -->|".py"| PyAnalyzer["PythonASTAnalyzer"] - - CAnalyzer --> Output["Node and CallRelationship lists"] - CppAnalyzer --> Output - CSharpAnalyzer --> Output - JavaAnalyzer --> Output - JSAnalyzer --> Output - TSAnalyzer --> Output - PHPAnalyzer --> Output - PyAnalyzer --> Output - - Output --> Builder["Dependency Graph Builder"] - Output --> CallGraph["Call Graph Analyzer"] -``` - -### Component identity (FQDN scheme) - -Every analyzer computes a **component ID** for each extracted node so that entities can be uniquely referenced across the whole repository, following the pattern: - -```text -:: -::. -``` - -For example, a method `save` inside class `UserRepository` located at `src/data/user_repository.py` becomes: - -```text -src.data.user_repository::UserRepository.save -``` - -The module path is derived by taking the file path relative to the repository root, stripping the language-specific file extension, and replacing path separators with dots. This scheme is shared by all analyzers so that downstream consumers (dependency graph construction, clustering, and documentation generation) can treat component IDs uniformly regardless of source language. - -### Two-phase extraction - -Most analyzers separate **node extraction** from **relationship extraction** into two full tree traversals: - -- The first traversal builds a lookup table of top-level names (`top_level_nodes` / `all_entities`) mapping simple names to their `Node` objects. -- The second traversal re-walks the tree, and whenever it encounters a call, instantiation, inheritance clause, or type reference, it looks up the target name in that table to decide whether the relationship is *locally resolved* (target found in the same file) or *unresolved* (left for the cross-file resolution stage to handle later, similar to how the call-graph analyzer resolves cross-file callees). - -This is important because a single file cannot know about declarations in other files; each analyzer only guarantees **intra-file** resolution and marks everything else with `is_resolved=False`, leaving further resolution to the broader dependency analysis pipeline. - -## Sub-modules - -The eight language analyzers are grouped into five documentation sub-modules based on shared structural patterns and language families: - -- [C Family Analyzers](c_family_analyzers.md) — `TreeSitterCAnalyzer`, `TreeSitterCppAnalyzer`, `TreeSitterCSharpAnalyzer`. These three analyzers share a nearly identical single-pass extraction structure with growing structural complexity (functions only → classes/namespaces → interfaces/records/delegates). -- [Java Analyzer](java_analyzer.md) — `TreeSitterJavaAnalyzer`. Handles Java's rich type system: classes, interfaces, enums, records, annotations, and detailed method/field/object-creation relationship extraction, including local-variable type inference. -- [JavaScript & TypeScript Analyzers](javascript_typescript_analyzers.md) — `TreeSitterJSAnalyzer`, `TreeSitterTSAnalyzer`. Both analyze ECMAScript-family syntax (functions, arrow functions, classes, JSDoc/TypeScript type annotations) but differ in how strictly they validate "true top-level" declarations versus nested ones. -- [PHP Analyzer](php_analyzer.md) — `TreeSitterPHPAnalyzer` (with its `NamespaceResolver` helper). Unique among the analyzers in that it must resolve PHP namespaces and `use` statement aliases to fully qualified class names before building relationships. -- [Python Analyzer](python_analyzer.md) — `PythonASTAnalyzer`. The only analyzer built on Python's native `ast` module rather than tree-sitter, using the visitor pattern (`ast.NodeVisitor`) instead of manual tree traversal. - -## Node & Relationship Type Coverage - -| Language | Node Types Extracted | Relationship Types Extracted | -|---|---|---| -| C | function, struct, variable | calls, global-variable usage | -| C++ | class, struct, function, method, namespace, variable | calls, inherits, creates (`new`), uses (variable) | -| C# | class, abstract class, static class, interface, struct, enum, record, delegate | inherits (base list), property/field/parameter type usage | -| Java | class, abstract class, interface, enum, record, annotation, method | extends, implements, field type use, method invocation, object creation | -| JavaScript | class, abstract class, interface, function, generator function, arrow function, method | calls, inheritance, JSDoc-derived type dependencies | -| TypeScript | function, class, abstract class, interface, type alias, enum, variable, export statement | calls, `new` instantiation, member access, type annotations/arguments, inheritance/implements | -| PHP | class, abstract class, interface, trait, enum, function, method | `use` imports, extends, implements, `new`, static (`::`) calls, constructor property promotion | -| Python | class, function | calls, class inheritance | - -## Relationship to the Rest of the System - -These analyzers do not run standalone — they are invoked by the dependency parsing stage of the broader dependency analyzer, which selects the appropriate analyzer per file based on its extension, aggregates the resulting nodes and relationships across the whole repository, and hands them off to graph construction and call-graph resolution. The `Node` and `CallRelationship` data structures they produce are defined in the shared dependency analyzer models, ensuring a single consistent contract across every language-specific analyzer. diff --git a/docs/reference/architecture/cli-core/cli-core.md b/docs/reference/architecture/cli-core/cli-core.md deleted file mode 100644 index 794ff033..00000000 --- a/docs/reference/architecture/cli-core/cli-core.md +++ /dev/null @@ -1,103 +0,0 @@ -# Cli Core - -## Overview - -The Cli Core module is the command-line interface layer of CodeWiki. It is the entry point that end users interact with when running `codewiki` from a terminal: it manages persistent user configuration, wraps git repository operations, drives the backend documentation pipeline with progress reporting, renders a static HTML viewer for GitHub Pages, and provides shared logging/progress utilities used throughout the CLI experience. - -Rather than reimplementing documentation generation, Cli Core acts as a **thin orchestration and presentation layer** on top of the backend engine (see the [Backend Core](backend-core.md) module). It translates CLI-specific concerns — credential storage, terminal progress bars, colored logging, git branch management, and static site generation — into calls against the backend's `DocumentationGenerator` and `Config` primitives. - -## Responsibilities - -- **Configuration persistence**: securely store LLM provider credentials (cluster/main/fallback API keys) in the OS keyring, and persist non-sensitive settings (models, base URLs, token limits, agent instructions) to `~/.codewiki/config.json`. -- **Job orchestration**: adapt the backend's async documentation pipeline into a CLI-friendly staged workflow (dependency analysis → module clustering → documentation generation → optional HTML generation → finalization) with verbose/non-verbose progress reporting. -- **Git integration**: detect repository state, verify a clean working tree, create timestamped documentation branches, commit generated docs, and compute GitHub PR/Pages URLs. -- **Static site generation**: render a self-contained `index.html` viewer for GitHub Pages by combining a template with the generated `module_tree.json` and `metadata.json`. -- **Job/result modeling**: define typed data models (`DocumentationJob`, `LLMConfig`, `JobStatistics`, `GenerationOptions`, `JobStatus`) that describe a documentation run's inputs, progress, and outputs, with JSON (de)serialization for the `metadata.json` artifact. -- **Terminal UX utilities**: colored logging and multi-stage progress tracking with ETA estimation, shared across all CLI commands. - -## Architecture - -```mermaid -flowchart TD - User["CLI User"] -->|"codewiki generate"| ConfigMgr["ConfigManager"] - ConfigMgr -->|"loads/saves"| ConfigFile[("~/.codewiki/config.json")] - ConfigMgr -->|"stores API keys"| Keyring[("OS Keyring")] - ConfigMgr -->|"provides"| Config["Configuration + AgentInstructions"] - - Config -->|"to_backend_config()"| CLIDocGen["CLIDocumentationGenerator"] - GitMgr["GitManager"] -->|"branch/commit info"| CLIDocGen - - CLIDocGen -->|"drives"| BackendGen["Backend DocumentationGenerator"] - CLIDocGen -->|"tracks progress via"| Progress["ProgressTracker"] - CLIDocGen -->|"builds"| Job["DocumentationJob"] - CLIDocGen -->|"optional"| HTMLGen["HTMLGenerator"] - - HTMLGen -->|"reads"| ModuleTree[("module_tree.json")] - HTMLGen -->|"reads"| Metadata[("metadata.json")] - HTMLGen -->|"writes"| IndexHTML[("index.html")] - - Job -->|"serializes to"| Metadata - - Logger["CLILogger"] -.->|"used by"| CLIDocGen - Logger -.->|"used by"| ConfigMgr - Logger -.->|"used by"| GitMgr - - BackendGen -->|"belongs to"| BackendCore["Backend Core module"] - - style BackendCore fill:#eee,stroke:#999,stroke-dasharray: 5 5 -``` - -At a high level, a CLI command (e.g. `generate`) loads user settings via `ConfigManager`, optionally inspects/manipulates the repository via `GitManager`, and then hands control to `CLIDocumentationGenerator`, which converts CLI configuration into a backend `Config` object and drives the [Backend Core](backend-core.md) pipeline stage by stage, reporting progress through `ProgressTracker`. The resulting `DocumentationJob` captures statistics and status, which are persisted as `metadata.json`. If HTML output is requested, `HTMLGenerator` renders a static viewer from the generated `module_tree.json` and `metadata.json`. - -## Sub-modules - -Cli Core is organized into the following functional areas: - -| Sub-module | Responsibility | -|---|---| -| [Generation](cli-core/generation/generation.md) | Adapts the backend documentation pipeline for CLI use with staged progress reporting and logging configuration. | -| [Configuration](cli-core/configuration/configuration.md) | Manages persistent CLI settings and secure API key storage via the OS keyring; defines the configuration data model. | -| [Job Models](cli-core/job_models/job_models.md) | Typed data models describing a documentation job's status, statistics, and LLM configuration, with JSON serialization. | -| [Git Integration](cli-core/git_integration/git_integration.md) | Wraps git operations needed for documentation branch workflows (clean-check, branch creation, commit, remote/PR URL detection). | -| [Html Generation](cli-core/html_generation/html_generation.md) | Renders a static, self-contained HTML documentation viewer for GitHub Pages. | -| [Utils](cli-core/utils/utils.md) | Shared terminal UX helpers: colored logging and multi-stage progress tracking with ETA. | - -### Generation - -The [Generation](cli-core/generation/generation.md) sub-module contains `CLIDocumentationGenerator`, the central adapter that bridges CLI configuration and the backend engine. It normalizes additional source paths, builds a backend `Config`, configures backend logging with colored output, and runs the five-stage pipeline (dependency analysis, module clustering, documentation generation, optional HTML generation, finalization), reporting progress and raising `APIError` on failure. - -### Configuration - -The [Configuration](cli-core/configuration/configuration.md) sub-module contains `ConfigManager` together with the `Configuration` and `AgentInstructions` data models. `ConfigManager` persists non-sensitive settings to `~/.codewiki/config.json` and stores per-provider API keys (cluster/main/fallback) in the OS keyring, falling back gracefully when the keyring is unavailable. `Configuration.to_backend_config()` is the bridge that converts persisted CLI settings into a backend `Config` instance for a specific run. - -### Job Models - -The [Job Models](cli-core/job_models/job_models.md) sub-module defines `DocumentationJob` and its supporting types (`JobStatus`, `JobStatistics`, `GenerationOptions`, `LLMConfig`). These models track a single documentation run's lifecycle (pending → running → completed/failed), record statistics such as files analyzed and leaf nodes, and support round-trip JSON serialization used for the `metadata.json` artifact. - -### Git Integration - -The [Git Integration](cli-core/git_integration/git_integration.md) sub-module contains `GitManager`, which wraps the `git` Python library to support the "create documentation branch and commit" workflow: checking for a clean working directory, generating timestamped branch names, committing generated docs, and deriving GitHub remote/PR URLs. - -### Html Generation - -The [Html Generation](cli-core/html_generation/html_generation.md) sub-module contains `HTMLGenerator`, which loads `module_tree.json` and `metadata.json` from a documentation output directory, populates a bundled HTML template, and writes a self-contained `index.html` suitable for GitHub Pages hosting. - -### Utils - -The [Utils](cli-core/utils/utils.md) sub-module contains `CLILogger`, `ProgressTracker`, and `ModuleProgressBar` — shared terminal presentation helpers used across the CLI for colored log output and multi-stage progress bars with ETA estimation. - -## Relationship to Other Modules - -- **Backend Core**: Cli Core does not implement documentation generation itself; it delegates to the backend's `DocumentationGenerator`, dependency analyzer, and LLM services. See the [Backend Core](backend-core.md) module for details on dependency analysis, clustering, and agent orchestration. -- **Config Core**: The backend-facing `Config` object that `Configuration.to_backend_config()` produces is defined in the shared configuration module. See [Config Core](config-core.md) for details on the runtime configuration surface consumed by the backend. - -## Error Handling - -Cli Core raises a small hierarchy of typed exceptions used to communicate failures to the CLI entry point with distinct exit codes: - -- `ConfigurationError` — raised by `ConfigManager` when configuration cannot be loaded/saved or the keychain is unavailable. -- `RepositoryError` — raised by `GitManager` when git operations fail (e.g., dirty working directory, invalid repository). -- `FileSystemError` — raised by `HTMLGenerator` and configuration I/O helpers when reading/writing files fails. -- `APIError` — raised by `CLIDocumentationGenerator` when a backend/LLM stage (dependency analysis, clustering, documentation generation) fails. - -Each exception type maps to a dedicated process exit code, allowing CLI scripts and CI pipelines to distinguish between configuration problems, repository issues, filesystem errors, and API failures. diff --git a/docs/reference/architecture/cli-core/cli-core/configuration/configuration.md b/docs/reference/architecture/cli-core/cli-core/configuration/configuration.md deleted file mode 100644 index 7ed25c0a..00000000 --- a/docs/reference/architecture/cli-core/cli-core/configuration/configuration.md +++ /dev/null @@ -1,249 +0,0 @@ -# Configuration - -The Configuration module provides the persistent settings layer for the CodeWiki CLI. It defines the data models that describe user preferences (LLM providers, model names, token limits, temperature settings, and documentation-generation instructions) and the manager responsible for reading and writing those settings safely to disk and to the operating system's secure credential store. - -This module is a child of [Cli Core](../../cli-core.md) and works alongside sibling modules such as the job/runtime models, generation pipeline, and utility helpers to support the CLI's end-to-end workflow. - -## Purpose and Scope - -The Configuration module answers three core questions for the CLI: - -1. **What settings does the user have configured?** — captured by the `Configuration` dataclass. -2. **How should the documentation agent behave for a given run?** — captured by `AgentInstructions`. -3. **How are these settings persisted, validated, and loaded securely?** — handled by `ConfigManager`. - -Sensitive values (API keys) are never written to plaintext configuration files. Instead, they are stored using the system keyring (macOS Keychain, Windows Credential Manager, or Linux Secret Service) and only non-sensitive settings are persisted to `~/.codewiki/config.json`. - -## Core Components - -| Component | Responsibility | -|---|---| -| `Configuration` | Dataclass representing all persistent CLI settings: model names, base URLs, API versions, token/temperature limits, and clustering parameters. | -| `AgentInstructions` | Dataclass representing optional, user-customizable instructions for the documentation agent (file filters, focus modules, doc type, free-form instructions). | -| `ConfigManager` | Orchestrates loading/saving `Configuration` to `~/.codewiki/config.json` and securely storing/retrieving API keys via keyring. | - -## Architecture Overview - -```mermaid -flowchart TD - subgraph ConfigModule["Configuration Module"] - CM["ConfigManager"] - Cfg["Configuration"] - AI["AgentInstructions"] - end - - FS["~/.codewiki/config.json"] - KR["System Keyring"] - Backend["Backend Config"] - - CM -->|"load() / save()"| FS - CM -->|"get/set API keys"| KR - CM -->|"holds"| Cfg - Cfg -->|"has one"| AI - Cfg -->|"to_backend_config()"| Backend -``` - -`Configuration` is a plain, serializable dataclass with no direct dependency on keyring or the filesystem — those concerns are owned exclusively by `ConfigManager`. This separation keeps the data model easy to test and reuse, while `ConfigManager` acts as the single access point for persistence. - -## Data Model: Configuration - -`Configuration` captures all settings needed to drive a documentation-generation run, organized into three provider "roles": - -- **cluster** — model used for module clustering/decomposition -- **main** — primary model used for documentation generation -- **fallback** — fallback model used when the main model fails or is rate-limited - -For each role, the model tracks: -- Model name (`*_model`) -- Base URL (`*_base_url`, optional — for OpenAI-compatible or self-hosted endpoints) -- API version (`*_api_version`, optional) -- Max tokens (`*_max_tokens`) -- Temperature (`*_temperature`) and whether temperature is supported (`*_temperature_supported`) -- Max-token parameter field name (`*_max_token_field`, e.g., `"max_tokens"` vs. provider-specific names) - -In addition to per-provider settings, `Configuration` holds shared clustering parameters (`max_token_per_module`, `max_token_per_leaf_module`, `max_depth`) and a `default_output` directory. API keys (`cluster_api_key`, `main_api_key`, `fallback_api_key`) exist on the dataclass only as **runtime-only fields** — they are never serialized by `to_dict()` and are populated separately from the keyring by `ConfigManager`. - -```mermaid -classDiagram - class Configuration { - +str main_model - +str cluster_model - +str fallback_model - +str default_output - +str cluster_base_url - +str main_base_url - +str fallback_base_url - +int cluster_max_tokens - +int main_max_tokens - +int fallback_max_tokens - +float cluster_temperature - +float main_temperature - +float fallback_temperature - +bool cluster_temperature_supported - +bool main_temperature_supported - +bool fallback_temperature_supported - +int max_token_per_module - +int max_token_per_leaf_module - +int max_depth - +AgentInstructions agent_instructions - +validate() - +to_dict() dict - +from_dict(data) Configuration - +is_complete() bool - +to_backend_config(...) Config - } - - class AgentInstructions { - +List~str~ include_patterns - +List~str~ exclude_patterns - +List~str~ focus_modules - +str doc_type - +str custom_instructions - +to_dict() dict - +from_dict(data) AgentInstructions - +is_empty() bool - +get_prompt_addition() str - } - - Configuration "1" *-- "1" AgentInstructions : agent_instructions -``` - -### Validation - -`Configuration.validate()` performs field-level checks: -- Base URLs (when set) are validated via `validate_url`. -- Model names for all three roles are validated via `validate_model_name`. - -`from_dict()` performs defensive type coercion when loading from JSON (which may contain string-typed numbers/booleans due to manual editing), converting and range-checking integers (e.g., token limits), floats (temperature, bounded `0.0`–`2.0`), and booleans, raising `ValueError` on invalid input. - -### Serialization Rules - -- `to_dict()` only emits optional fields (`*_base_url`, `*_api_version`, `agent_instructions`) when they are set/non-empty, keeping the persisted JSON minimal. -- API keys are **excluded** from `to_dict()` entirely — they are never written to `config.json`. - -## Data Model: AgentInstructions - -`AgentInstructions` lets users customize how the documentation agent analyzes a repository and generates content: - -- `include_patterns` / `exclude_patterns` — glob-style file filters (e.g., `["*.cs"]`, `["*Tests*"]`) -- `focus_modules` — modules that should receive more detailed documentation -- `doc_type` — a preset documentation style (`api`, `architecture`, `user-guide`, `developer`) or a free-form type -- `custom_instructions` — arbitrary additional guidance passed to the LLM - -`get_prompt_addition()` translates these fields into a natural-language instruction block that is merged into the agent's prompt at generation time. `is_empty()` allows callers to distinguish "no customization" from "customization with all-default values." - -Instructions can be set persistently (stored in `config.json` as part of `Configuration`) or supplied at runtime for a single job; `Configuration.to_backend_config()` merges the two, with runtime instructions taking precedence field-by-field. - -## ConfigManager: Persistence and Secure Storage - -`ConfigManager` is the sole component responsible for reading and writing configuration state. It manages two independent storage backends: - -1. **JSON file** (`~/.codewiki/config.json`) — non-sensitive settings, versioned with a `CONFIG_VERSION` marker for future migrations. -2. **System keyring** (via the `keyring` library) — the three provider API keys (`cluster_api_key`, `main_api_key`, `fallback_api_key`), stored under a shared `codewiki` service name with distinct account identifiers. - -```mermaid -classDiagram - class ConfigManager { - -Optional~str~ _cluster_api_key - -Optional~str~ _main_api_key - -Optional~str~ _fallback_api_key - -Optional~Configuration~ _config - -bool _keyring_available - +load() bool - +save(...) void - +get_cluster_api_key() Optional~str~ - +get_main_api_key() Optional~str~ - +get_fallback_api_key() Optional~str~ - +get_config() Optional~Configuration~ - +is_configured() bool - +delete_api_keys() void - +clear() void - +keyring_available bool - +config_file_path Path - } - ConfigManager --> Configuration : loads/saves -``` - -### Load Flow - -```mermaid -sequenceDiagram - participant Caller - participant CM as ConfigManager - participant FS as "config.json" - participant KR as "System Keyring" - - Caller->>CM: load() - CM->>FS: check exists / read - alt file missing - FS-->>CM: not found - CM-->>Caller: False - else file present - FS-->>CM: JSON content - CM->>CM: Configuration.from_dict(data) - CM->>KR: get_password(cluster_api_key) - CM->>KR: get_password(main_api_key) - CM->>KR: get_password(fallback_api_key) - KR-->>CM: key values (or None) - CM-->>Caller: True - end -``` - -### Save Flow - -```mermaid -sequenceDiagram - participant Caller - participant CM as ConfigManager - participant FS as "config.json" - participant KR as "System Keyring" - - Caller->>CM: save(fields..., api_keys...) - CM->>FS: ensure_directory(CONFIG_DIR) - alt no in-memory config - CM->>CM: load() existing or create default Configuration - end - CM->>CM: apply provided field updates - CM->>CM: Configuration.validate() (if models set) - CM->>KR: set_password(...) for each provided API key - CM->>FS: write JSON (version + Configuration.to_dict()) - CM-->>Caller: done (or raises ConfigurationError) -``` - -Key behaviors: -- **Partial updates**: `save()` accepts every field as an optional keyword argument; only provided values overwrite the in-memory `Configuration`, allowing incremental configuration (e.g., `codewiki configure` sub-commands that set one field at a time). -- **Validation gate**: full validation (`Configuration.validate()`) only runs once `main_model` and `cluster_model` are both set, avoiding premature failures during multi-step setup. -- **Keyring failures surface as `ConfigurationError`**, with an actionable message when the OS keychain is unavailable or misconfigured. -- **`is_configured()`** combines two checks: all three API keys must be retrievable from keyring, and `Configuration.is_complete()` must be true (all three model names set). -- **`clear()`** performs a full reset — deleting API keys from keyring and removing `config.json` — used by commands like `codewiki configure --reset`. - -## Bridging to Runtime Execution: to_backend_config - -`Configuration.to_backend_config()` is the seam between this module's persistent settings and the runtime configuration consumed by the documentation-generation backend. It: - -1. Fetches any missing API keys from the keyring via a fresh `ConfigManager` instance (if not explicitly passed in). -2. Merges `runtime_instructions` (per-invocation `AgentInstructions`) over the persisted `agent_instructions`, with runtime values taking precedence field-by-field. -3. Constructs and returns a backend `Config` object (`Config.from_cli(...)`) populated with all model, token, temperature, and clustering settings, ready to drive a documentation job. - -```mermaid -flowchart LR - A["CLI Command"] --> B["ConfigManager.load()"] - B --> C["Configuration"] - C --> D["Configuration.to_backend_config()"] - D --> E["ConfigManager (keyring lookup for missing keys)"] - D --> F["Merge AgentInstructions (runtime over persisted)"] - D --> G["Config.from_cli(...)"] - G --> H["Backend Config"] -``` - -This bridging pattern keeps the CLI's persistent, user-facing settings model (`Configuration`) decoupled from the backend's execution-time `Config` model. The backend `Config` class itself is documented in the [Config Core](../../config-core.md) module. - -## Relationship to Other CLI Modules - -- **[Job Models](../job_models/job_models.md)** — represents the runtime state of an in-progress documentation job (`DocumentationJob`, `JobStatus`, `LLMConfig`, `GenerationOptions`, `JobStatistics`). Where `Configuration` describes *persistent user preferences*, the job models describe the *live execution* of a single generation run, often derived from a `Configuration` via `to_backend_config()`. -- **[Generation](../generation/generation.md)** — the `CLIDocumentationGenerator` adapter consumes a resolved `Configuration`/backend `Config` to drive the documentation pipeline. -- **[Utils](../utils/utils.md)** — provides shared filesystem (`ensure_directory`, `safe_read`, `safe_write`) and error types (`ConfigurationError`, `FileSystemError`) used internally by `ConfigManager`. - -## Summary - -The Configuration module is the trust boundary for user secrets and the single source of truth for persistent CLI preferences. `Configuration` and `AgentInstructions` define what can be configured; `ConfigManager` defines how those settings are safely loaded, validated, updated, and translated into a form the backend documentation pipeline can execute against. diff --git a/docs/reference/architecture/cli-core/cli-core/generation/generation.md b/docs/reference/architecture/cli-core/cli-core/generation/generation.md deleted file mode 100644 index 17120aeb..00000000 --- a/docs/reference/architecture/cli-core/cli-core/generation/generation.md +++ /dev/null @@ -1,167 +0,0 @@ -# Generation - -The Generation module is the central orchestration layer of the CodeWiki CLI. It bridges the command-line interface with the backend documentation engine, coordinating the full end-to-end pipeline that turns a source code repository into a structured, hierarchical documentation set — including dependency analysis, LLM-driven module clustering, documentation generation, optional Mermaid diagram extraction, and optional HTML output. - -Its single core component, `CLIDocumentationGenerator`, acts as an **adapter**: it wraps the backend's `DocumentationGenerator` (from [Backend Core](../../backend-core.md)) with CLI-specific concerns such as progress reporting, colored logging, verbose diagnostics, and job lifecycle tracking via the [Job Models](../job_models/job_models.md) module. - -## Purpose and Role in the System - -As a child of [Cli Core](../../cli-core.md), the Generation module is invoked whenever a user runs a documentation generation command. It does not implement dependency analysis, clustering, or documentation writing itself — those responsibilities belong to `backend-core`. Instead, Generation is responsible for: - -1. **Translating CLI configuration** into the backend's `Config` object (model selection, API keys, base URLs, token limits, clustering depth, agent instructions, additional source paths). -2. **Orchestrating the multi-stage pipeline** (dependency analysis → module clustering → documentation generation → optional HTML generation → finalization) with a five-stage `ProgressTracker`. -3. **Tracking job state** using a `DocumentationJob` model, recording statistics, generated files, and success/failure outcomes. -4. **Handling synthetic module fallback** to avoid context-window overflows when the LLM clustering step returns an empty module tree. -5. **Optionally generating an HTML viewer** via [Html Generation](../html_generation/html_generation.md) and extracting Mermaid diagrams into a dedicated directory. - -## Architecture Overview - -```mermaid -flowchart TD - CLI["CLI Command Entry Point"] --> CDG["CLIDocumentationGenerator"] - CDG --> PT["ProgressTracker (utils)"] - CDG --> Job["DocumentationJob (job_models)"] - CDG --> BC["BackendConfig (config-core)"] - CDG --> DG["DocumentationGenerator (backend-core)"] - DG --> GB["DependencyGraphBuilder"] - DG --> AO["AgentOrchestrator"] - CDG --> CM["cluster_modules (backend-core)"] - CDG --> HG["HTMLGenerator (html_generation)"] - CDG --> LOG["ColoredFormatter (backend-core logging)"] -``` - -## Core Component - -### CLIDocumentationGenerator - -`CLIDocumentationGenerator` (`codewiki/cli/adapters/doc_generator.py`) is instantiated with: - -- `repo_path` — path to the repository being documented -- `output_dir` — target directory for generated docs -- `config` — a dictionary of LLM/model configuration (models, API keys, base URLs, token limits, temperature settings, clustering parameters, agent instructions, additional source paths) -- `verbose` — whether to emit detailed diagnostic output -- `generate_html` — whether to produce an HTML viewer (`index.html`) -- `diagrams_dir` — optional separate output directory for extracted Mermaid diagrams - -On construction, it: -- Creates a `ProgressTracker` configured for 5 pipeline stages. -- Creates a `DocumentationJob` and populates its metadata (repository path/name, output directory, `LLMConfig`). -- Configures backend logging by attaching a `ColoredFormatter`-based handler to the `codewiki.src.be` logger namespace, switching verbosity between `INFO` (verbose) and `WARNING` (quiet) levels, and disabling propagation to avoid duplicate log lines. - -#### Key Responsibilities - -| Responsibility | Method | -|---|---| -| Build backend configuration from CLI config dict | `generate()` | -| Run the full async generation pipeline | `_run_backend_generation()` | -| Generate the HTML viewer | `_run_html_generation()` | -| Ensure job metadata file exists | `_finalize_job()` | -| Configure colored backend logging | `_configure_backend_logging()` | - -## Generation Pipeline - -The `generate()` method is the public entry point. It is synchronous from the caller's perspective but internally drives an `asyncio`-based backend pipeline. It returns a completed `DocumentationJob` (see [Job Models](../job_models/job_models.md)) or raises an `APIError` on failure. - -```mermaid -sequenceDiagram - participant Caller as "CLI Command" - participant CDG as "CLIDocumentationGenerator" - participant BC as "BackendConfig" - participant DG as "DocumentationGenerator" - participant CM as "cluster_modules" - participant HG as "HTMLGenerator" - participant Job as "DocumentationJob" - - Caller->>CDG: generate() - CDG->>Job: start() - CDG->>BC: Config.from_cli(...) - CDG->>DG: _run_backend_generation(backend_config) - DG->>DG: graph_builder.build_dependency_graph() - DG->>CM: cluster_modules(leaf_nodes, components, config) - Note over DG,CM: Synthetic module fallback if tree is empty - DG->>DG: generate_module_documentation(components, leaf_nodes) - opt diagrams_dir configured - DG->>DG: extract_and_save_mermaid_diagrams() - end - opt generate_html is true - CDG->>HG: generate(output_path, ...) - end - CDG->>CDG: _finalize_job() - CDG->>Job: complete() - CDG-->>Caller: DocumentationJob -``` - -### Stage 1 — Dependency Analysis - -Instantiates the backend `DocumentationGenerator` (which internally wires up `DependencyGraphBuilder` and `AgentOrchestrator` from [Backend Core](../../backend-core.md)) and calls `doc_generator.graph_builder.build_dependency_graph()`. The result is a map of `components` and a list of `leaf_nodes`. Statistics (`total_files_analyzed`, `leaf_nodes`) are recorded on the `DocumentationJob`. Failures are wrapped as `APIError("Dependency analysis failed: ...")`. - -### Stage 2 — Module Clustering - -Loads a cached `first_module_tree.json` if present, otherwise calls `cluster_modules(leaf_nodes, components, backend_config)` to invoke the LLM clustering model. The result is cached to disk (`first_module_tree_path`) and then persisted as the working module tree (`module_tree_path`). - -**Synthetic Module Patch**: If the module tree ends up empty despite having leaf nodes (to prevent an LLM "whole-repo" fallback that could exceed API context limits), the generator batches leaf nodes into synthetic modules of a configurable size (`CODEWIKI_MAX_FILES_PER_MODULE`, default `5`) and re-persists the tree. This safeguard applies even when loading from cache, closing a previously identified cache-bypass gap. - -The final module count is stored on the `DocumentationJob.module_count` field. - -### Stage 3 — Documentation Generation - -Calls `doc_generator.generate_module_documentation(components, leaf_nodes)`, which performs the topologically-ordered (leaf-first) generation of module documentation via the `AgentOrchestrator`. After generation: - -- `doc_generator.create_documentation_metadata(...)` writes `metadata.json`. -- Generated `.md` and `.json` files in the output directory are collected into `DocumentationJob.files_generated`. -- If `diagrams_dir` was configured, Mermaid diagrams embedded in generated markdown are extracted via `extract_and_save_mermaid_diagrams` and indexed with `create_diagrams_readme`. - -### Stage 4 — HTML Generation (Optional) - -If `generate_html=True`, `_run_html_generation()` uses the [Html Generation](../html_generation/html_generation.md) module's `HTMLGenerator` to detect repository info (name, URL, GitHub Pages URL) and render `index.html`, auto-loading the module tree and metadata from the output directory. The generated file is appended to `DocumentationJob.files_generated`. - -### Stage 5 — Finalization - -`_finalize_job()` verifies that `metadata.json` exists in the output directory; if the backend did not already write it, the generator writes the job's own JSON representation (`DocumentationJob.to_json()`) as a fallback. - -## Configuration Mapping - -`generate()` builds the backend `Config` object (see [Config Core](../../config-core.md)) via `Config.from_cli(...)`, translating the CLI's flat configuration dictionary into per-provider settings: - -```mermaid -flowchart LR - subgraph CLIConfig["CLI config dict"] - A1["main_model / cluster_model / fallback_model"] - A2["*_api_key"] - A3["*_base_url"] - A4["*_api_version"] - A5["*_max_tokens / *_temperature"] - A6["max_token_per_module / max_depth"] - A7["agent_instructions"] - A8["additional_paths"] - end - CLIConfig --> FromCLI["Config.from_cli()"] - FromCLI --> BackendConfig["Backend Config instance"] - BackendConfig --> DG2["DocumentationGenerator"] -``` - -Additional source paths supplied in the CLI config are normalized to absolute paths (resolved relative to `repo_path`) before being passed through as `additional_source_paths`. - -In verbose mode, the generator prints a detailed configuration summary (model names, base URLs, token limits, module settings, additional paths, and a preview of custom agent instructions) before kicking off Stage 1. - -## Progress Tracking and Logging - -The Generation module relies on utilities from the [Utils](../utils/utils.md) module: - -- **`ProgressTracker`**: manages a 5-stage weighted progress model (Dependency Analysis 40%, Module Clustering 20%, Documentation Generation 30%, HTML Generation 5%, Finalization 5%), providing `start_stage`, `update_stage`, `complete_stage`, elapsed-time formatting, and ETA estimation. -- **`CLILogger`**: complements verbose/non-verbose console output alongside `ProgressTracker`'s stage banners. - -Backend logs are captured by attaching a handler using `ColoredFormatter` (from `backend-core`'s Logging Config child module) directly to the `codewiki.src.be` logger, ensuring consistent colored output regardless of whether the CLI or backend emitted the log line, while preventing duplicate propagation to the root logger. - -## Error Handling - -All backend-facing calls (`build_dependency_graph`, `cluster_modules`, `generate_module_documentation`) are wrapped in `try`/`except` blocks that convert unexpected exceptions into `APIError` with contextual messages (e.g., `"Dependency analysis failed: ..."`). At the top level, `generate()` catches both `APIError` and generic `Exception`, marks the `DocumentationJob` as failed via `job.fail(str(e))`, and re-raises so the CLI layer can present the error to the user. - -## Relationship to Other Modules - -- **[Cli Core](../../cli-core.md)**: parent module; Generation is one of its functional children alongside [Configuration](../configuration/configuration.md), [Job Models](../job_models/job_models.md), [Git Integration](../git_integration/git_integration.md), [Html Generation](../html_generation/html_generation.md), and [Utils](../utils/utils.md). -- **[Job Models](../job_models/job_models.md)**: supplies `DocumentationJob`, `LLMConfig`, `JobStatus`, and `JobStatistics`, which Generation populates throughout the pipeline. -- **[Html Generation](../html_generation/html_generation.md)**: invoked optionally at Stage 4 to render the documentation as a browsable HTML site. -- **[Utils](../utils/utils.md)**: provides `ProgressTracker` for stage-based progress reporting. -- **[Backend Core](../../backend-core.md)**: supplies the actual documentation engine (`DocumentationGenerator`, `AgentOrchestrator`, `DependencyGraphBuilder`) that Generation orchestrates but does not reimplement. -- **[Config Core](../../config-core.md)**: supplies the `Config` class used to translate CLI settings into the backend's runtime configuration. diff --git a/docs/reference/architecture/cli-core/cli-core/git_integration/git_integration.md b/docs/reference/architecture/cli-core/cli-core/git_integration/git_integration.md deleted file mode 100644 index 41acdeb2..00000000 --- a/docs/reference/architecture/cli-core/cli-core/git_integration/git_integration.md +++ /dev/null @@ -1,147 +0,0 @@ -# Git Integration - -The Git Integration module provides the `GitManager` component, which encapsulates all git repository operations required by the CodeWiki CLI's documentation workflow. It is responsible for validating that a target directory is a git repository, inspecting working-directory cleanliness, creating dedicated documentation branches, committing generated documentation, and deriving useful metadata such as remote URLs, branch names, commit hashes, and GitHub pull-request links. - -This module is a child of the [Cli Core](../../cli-core.md) module and is used primarily by the CLI's generation workflow (see [Generation](../generation/generation.md)) to safely manage version control state before, during, and after documentation is generated. - -## Purpose and Scope - -When the CLI is run with branch-creation options (e.g. `--create-branch`), it must: - -1. Confirm the target path is inside a valid git repository. -2. Ensure there are no uncommitted changes that could be silently mixed into a documentation commit (unless the user explicitly forces the operation). -3. Create a uniquely named, timestamped branch dedicated to documentation output. -4. Stage and commit the generated documentation files. -5. Surface repository metadata (current branch, commit hash, remote URL) back to the CLI so it can print helpful summaries. -6. Build a ready-to-open GitHub pull-request URL if the remote is hosted on GitHub. - -All of this logic is centralized in the `GitManager` class, keeping git operations isolated from the rest of the CLI so that other components (progress reporting, configuration, documentation generation) do not need to know about the underlying git library (`GitPython`). - -## Core Component: GitManager - -`GitManager` wraps a `git.Repo` instance (from the `GitPython` library) and exposes a small, purpose-built API surface consumed by the CLI layer. - -### Initialization - -On construction, `GitManager` resolves the given `repo_path` to an absolute path and attempts to open it as a git repository using `git.Repo(repo_path, search_parent_directories=True)`. The `search_parent_directories=True` flag allows the manager to locate a `.git` directory even if `repo_path` points to a subdirectory of the repository (similar to how the native `git` command walks upward from the current directory). - -If no valid repository is found, a `RepositoryError` is raised with actionable guidance (e.g., suggesting `git init`). `RepositoryError` is a CLI-specific exception type that carries a dedicated exit code (`EXIT_REPOSITORY_ERROR`), allowing the CLI's top-level error handler to report failures consistently and set correct process exit statuses. - -### Key Responsibilities - -#### 1. Working Directory Cleanliness Check - -`check_clean_working_directory()` inspects the repository for uncommitted modifications or untracked files using `repo.is_dirty(untracked_files=True)`. When dirty, it builds a human-readable summary listing up to three modified and three untracked files (with a "... and N more" suffix for longer lists), returning a `(is_clean, status_message)` tuple. This check underpins the safety guarantee that documentation generation will not accidentally commit unrelated in-progress changes. - -#### 2. Documentation Branch Creation - -`create_documentation_branch(force: bool = False)` creates and checks out a new branch named `docs/codewiki-` (format `%Y%m%d-%H%M%S`). Before creating the branch it re-runs the cleanliness check unless `force=True`, raising a detailed `RepositoryError` with copy-pastable remediation commands (`git status`, `git add -A && git commit`, `git stash`) if the working directory is dirty. - -To guard against timestamp collisions, it checks existing branch names and appends an incrementing numeric suffix (`-1`, `-2`, ...) if a collision would otherwise occur. Branch creation and checkout are performed via `repo.create_head(...)` followed by `.checkout()`; any underlying `GitCommandError` is translated into a `RepositoryError`. - -#### 3. Committing Documentation - -`commit_documentation(docs_path: Path, message: Optional[str] = None)` stages the documentation output directory via `repo.index.add([str(docs_path)])` and commits it with either a caller-supplied message or the default `"Add generated documentation\n\nGenerated by CodeWiki CLI"`. It returns the resulting commit's `hexsha`. Failures during add/commit are wrapped in `RepositoryError`. - -#### 4. Repository Metadata Accessors - -- `get_remote_url(remote_name="origin")` — returns the URL of the named remote, or `None` if it does not exist. -- `get_current_branch()` — returns the active branch name, or the literal string `"HEAD"` when in a detached-HEAD state (caught via `TypeError` from `GitPython`). -- `get_commit_hash()` — returns the current `HEAD` commit's hex SHA. -- `branch_exists(branch_name)` — returns whether a branch with the given name already exists locally. - -#### 5. GitHub Pull-Request URL Construction - -`get_github_pr_url(branch_name)` derives a ready-to-use GitHub "compare" URL (`/compare/`) from the `origin` remote, but only when that remote points to `github.com`. It normalizes the remote URL by: -- Stripping a trailing slash and any `.git` suffix. -- Converting SSH-style remotes (`git@github.com:org/repo`) into HTTPS form (`https://github.com/org/repo`). - -If there is no remote or it is not a GitHub remote, it returns `None`, allowing the CLI to gracefully skip printing a PR link for non-GitHub repositories. - -## Error Handling - -All git-related failures surface as `RepositoryError` (defined outside this module in the CLI's shared error utilities), rather than leaking raw `GitCommandError` or `InvalidGitRepositoryError` exceptions from `GitPython`. This gives the CLI a single, well-known exception type to catch at the top level and translate into a clean error message and a dedicated process exit code, keeping git-library implementation details out of user-facing error handling. - -## Architecture - -```mermaid -classDiagram - class GitManager { - +Path repo_path - +Repo repo - +__init__(repo_path) - +check_clean_working_directory() Tuple - +create_documentation_branch(force) str - +commit_documentation(docs_path, message) str - +get_remote_url(remote_name) str - +get_current_branch() str - +get_commit_hash() str - +branch_exists(branch_name) bool - +get_github_pr_url(branch_name) str - } - class RepositoryError { - +__init__(message) - } - GitManager ..> RepositoryError : raises -``` - -## Interaction with the CLI Generation Workflow - -`GitManager` is typically instantiated by the CLI's generation flow when branch-based documentation workflows are requested. The sequence below illustrates a typical end-to-end interaction: verifying repository state, creating a branch, running documentation generation (handled by `CLIDocumentationGenerator` in the [Generation](../generation/generation.md) module), committing the results, and reporting a PR link. - -```mermaid -sequenceDiagram - participant CLI as "CLI Entry Point" - participant GM as "GitManager" - participant DocGen as "CLIDocumentationGenerator" - participant Repo as "git.Repo" - - CLI->>GM: "GitManager(repo_path)" - GM->>Repo: "git.Repo(repo_path)" - CLI->>GM: "check_clean_working_directory()" - GM-->>CLI: "(is_clean, status_message)" - alt "Working directory dirty and not forced" - GM-->>CLI: "raise RepositoryError" - else "Clean or forced" - CLI->>GM: "create_documentation_branch(force)" - GM->>Repo: "create_head(branch_name)" - GM->>Repo: "checkout()" - GM-->>CLI: "branch_name" - CLI->>DocGen: "generate()" - DocGen-->>CLI: "DocumentationJob" - CLI->>GM: "commit_documentation(docs_path, message)" - GM->>Repo: "index.add([docs_path])" - GM->>Repo: "index.commit(message)" - GM-->>CLI: "commit_hash" - CLI->>GM: "get_github_pr_url(branch_name)" - GM-->>CLI: "pr_url or None" - end -``` - -## Process Flow: Documentation Branch Creation - -```mermaid -flowchart TD - Start["create_documentation_branch(force)"] --> CheckForce{{force?}} - CheckForce -->|"false"| CheckClean["check_clean_working_directory()"] - CheckForce -->|"true"| GenName["Generate timestamped branch name"] - CheckClean --> IsClean{{"is_clean?"}} - IsClean -->|"false"| RaiseErr["raise RepositoryError with remediation steps"] - IsClean -->|"true"| GenName - GenName --> CheckExists{{"branch name exists?"}} - CheckExists -->|"true"| AppendCounter["Append incrementing counter suffix"] - AppendCounter --> CheckExists - CheckExists -->|"false"| CreateBranch["repo.create_head(branch_name)"] - CreateBranch --> Checkout["new_branch.checkout()"] - Checkout --> ReturnName["return branch_name"] -``` - -## Relationship to Other Modules - -- **[Cli Core](../../cli-core.md)** — the parent module; `GitManager` is one of the core services orchestrated alongside configuration management, job models, and progress tracking to deliver the full `codewiki generate` experience. -- **[Generation](../generation/generation.md)** — `CLIDocumentationGenerator` performs the actual documentation generation (dependency analysis, clustering, LLM-driven writing, optional HTML output). The CLI entry point coordinates `GitManager` and `CLIDocumentationGenerator` together: creating a branch before generation, and committing the resulting files afterward. -- **[Job Models](../job_models/job_models.md)** — while `GitManager` itself does not depend on `DocumentationJob`, the surrounding CLI workflow often records git-derived metadata (branch name, commit hash) alongside job statistics for reporting purposes. - -## Dependencies - -`GitManager` depends on the third-party `GitPython` library (imported as `git`) for all low-level repository interactions, including `git.Repo`, `git.InvalidGitRepositoryError`, and `git.exc.GitCommandError`. It also depends on the CLI's shared `RepositoryError` exception type for consistent error reporting across the CLI. diff --git a/docs/reference/architecture/cli-core/cli-core/html_generation/html_generation.md b/docs/reference/architecture/cli-core/cli-core/html_generation/html_generation.md deleted file mode 100644 index f11bd2a0..00000000 --- a/docs/reference/architecture/cli-core/cli-core/html_generation/html_generation.md +++ /dev/null @@ -1,131 +0,0 @@ -# Html Generation - -The Html Generation module produces a self-contained, static HTML documentation viewer suitable for GitHub Pages (or any static file host). It is the final, optional presentation layer of the CodeWiki CLI pipeline: after the [Generation](../generation/generation.md) stage has produced Markdown documentation files, a module tree, and metadata, the Html Generation module packages that output into a single browsable `index.html` file with embedded configuration, styles, and client-side rendering logic. - -## Purpose and Scope - -The `HTMLGenerator` class is the sole core component of this module. Its responsibilities are: - -- **Template loading** — reads a static HTML template (`viewer_template.html`) shipped with the package. -- **Data discovery** — auto-loads `module_tree.json` and `metadata.json` from a documentation output directory when explicit data is not supplied. -- **Placeholder substitution** — injects title, repository link, embedded JSON data (module tree, metadata, config), and an "info panel" HTML fragment into the template. -- **Repository introspection** — inspects a local git repository to derive a display name, remote URL, and a predicted GitHub Pages URL. -- **Atomic file output** — writes the final HTML file safely using the shared filesystem utilities. - -This module has no knowledge of *how* documentation content was produced; it only consumes the artifacts (`module_tree.json`, `metadata.json`, generated Markdown files) that the [Generation](../generation/generation.md) module and the backend documentation pipeline produce. This keeps Html Generation a pure "rendering/packaging" concern, decoupled from LLM orchestration, dependency analysis, and job tracking. - -## Core Component - -### HTMLGenerator - -`HTMLGenerator` (`codewiki/cli/html_generator.py`) encapsulates all HTML viewer generation logic. - -| Method | Responsibility | -|---|---| -| `__init__(template_dir)` | Resolves the template directory, defaulting to the package's `templates/github_pages` folder. | -| `load_module_tree(docs_dir)` | Reads `module_tree.json` from the docs directory; falls back to a minimal single-node structure if the file is missing. | -| `load_metadata(docs_dir)` | Reads `metadata.json`; returns `None` (non-critical) if missing or unparsable. | -| `generate(...)` | Orchestrates the full generation: auto-loads data, builds the info panel, computes paths/links, serializes JSON, performs placeholder substitution, and writes the output file. | -| `_build_info_content(metadata)` | Builds an HTML fragment (model name, generation timestamp, commit hash, component count, max depth) displayed in the viewer's info panel. | -| `_escape_html(text)` | Escapes HTML-sensitive characters to prevent malformed markup when embedding user/repo-derived strings (e.g., title). | -| `detect_repository_info(repo_path)` | Uses `GitPython` to read the repository name, normalize the remote URL (including `git@github.com:` SSH URLs), and compute the expected `https://.github.io//` Pages URL. | - -## Architecture - -```mermaid -flowchart TD - Docs["Documentation Output Directory"] --> ModuleTreeJson["module_tree.json"] - Docs --> MetadataJson["metadata.json"] - Template["viewer_template.html"] --> Generator["HTMLGenerator"] - ModuleTreeJson --> Generator - MetadataJson --> Generator - RepoPath["Repository Path"] --> DetectInfo["detect_repository_info()"] - DetectInfo --> Generator - Generator -->|"safe_write()"| IndexHtml["index.html"] - Generator -->|"on failure"| FSError["FileSystemError"] -``` - -### Dependencies - -- **Shared filesystem helpers** — `safe_read` / `safe_write` provide atomic, encoding-safe file I/O used to load the template/JSON files and write the final HTML output. See the [Utils](../utils/utils.md) module for other shared CLI utilities such as logging and progress tracking. -- **`FileSystemError`** — raised when the template file is missing or when reading/writing fails, allowing the CLI layer to surface a consistent error type. -- **`git` (GitPython)** — used only within `detect_repository_info` to introspect the local repository; failures are caught and silently ignored, degrading gracefully to a viewer without repository links. - -## Integration with the CLI Pipeline - -Html Generation is invoked as an optional, final stage by `CLIDocumentationGenerator` (from the [Generation](../generation/generation.md) module), which drives the overall CLI workflow: dependency analysis → module clustering → documentation generation → **HTML generation** → job finalization. - -```mermaid -sequenceDiagram - participant CLIGen as "CLIDocumentationGenerator" - participant HTMLGen as "HTMLGenerator" - participant FS as "safe_read / safe_write" - participant Git as "GitPython" - - CLIGen->>HTMLGen: HTMLGenerator() - CLIGen->>HTMLGen: detect_repository_info(repo_path) - HTMLGen->>Git: Repo(repo_path) - Git-->>HTMLGen: remote URL, name - HTMLGen-->>CLIGen: name, url, github_pages_url - CLIGen->>HTMLGen: generate(output_path, title, repository_url, docs_dir) - HTMLGen->>FS: safe_read(module_tree.json) - HTMLGen->>FS: safe_read(metadata.json) - HTMLGen->>FS: safe_read(viewer_template.html) - HTMLGen->>HTMLGen: _build_info_content(metadata) - HTMLGen->>HTMLGen: substitute placeholders - HTMLGen->>FS: safe_write(index.html) - HTMLGen-->>CLIGen: index.html written -``` - -This corresponds to the `_run_html_generation` step inside `CLIDocumentationGenerator.generate()`: it is only executed when the CLI was invoked with `generate_html=True`, after the backend has produced Markdown files, `module_tree.json`, and `metadata.json` in the output directory. On success, `"index.html"` is appended to `DocumentationJob.files_generated` (see the [Job Models](../job_models/job_models.md) module). - -## Generation Flow - -```mermaid -flowchart TD - Start["generate() called"] --> CheckDocsDir{{"docs_dir provided?"}} - CheckDocsDir -->|"yes"| AutoLoadTree["load_module_tree(docs_dir)"] - CheckDocsDir -->|"yes"| AutoLoadMeta["load_metadata(docs_dir)"] - CheckDocsDir -->|"no"| UseProvided["use provided module_tree/metadata"] - AutoLoadTree --> Defaults - AutoLoadMeta --> Defaults - UseProvided --> Defaults - Defaults["Apply defaults for module_tree/config"] --> LoadTemplate["Load viewer_template.html"] - LoadTemplate -->|"missing"| RaiseErr["raise FileSystemError"] - LoadTemplate -->|"found"| BuildInfo["_build_info_content(metadata)"] - BuildInfo --> BuildRepoLink["Build repository link HTML"] - BuildRepoLink --> ComputeBasePath["Compute docs_base_path"] - ComputeBasePath --> SerializeJson["Serialize config/module_tree/metadata to JSON"] - SerializeJson --> Replace["Replace template placeholders"] - Replace --> WriteOut["safe_write(output_path)"] - WriteOut --> Done["index.html generated"] -``` - -### Template Placeholders - -The generator performs a straightforward string substitution over the template file. The following placeholders are populated by `generate()`: - -| Placeholder | Source | -|---|---| -| `{{TITLE}}` | Escaped `title` argument (e.g., repository name) | -| `{{REPO_LINK}}` | HTML anchor to `repository_url`, empty if not provided | -| `{{SHOW_INFO}}` | `"block"` or `"none"` depending on whether info content was built | -| `{{INFO_CONTENT}}` | HTML fragment from `_build_info_content` (model, timestamp, commit, stats) | -| `{{CONFIG_JSON}}` | JSON-serialized `config` dictionary | -| `{{MODULE_TREE_JSON}}` | JSON-serialized module tree structure | -| `{{METADATA_JSON}}` | JSON-serialized metadata, or the literal `null` | -| `{{DOCS_BASE_PATH}}` | Relative path from the output file to the docs directory | - -## Error Handling - -- Missing `viewer_template.html` raises `FileSystemError`, propagated up to the CLI layer. -- Missing or malformed `module_tree.json` is either substituted with a minimal fallback structure (`load_module_tree`) or, on unexpected read/parse errors, raises `FileSystemError`. -- Missing or malformed `metadata.json` is treated as non-critical: `load_metadata` swallows exceptions and returns `None`, resulting in the info panel being hidden (`{{SHOW_INFO}} = "none"`). -- Git introspection failures in `detect_repository_info` are caught broadly, so the generator always returns a usable (if partially empty) info dictionary rather than failing the whole pipeline. - -## Relationship to Other Modules - -- **[Generation](../generation/generation.md)** — the `CLIDocumentationGenerator` orchestrates when and how `HTMLGenerator` is invoked as part of the overall documentation generation job. -- **[Job Models](../job_models/job_models.md)** — the generated `index.html` file path is recorded in the `DocumentationJob.files_generated` list. -- **[Utils](../utils/utils.md)** — shares low-level filesystem helpers (`safe_read`/`safe_write`) and error types used throughout the CLI. -- **[Cli Core](../../cli-core.md)** — the parent module that aggregates Html Generation alongside generation, configuration, job models, git integration, and shared utilities into the full CLI toolchain. diff --git a/docs/reference/architecture/cli-core/cli-core/job_models/job_models.md b/docs/reference/architecture/cli-core/cli-core/job_models/job_models.md deleted file mode 100644 index fb94a70c..00000000 --- a/docs/reference/architecture/cli-core/cli-core/job_models/job_models.md +++ /dev/null @@ -1,244 +0,0 @@ -# Job Models - -## Introduction - -The Job Models module defines the core data structures used to represent, track, and persist documentation generation jobs within the CodeWiki CLI. It is a foundational, dependency-free module that provides typed dataclasses and enums for job state, configuration, and statistics — enabling consistent serialization, deserialization, and status tracking across the CLI documentation pipeline. - -This module is a child of the [Cli Core](../../cli-core.md) module and is consumed by sibling modules such as [Generation](../generation/generation.md), [Configuration](../configuration/configuration.md), [Utils](../utils/utils.md), and other CLI orchestration components that need to create, update, or persist job records. - -## Purpose and Scope - -The Job Models module has a single, focused responsibility: **define the shape of a documentation job and its lifecycle**. It does not perform any I/O, orchestration, or business logic beyond simple state transitions and serialization helpers. This keeps the module lightweight, easily testable, and safe to import from anywhere in the CLI codebase without introducing circular dependencies. - -Key responsibilities: -- Define the `JobStatus` enum representing the lifecycle states of a job. -- Define `GenerationOptions` for user-configurable generation behavior (branching, GitHub Pages, caching, output paths). -- Define `LLMConfig` for capturing which language models and endpoint were used for a run. -- Define `JobStatistics` for capturing quantitative results of a run (files analyzed, tree depth, tokens consumed). -- Define `DocumentationJob`, the aggregate root that ties together identity, timing, status, configuration, and results for a single documentation generation run. -- Provide robust `to_dict()` / `to_json()` / `from_dict()` methods for safe persistence and rehydration, including defensive coercion of malformed or partial data. - -## Core Components - -### JobStatus - -`JobStatus` is a `str`-backed `Enum` representing the possible lifecycle states of a documentation job: - -| Value | Meaning | -|---|---| -| `pending` | Job has been created but not yet started | -| `running` | Job is actively executing | -| `completed` | Job finished successfully | -| `failed` | Job terminated with an error | - -Because it inherits from both `str` and `Enum`, instances serialize naturally to plain strings in JSON output while still supporting type-safe comparisons in code (e.g., `job.status == JobStatus.RUNNING`). - -### GenerationOptions - -`GenerationOptions` is a dataclass capturing user-facing flags that control how a documentation run behaves: - -- `create_branch` (`bool`): Whether to create a dedicated git branch for the generated docs. Consumed by the [Git Integration](../git_integration/git_integration.md) module. -- `github_pages` (`bool`): Whether to prepare output for GitHub Pages publishing. -- `no_cache` (`bool`): Whether to bypass any caching layer during generation. -- `custom_output` (`Optional[str]`): An optional override for the output directory. - -### LLMConfig - -`LLMConfig` records which language models and endpoint were used to produce a job's output: - -- `main_model` (`str`): The primary model used for documentation generation. -- `cluster_model` (`str`): The model used for clustering/module grouping decisions. -- `base_url` (`str`): The API base URL for the LLM provider. - -This is a plain, required-field dataclass (no defaults), reflecting that when an `LLMConfig` is attached to a job, all three fields are expected to be known. - -### JobStatistics - -`JobStatistics` aggregates quantitative metrics collected during a documentation run: - -- `total_files_analyzed` (`int`): Count of source files processed. -- `leaf_nodes` (`int`): Count of leaf-level components/modules identified. -- `max_depth` (`int`): Maximum depth of the module hierarchy produced. -- `total_tokens_used` (`int`): Total LLM tokens consumed for the run. - -All fields default to `0`, so a freshly created job has a valid, zeroed-out statistics object even before execution begins. - -### DocumentationJob - -`DocumentationJob` is the central aggregate of this module. It represents a single documentation generation run end-to-end, combining identity, repository context, git metadata, timing, status, and nested configuration/statistics objects. - -**Fields:** - -| Field | Type | Description | -|---|---|---| -| `job_id` | `str` | UUID4 identifier, auto-generated if not supplied | -| `repository_path` | `str` | Absolute filesystem path to the repository being documented | -| `repository_name` | `str` | Human-readable repository name | -| `output_directory` | `str` | Destination directory for generated docs | -| `commit_hash` | `str` | Git commit SHA the job was run against | -| `branch_name` | `Optional[str]` | Git branch name, if applicable | -| `timestamp_start` | `str` | ISO-format start timestamp, auto-populated | -| `timestamp_end` | `Optional[str]` | ISO-format end timestamp, set on completion/failure | -| `status` | `JobStatus` | Current lifecycle status | -| `error_message` | `Optional[str]` | Populated when the job fails | -| `files_generated` | `List[str]` | Paths of documentation files produced | -| `module_count` | `int` | Number of modules documented | -| `generation_options` | `GenerationOptions` | Options selected for this run | -| `llm_config` | `Optional[LLMConfig]` | LLM configuration used, if known | -| `statistics` | `JobStatistics` | Quantitative results of the run | - -**Lifecycle methods:** - -- `start()` — Transitions status to `RUNNING` and refreshes `timestamp_start`. -- `complete()` — Transitions status to `COMPLETED` and sets `timestamp_end`. -- `fail(error_message)` — Transitions status to `FAILED`, records the error, and sets `timestamp_end`. - -**Serialization methods:** - -- `to_dict()` — Produces a JSON-serializable `dict` representation, flattening nested dataclasses (`GenerationOptions`, `LLMConfig`, `JobStatistics`) into plain dicts and converting `JobStatus` to its string value. -- `to_json()` — Convenience wrapper around `to_dict()` that returns a pretty-printed JSON string (2-space indent). -- `from_dict(data)` — Classmethod that reconstructs a `DocumentationJob` from a raw dictionary (e.g., loaded from a JSON file), using defensive coercion helpers to tolerate missing, malformed, or partial fields. - -### Defensive Coercion Helpers - -The module defines several private module-level helper functions used exclusively by `DocumentationJob.from_dict()` to safely rebuild nested objects from untrusted or partial data: - -- `_coerce_job_status(value, default)` — Converts a raw value to a valid `JobStatus`, falling back to `JobStatus.PENDING` (or a supplied default) if the value is `None` or not a recognized status string. -- `_coerce_int(value, default)` — Safely converts a value to `int`, returning a default (`0`) on `None`, `TypeError`, or `ValueError`. -- `_coerce_generation_options(value)` — Rebuilds a `GenerationOptions` from a dict, defaulting missing keys. -- `_coerce_llm_config(value)` — Rebuilds an `LLMConfig` from a dict, or returns `None` if no data is present. -- `_coerce_statistics(value)` — Rebuilds a `JobStatistics` from a dict, defaulting missing/invalid numeric fields to `0`. - -This coercion layer makes `DocumentationJob.from_dict()` resilient to schema drift (e.g., loading job files written by an older version of the CLI) without raising exceptions during deserialization. - -## Architecture - -### Component Structure - -```mermaid -classDiagram - class JobStatus { - <> - PENDING - RUNNING - COMPLETED - FAILED - } - class GenerationOptions { - +bool create_branch - +bool github_pages - +bool no_cache - +str custom_output - } - class LLMConfig { - +str main_model - +str cluster_model - +str base_url - } - class JobStatistics { - +int total_files_analyzed - +int leaf_nodes - +int max_depth - +int total_tokens_used - } - class DocumentationJob { - +str job_id - +str repository_path - +str repository_name - +str output_directory - +str commit_hash - +str branch_name - +str timestamp_start - +str timestamp_end - +JobStatus status - +str error_message - +List files_generated - +int module_count - +GenerationOptions generation_options - +LLMConfig llm_config - +JobStatistics statistics - +start() - +complete() - +fail(error_message) - +to_dict() dict - +to_json() str - +from_dict(data) DocumentationJob - } - DocumentationJob --> JobStatus : "status" - DocumentationJob --> GenerationOptions : "generation_options" - DocumentationJob --> LLMConfig : "llm_config (optional)" - DocumentationJob --> JobStatistics : "statistics" -``` - -### Job Lifecycle State Machine - -```mermaid -stateDiagram-v2 - [*] --> PENDING : "DocumentationJob() created" - PENDING --> RUNNING : "start()" - RUNNING --> COMPLETED : "complete()" - RUNNING --> FAILED : "fail(error_message)" - COMPLETED --> [*] - FAILED --> [*] -``` - -### Serialization / Deserialization Flow - -```mermaid -sequenceDiagram - participant Caller as "CLI Component" - participant Job as "DocumentationJob" - participant Coerce as "Coercion Helpers" - - Caller->>Job: "DocumentationJob(...)" - Job-->>Caller: "job instance (status=PENDING)" - Caller->>Job: "job.start()" - Caller->>Job: "job.to_dict() / job.to_json()" - Job-->>Caller: "dict / JSON string" - Note over Caller: "Persisted to disk or transmitted" - - Caller->>Job: "DocumentationJob.from_dict(raw_data)" - Job->>Coerce: "_coerce_job_status(raw_data.status)" - Job->>Coerce: "_coerce_int(raw_data.module_count)" - Job->>Coerce: "_coerce_generation_options(raw_data.generation_options)" - Job->>Coerce: "_coerce_llm_config(raw_data.llm_config)" - Job->>Coerce: "_coerce_statistics(raw_data.statistics)" - Coerce-->>Job: "validated nested objects" - Job-->>Caller: "reconstructed DocumentationJob" -``` - -## Integration with Other Modules - -The Job Models module is intentionally dependency-free (it only imports from the Python standard library), which allows it to be safely imported by any layer of the CLI without risk of circular imports: - -- **[Cli Core](../../cli-core.md)** (parent module): Exposes `DocumentationJob`, `GenerationOptions`, `JobStatistics`, `JobStatus`, and `LLMConfig` as part of its public component surface, making them available to all CLI subsystems. -- **[Generation](../generation/generation.md)**: The `CLIDocumentationGenerator` creates and drives `DocumentationJob` instances through their lifecycle (`start()` → `complete()`/`fail()`), populating `files_generated`, `module_count`, and `statistics` as generation proceeds. -- **[Configuration](../configuration/configuration.md)**: `ConfigManager` and the `Configuration`/`AgentInstructions` models supply values (such as model names and base URLs) that are captured into a job's `LLMConfig`. -- **[Git Integration](../git_integration/git_integration.md)**: `GitManager` supplies `commit_hash` and `branch_name` values and consults `GenerationOptions.create_branch` to decide whether to create a dedicated branch for generated documentation. -- **[Utils](../utils/utils.md)**: `CLILogger` and `ProgressTracker`/`ModuleProgressBar` report on job progress using the status and statistics captured on `DocumentationJob` instances. -- **[Html Generation](../html_generation/html_generation.md)**: `HTMLGenerator` may reference `files_generated` and `output_directory` from a completed job when producing rendered output. - -```mermaid -flowchart TD - JobModels["Job Models"] - CliCore["Cli Core (parent)"] - Generation["Generation"] - Configuration["Configuration"] - GitIntegration["Git Integration"] - HtmlGeneration["Html Generation"] - Utils["Utils"] - - CliCore --> JobModels - Generation -->|"creates & drives"| JobModels - Configuration -->|"populates LLMConfig"| JobModels - GitIntegration -->|"populates commit_hash / branch_name"| JobModels - HtmlGeneration -->|"reads files_generated"| JobModels - Utils -->|"reports on status/statistics"| JobModels -``` - -## Design Notes - -- **Immutability of shape, mutability of state**: While the dataclasses themselves are mutable (no `frozen=True`), the module encourages controlled mutation through explicit lifecycle methods (`start()`, `complete()`, `fail()`) rather than direct field assignment, keeping status transitions predictable and auditable. -- **Safe round-tripping**: The combination of `to_dict()`/`to_json()` and `from_dict()` with defensive coercion helpers ensures that job records can be persisted to disk (e.g., as job history or resumable state) and reloaded reliably even if the on-disk schema is incomplete or slightly out of date. -- **String-backed enum**: `JobStatus` extending `str` means job status values serialize directly as readable strings in JSON without requiring custom encoders, simplifying integration with logging, CLI output, and external tooling. -- **Zero external dependencies**: The module relies only on `dataclasses`, `datetime`, `typing`, `enum`, `uuid`, and `json` from the standard library, making it trivially testable and reusable across CLI contexts. diff --git a/docs/reference/architecture/cli-core/cli-core/utils/utils.md b/docs/reference/architecture/cli-core/cli-core/utils/utils.md deleted file mode 100644 index 339a3fd3..00000000 --- a/docs/reference/architecture/cli-core/cli-core/utils/utils.md +++ /dev/null @@ -1,168 +0,0 @@ -# Utils - -The Utils module provides the foundational output and progress-reporting primitives used throughout the [Cli Core](../../cli-core.md) subsystem. It contains no business logic of its own; instead it supplies two small, dependency-light building blocks — a colored console logger and a multi-stage progress tracker — that every other CLI component (job orchestration, generation, git integration, HTML generation) relies on to communicate status to the end user. - -Because these utilities sit at the bottom of the CLI dependency graph, they are intentionally free of imports from other CLI modules. This keeps them reusable, easy to test, and safe to import from anywhere in the CLI without risking circular dependencies. - -## Purpose and Scope - -The module addresses two related but distinct concerns: - -1. **Structured, leveled console output** — via `CLILogger`, which standardizes how informational, success, warning, error, debug, and step messages are rendered to the terminal (including color and timestamps), and quiets noisy third-party HTTP/SDK loggers. -2. **Progress and ETA reporting** — via `ProgressTracker` (stage-weighted progress across the overall documentation pipeline) and `ModuleProgressBar` (per-module progress during the module-by-module generation phase). - -These two concerns are complementary: `CLILogger` handles ad-hoc textual messages, while the progress classes handle numeric/visual progress state and time estimation for long-running operations. - -## Component Overview - -```mermaid -classDiagram - class CLILogger { - +bool verbose - +datetime start_time - +debug(message) void - +info(message) void - +success(message) void - +warning(message) void - +error(message) void - +step(message, step, total) void - +elapsed_time() str - } - - class ProgressTracker { - +int total_stages - +int current_stage - +float stage_progress - +float start_time - +bool verbose - +STAGE_WEIGHTS dict - +STAGE_NAMES dict - +start_stage(stage, description) void - +update_stage(progress, message) void - +complete_stage(message) void - +get_overall_progress() float - +get_eta() str - } - - class ModuleProgressBar { - +int total_modules - +int current_module - +bool verbose - +bar - +update(module_name, cached) void - +finish() void - } -``` - -### CLILogger - -`CLILogger` (`codewiki/cli/utils/logging.py`) is the standard mechanism for writing user-facing output during a CLI run. It wraps [Click](https://click.palletsprojects.com/)'s `echo`/`secho` helpers and adds: - -- **Leveled methods**: `debug`, `info`, `success`, `warning`, `error`, and `step`, each with a distinct color and symbol (✓ for success, ⚠️ for warning, ✗ for error) so terminal output is easy to scan. -- **Verbose-only debug output**: `debug()` messages are only rendered when `verbose=True`, and are timestamped for correlation with other logs. -- **Step announcements**: `step()` renders either a `[current/total]` prefix (when both `step` and `total` are supplied) or a generic arrow (`→`) prefix, useful for announcing pipeline phases. -- **Elapsed time tracking**: `elapsed_time()` reports time since the logger was constructed, formatted as `Xm Ys` or `Ys`. - -The module also exposes two module-level helpers: - -- `quiet_third_party_loggers(level=logging.WARNING)` — caps the log level of noisy third-party libraries (`httpx`, `openai`, `openai._base_client`, `anthropic`) so that a documentation run does not produce thousands of INFO-level HTTP request lines in CI output. This is applied explicitly rather than as an import-time side effect, so its behavior is visible at the call site. -- `create_logger(verbose=False)` — the standard factory used by CLI entry points to construct a properly configured `CLILogger`, automatically invoking `quiet_third_party_loggers()` first. - -### ProgressTracker - -`ProgressTracker` (`codewiki/cli/utils/progress.py`) models the overall documentation generation pipeline as five weighted stages: - -| Stage | Name | Weight | -|-------|------|--------| -| 1 | Dependency Analysis | 40% | -| 2 | Module Clustering | 20% | -| 3 | Documentation Generation | 30% | -| 4 | HTML Generation (optional) | 5% | -| 5 | Finalization | 5% | - -These weights (`STAGE_WEIGHTS`) reflect the relative time each stage is expected to consume, and are used to compute an aggregate `get_overall_progress()` value across the whole run, independent of how much intra-stage progress has been made. - -Key behaviors: - -- `start_stage(stage, description=None)` resets stage progress to `0.0`, records the stage start time, and prints a banner (verbose mode includes elapsed time; non-verbose mode shows a compact `[stage/total]` header). -- `update_stage(progress, message=None)` clamps `progress` to `[0.0, 1.0]` and, in verbose mode, prints an indented status message. -- `complete_stage(message=None)` sets stage progress to `1.0` and, in verbose mode, prints the stage's wall-clock duration plus any completion message. -- `get_overall_progress()` sums the weights of fully completed stages plus the weighted fraction of the current stage's progress. -- `get_eta()` extrapolates total run time from elapsed time and overall progress, returning a human-readable estimate (e.g., `"2m 15s"`, `"1h 5m"`, or `"< 1 min"`), or `None` if no progress has been made yet (avoiding a divide-by-zero). - -This stage model is the shared contract that pipeline-driving code (in the [Job Models](../job_models/job_models.md) and [Generation](../generation/generation.md) modules) uses to report high-level progress consistently. - -### ModuleProgressBar - -`ModuleProgressBar` (`codewiki/cli/utils/progress.py`) is a narrower, complementary tool used specifically during the "Documentation Generation" stage, where progress is naturally expressed as "N of M modules processed" rather than a continuous percentage. - -- In **non-verbose** mode, it wraps `click.progressbar` to render a live terminal progress bar with ETA and percentage, entering the context manager on construction (`__enter__`) and exiting it on `finish()` (`__exit__`). -- In **verbose** mode, no bar is drawn; instead, `update()` prints one line per module, showing whether the module was `✓ (cached)` or `⟳ (generating)`. -- `update(module_name, cached=False)` increments the internal module counter and reports progress using whichever mode is active. -- `finish()` safely closes the underlying `click.progressbar` context if one was opened. - -Because `ModuleProgressBar` owns a `click.progressbar` context manager internally, callers should ensure `finish()` is invoked (e.g., in a `finally` block) even if module generation raises an exception, to avoid leaving the terminal progress bar in an inconsistent state. - -## Interaction with the Broader CLI Pipeline - -The Utils module is consumed — not the consumer. Higher-level orchestration code (such as the documentation generation flow) drives both `ProgressTracker` and `ModuleProgressBar` in tandem: `ProgressTracker` reports macro-level stage progress across the whole run, while `ModuleProgressBar` provides fine-grained visibility during the module generation stage specifically. `CLILogger` is used throughout for all textual status, warnings, and errors. - -```mermaid -sequenceDiagram - participant Pipeline as "CLI Pipeline" - participant Logger as "CLILogger" - participant Tracker as "ProgressTracker" - participant ModBar as "ModuleProgressBar" - - Pipeline->>Logger: create_logger(verbose) - Pipeline->>Tracker: start_stage(1, "Dependency Analysis") - Pipeline->>Logger: info("Analyzing repository...") - Tracker-->>Pipeline: update_stage(progress) - Pipeline->>Tracker: complete_stage() - - Pipeline->>Tracker: start_stage(3, "Documentation Generation") - Pipeline->>ModBar: new ModuleProgressBar(total_modules) - loop "for each module" - Pipeline->>ModBar: update(module_name, cached) - Pipeline->>Logger: debug("module details") - end - Pipeline->>ModBar: finish() - Pipeline->>Tracker: complete_stage("Generation finished") - - Pipeline->>Logger: success("Documentation generated") -``` - -## Typical Usage Flow - -```mermaid -flowchart TD - A["create_logger(verbose)"] --> B["ProgressTracker(total_stages=5)"] - B --> C["tracker.start_stage(1, 'Dependency Analysis')"] - C --> D["perform analysis; call tracker.update_stage()"] - D --> E["tracker.complete_stage()"] - E --> F["tracker.start_stage(3, 'Documentation Generation')"] - F --> G["ModuleProgressBar(total_modules)"] - G --> H["for each module: bar.update(name, cached)"] - H --> I["bar.finish()"] - I --> J["tracker.complete_stage()"] - J --> K["logger.success('Done')"] -``` - -## Design Notes - -- **No cross-module coupling**: Neither `CLILogger` nor the progress classes depend on other CLI modules such as [Configuration](../configuration/configuration.md), [Job Models](../job_models/job_models.md), or [Generation](../generation/generation.md). This keeps them trivially reusable and testable in isolation. -- **Verbose vs. non-verbose modes are first-class**: Every class exposes a `verbose` flag that changes rendering strategy (timestamped multi-line logs vs. compact single-line/progress-bar output), rather than relying on global logging configuration. -- **Explicit side-effect application**: `quiet_third_party_loggers()` is called from `create_logger()` rather than at module import time, so that importing the `utils` package never silently mutates global logging state — the effect only occurs when a caller explicitly requests a logger. -- **Time estimation is defensive**: `get_eta()` guards against division by zero when no progress has been recorded, returning `None` instead of raising or producing a nonsensical estimate. - -## Relationship to Parent and Sibling Modules - -The Utils module is a child of [Cli Core](../../cli-core.md), alongside: - -- [Generation](../generation/generation.md) — the documentation generation adapter that most heavily relies on `ProgressTracker` and `ModuleProgressBar` to report pipeline progress. -- [Configuration](../configuration/configuration.md) — CLI configuration models and the config manager. -- [Job Models](../job_models/job_models.md) — data models representing documentation jobs, statuses, and statistics that the progress-tracking stages correspond to. -- [Git Integration](../git_integration/git_integration.md) — git operations invoked during pipeline execution, whose status is typically reported via `CLILogger`. -- [Html Generation](../html_generation/html_generation.md) — the optional HTML rendering stage represented as Stage 4 in `ProgressTracker.STAGE_NAMES`. - -These sibling modules depend on Utils for consistent status reporting; Utils itself has no dependency on them. diff --git a/docs/reference/architecture/cli-core/cli_core/utils/utils.md b/docs/reference/architecture/cli-core/cli_core/utils/utils.md deleted file mode 100644 index ac1b4829..00000000 --- a/docs/reference/architecture/cli-core/cli_core/utils/utils.md +++ /dev/null @@ -1,168 +0,0 @@ -# Utils - -The Utils module provides the foundational output and progress-reporting primitives used throughout the [CLI Core](../cli_core.md) subsystem. It contains no business logic of its own; instead it supplies two small, dependency-light building blocks — a colored console logger and a multi-stage progress tracker — that every other CLI component (job orchestration, generation, git integration, HTML generation) relies on to communicate status to the end user. - -Because these utilities sit at the bottom of the CLI dependency graph, they are intentionally free of imports from other CLI modules. This keeps them reusable, easy to test, and safe to import from anywhere in the CLI without risking circular dependencies. - -## Purpose and Scope - -The module addresses two related but distinct concerns: - -1. **Structured, leveled console output** — via `CLILogger`, which standardizes how informational, success, warning, error, debug, and step messages are rendered to the terminal (including color and timestamps), and quiets noisy third-party HTTP/SDK loggers. -2. **Progress and ETA reporting** — via `ProgressTracker` (stage-weighted progress across the overall documentation pipeline) and `ModuleProgressBar` (per-module progress during the module-by-module generation phase). - -These two concerns are complementary: `CLILogger` handles ad-hoc textual messages, while the progress classes handle numeric/visual progress state and time estimation for long-running operations. - -## Component Overview - -```mermaid -classDiagram - class CLILogger { - +bool verbose - +datetime start_time - +debug(message) void - +info(message) void - +success(message) void - +warning(message) void - +error(message) void - +step(message, step, total) void - +elapsed_time() str - } - - class ProgressTracker { - +int total_stages - +int current_stage - +float stage_progress - +float start_time - +bool verbose - +STAGE_WEIGHTS dict - +STAGE_NAMES dict - +start_stage(stage, description) void - +update_stage(progress, message) void - +complete_stage(message) void - +get_overall_progress() float - +get_eta() str - } - - class ModuleProgressBar { - +int total_modules - +int current_module - +bool verbose - +bar - +update(module_name, cached) void - +finish() void - } -``` - -### CLILogger - -`CLILogger` (`codewiki/cli/utils/logging.py`) is the standard mechanism for writing user-facing output during a CLI run. It wraps [Click](https://click.palletsprojects.com/)'s `echo`/`secho` helpers and adds: - -- **Leveled methods**: `debug`, `info`, `success`, `warning`, `error`, and `step`, each with a distinct color and symbol (✓ for success, ⚠️ for warning, ✗ for error) so terminal output is easy to scan. -- **Verbose-only debug output**: `debug()` messages are only rendered when `verbose=True`, and are timestamped for correlation with other logs. -- **Step announcements**: `step()` renders either a `[current/total]` prefix (when both `step` and `total` are supplied) or a generic arrow (`→`) prefix, useful for announcing pipeline phases. -- **Elapsed time tracking**: `elapsed_time()` reports time since the logger was constructed, formatted as `Xm Ys` or `Ys`. - -The module also exposes two module-level helpers: - -- `quiet_third_party_loggers(level=logging.WARNING)` — caps the log level of noisy third-party libraries (`httpx`, `openai`, `openai._base_client`, `anthropic`) so that a documentation run does not produce thousands of INFO-level HTTP request lines in CI output. This is applied explicitly rather than as an import-time side effect, so its behavior is visible at the call site. -- `create_logger(verbose=False)` — the standard factory used by CLI entry points to construct a properly configured `CLILogger`, automatically invoking `quiet_third_party_loggers()` first. - -### ProgressTracker - -`ProgressTracker` (`codewiki/cli/utils/progress.py`) models the overall documentation generation pipeline as five weighted stages: - -| Stage | Name | Weight | -|-------|------|--------| -| 1 | Dependency Analysis | 40% | -| 2 | Module Clustering | 20% | -| 3 | Documentation Generation | 30% | -| 4 | HTML Generation (optional) | 5% | -| 5 | Finalization | 5% | - -These weights (`STAGE_WEIGHTS`) reflect the relative time each stage is expected to consume, and are used to compute an aggregate `get_overall_progress()` value across the whole run, independent of how much intra-stage progress has been made. - -Key behaviors: - -- `start_stage(stage, description=None)` resets stage progress to `0.0`, records the stage start time, and prints a banner (verbose mode includes elapsed time; non-verbose mode shows a compact `[stage/total]` header). -- `update_stage(progress, message=None)` clamps `progress` to `[0.0, 1.0]` and, in verbose mode, prints an indented status message. -- `complete_stage(message=None)` sets stage progress to `1.0` and, in verbose mode, prints the stage's wall-clock duration plus any completion message. -- `get_overall_progress()` sums the weights of fully completed stages plus the weighted fraction of the current stage's progress. -- `get_eta()` extrapolates total run time from elapsed time and overall progress, returning a human-readable estimate (e.g., `"2m 15s"`, `"1h 5m"`, or `"< 1 min"`), or `None` if no progress has been made yet (avoiding a divide-by-zero). - -This stage model is the shared contract that pipeline-driving code (in the [Job Models](../job_models/job_models.md) and generation layers) uses to report high-level progress consistently. - -### ModuleProgressBar - -`ModuleProgressBar` (`codewiki/cli/utils/progress.py`) is a narrower, complementary tool used specifically during the "Documentation Generation" stage, where progress is naturally expressed as "N of M modules processed" rather than a continuous percentage. - -- In **non-verbose** mode, it wraps `click.progressbar` to render a live terminal progress bar with ETA and percentage, entering the context manager on construction (`__enter__`) and exiting it on `finish()` (`__exit__`). -- In **verbose** mode, no bar is drawn; instead, `update()` prints one line per module, showing whether the module was `✓ (cached)` or `⟳ (generating)`. -- `update(module_name, cached=False)` increments the internal module counter and reports progress using whichever mode is active. -- `finish()` safely closes the underlying `click.progressbar` context if one was opened. - -Because `ModuleProgressBar` owns a `click.progressbar` context manager internally, callers should ensure `finish()` is invoked (e.g., in a `finally` block) even if module generation raises an exception, to avoid leaving the terminal progress bar in an inconsistent state. - -## Interaction with the Broader CLI Pipeline - -The Utils module is consumed — not the consumer. Higher-level orchestration code (such as the documentation generation flow) drives both `ProgressTracker` and `ModuleProgressBar` in tandem: `ProgressTracker` reports macro-level stage progress across the whole run, while `ModuleProgressBar` provides fine-grained visibility during the module generation stage specifically. `CLILogger` is used throughout for all textual status, warnings, and errors. - -```mermaid -sequenceDiagram - participant Pipeline as "CLI Pipeline" - participant Logger as "CLILogger" - participant Tracker as "ProgressTracker" - participant ModBar as "ModuleProgressBar" - - Pipeline->>Logger: create_logger(verbose) - Pipeline->>Tracker: start_stage(1, "Dependency Analysis") - Pipeline->>Logger: info("Analyzing repository...") - Tracker-->>Pipeline: update_stage(progress) - Pipeline->>Tracker: complete_stage() - - Pipeline->>Tracker: start_stage(3, "Documentation Generation") - Pipeline->>ModBar: new ModuleProgressBar(total_modules) - loop "for each module" - Pipeline->>ModBar: update(module_name, cached) - Pipeline->>Logger: debug("module details") - end - Pipeline->>ModBar: finish() - Pipeline->>Tracker: complete_stage("Generation finished") - - Pipeline->>Logger: success("Documentation generated") -``` - -## Typical Usage Flow - -```mermaid -flowchart TD - A["create_logger(verbose)"] --> B["ProgressTracker(total_stages=5)"] - B --> C["tracker.start_stage(1, 'Dependency Analysis')"] - C --> D["perform analysis; call tracker.update_stage()"] - D --> E["tracker.complete_stage()"] - E --> F["tracker.start_stage(3, 'Documentation Generation')"] - F --> G["ModuleProgressBar(total_modules)"] - G --> H["for each module: bar.update(name, cached)"] - H --> I["bar.finish()"] - I --> J["tracker.complete_stage()"] - J --> K["logger.success('Done')"] -``` - -## Design Notes - -- **No cross-module coupling**: Neither `CLILogger` nor the progress classes depend on other CLI modules such as [Configuration](../configuration/configuration.md), [Job Models](../job_models/job_models.md), or [Generation](../generation/generation.md). This keeps them trivially reusable and testable in isolation. -- **Verbose vs. non-verbose modes are first-class**: Every class exposes a `verbose` flag that changes rendering strategy (timestamped multi-line logs vs. compact single-line/progress-bar output), rather than relying on global logging configuration. -- **Explicit side-effect application**: `quiet_third_party_loggers()` is called from `create_logger()` rather than at module import time, so that importing the `utils` package never silently mutates global logging state — the effect only occurs when a caller explicitly requests a logger. -- **Time estimation is defensive**: `get_eta()` guards against division by zero when no progress has been recorded, returning `None` instead of raising or producing a nonsensical estimate. - -## Relationship to Parent and Sibling Modules - -The Utils module is a child of [CLI Core](../cli_core.md), alongside: - -- [Generation](../generation/generation.md) — the documentation generation adapter that most heavily relies on `ProgressTracker` and `ModuleProgressBar` to report pipeline progress. -- [Configuration](../configuration/configuration.md) — CLI configuration models and the config manager. -- [Job Models](../job_models/job_models.md) — data models representing documentation jobs, statuses, and statistics that the progress-tracking stages correspond to. -- [Git Integration](../git_integration/git_integration.md) — git operations invoked during pipeline execution, whose status is typically reported via `CLILogger`. -- [HTML Generation](../html_generation/html_generation.md) — the optional HTML rendering stage represented as Stage 4 in `ProgressTracker.STAGE_NAMES`. - -These sibling modules depend on Utils for consistent status reporting; Utils itself has no dependency on them. diff --git a/docs/reference/architecture/cli-core/configuration.md b/docs/reference/architecture/cli-core/configuration.md deleted file mode 100644 index dda61e21..00000000 --- a/docs/reference/architecture/cli-core/configuration.md +++ /dev/null @@ -1,249 +0,0 @@ -# Configuration - -The Configuration module provides the persistent settings layer for the CodeWiki CLI. It defines the data models that describe user preferences (LLM providers, model names, token limits, temperature settings, and documentation-generation instructions) and the manager responsible for reading and writing those settings safely to disk and to the operating system's secure credential store. - -This module is a child of [Cli Core](../cli-core.md) and works alongside sibling modules such as the job/runtime models, generation pipeline, and utility helpers to support the CLI's end-to-end workflow. - -## Purpose and Scope - -The Configuration module answers three core questions for the CLI: - -1. **What settings does the user have configured?** — captured by the `Configuration` dataclass. -2. **How should the documentation agent behave for a given run?** — captured by `AgentInstructions`. -3. **How are these settings persisted, validated, and loaded securely?** — handled by `ConfigManager`. - -Sensitive values (API keys) are never written to plaintext configuration files. Instead, they are stored using the system keyring (macOS Keychain, Windows Credential Manager, or Linux Secret Service) and only non-sensitive settings are persisted to `~/.codewiki/config.json`. - -## Core Components - -| Component | Responsibility | -|---|---| -| `Configuration` | Dataclass representing all persistent CLI settings: model names, base URLs, API versions, token/temperature limits, and clustering parameters. | -| `AgentInstructions` | Dataclass representing optional, user-customizable instructions for the documentation agent (file filters, focus modules, doc type, free-form instructions). | -| `ConfigManager` | Orchestrates loading/saving `Configuration` to `~/.codewiki/config.json` and securely storing/retrieving API keys via keyring. | - -## Architecture Overview - -```mermaid -flowchart TD - subgraph ConfigModule["Configuration Module"] - CM["ConfigManager"] - Cfg["Configuration"] - AI["AgentInstructions"] - end - - FS["~/.codewiki/config.json"] - KR["System Keyring"] - Backend["Backend Config"] - - CM -->|"load() / save()"| FS - CM -->|"get/set API keys"| KR - CM -->|"holds"| Cfg - Cfg -->|"has one"| AI - Cfg -->|"to_backend_config()"| Backend -``` - -`Configuration` is a plain, serializable dataclass with no direct dependency on keyring or the filesystem — those concerns are owned exclusively by `ConfigManager`. This separation keeps the data model easy to test and reuse, while `ConfigManager` acts as the single access point for persistence. - -## Data Model: Configuration - -`Configuration` captures all settings needed to drive a documentation-generation run, organized into three provider "roles": - -- **cluster** — model used for module clustering/decomposition -- **main** — primary model used for documentation generation -- **fallback** — fallback model used when the main model fails or is rate-limited - -For each role, the model tracks: -- Model name (`*_model`) -- Base URL (`*_base_url`, optional — for OpenAI-compatible or self-hosted endpoints) -- API version (`*_api_version`, optional) -- Max tokens (`*_max_tokens`) -- Temperature (`*_temperature`) and whether temperature is supported (`*_temperature_supported`) -- Max-token parameter field name (`*_max_token_field`, e.g., `"max_tokens"` vs. provider-specific names) - -In addition to per-provider settings, `Configuration` holds shared clustering parameters (`max_token_per_module`, `max_token_per_leaf_module`, `max_depth`) and a `default_output` directory. API keys (`cluster_api_key`, `main_api_key`, `fallback_api_key`) exist on the dataclass only as **runtime-only fields** — they are never serialized by `to_dict()` and are populated separately from the keyring by `ConfigManager`. - -```mermaid -classDiagram - class Configuration { - +str main_model - +str cluster_model - +str fallback_model - +str default_output - +str cluster_base_url - +str main_base_url - +str fallback_base_url - +int cluster_max_tokens - +int main_max_tokens - +int fallback_max_tokens - +float cluster_temperature - +float main_temperature - +float fallback_temperature - +bool cluster_temperature_supported - +bool main_temperature_supported - +bool fallback_temperature_supported - +int max_token_per_module - +int max_token_per_leaf_module - +int max_depth - +AgentInstructions agent_instructions - +validate() - +to_dict() dict - +from_dict(data) Configuration - +is_complete() bool - +to_backend_config(...) Config - } - - class AgentInstructions { - +List~str~ include_patterns - +List~str~ exclude_patterns - +List~str~ focus_modules - +str doc_type - +str custom_instructions - +to_dict() dict - +from_dict(data) AgentInstructions - +is_empty() bool - +get_prompt_addition() str - } - - Configuration "1" *-- "1" AgentInstructions : agent_instructions -``` - -### Validation - -`Configuration.validate()` performs field-level checks: -- Base URLs (when set) are validated via `validate_url`. -- Model names for all three roles are validated via `validate_model_name`. - -`from_dict()` performs defensive type coercion when loading from JSON (which may contain string-typed numbers/booleans due to manual editing), converting and range-checking integers (e.g., token limits), floats (temperature, bounded `0.0`–`2.0`), and booleans, raising `ValueError` on invalid input. - -### Serialization Rules - -- `to_dict()` only emits optional fields (`*_base_url`, `*_api_version`, `agent_instructions`) when they are set/non-empty, keeping the persisted JSON minimal. -- API keys are **excluded** from `to_dict()` entirely — they are never written to `config.json`. - -## Data Model: AgentInstructions - -`AgentInstructions` lets users customize how the documentation agent analyzes a repository and generates content: - -- `include_patterns` / `exclude_patterns` — glob-style file filters (e.g., `["*.cs"]`, `["*Tests*"]`) -- `focus_modules` — modules that should receive more detailed documentation -- `doc_type` — a preset documentation style (`api`, `architecture`, `user-guide`, `developer`) or a free-form type -- `custom_instructions` — arbitrary additional guidance passed to the LLM - -`get_prompt_addition()` translates these fields into a natural-language instruction block that is merged into the agent's prompt at generation time. `is_empty()` allows callers to distinguish "no customization" from "customization with all-default values." - -Instructions can be set persistently (stored in `config.json` as part of `Configuration`) or supplied at runtime for a single job; `Configuration.to_backend_config()` merges the two, with runtime instructions taking precedence field-by-field. - -## ConfigManager: Persistence and Secure Storage - -`ConfigManager` is the sole component responsible for reading and writing configuration state. It manages two independent storage backends: - -1. **JSON file** (`~/.codewiki/config.json`) — non-sensitive settings, versioned with a `CONFIG_VERSION` marker for future migrations. -2. **System keyring** (via the `keyring` library) — the three provider API keys (`cluster_api_key`, `main_api_key`, `fallback_api_key`), stored under a shared `codewiki` service name with distinct account identifiers. - -```mermaid -classDiagram - class ConfigManager { - -Optional~str~ _cluster_api_key - -Optional~str~ _main_api_key - -Optional~str~ _fallback_api_key - -Optional~Configuration~ _config - -bool _keyring_available - +load() bool - +save(...) void - +get_cluster_api_key() Optional~str~ - +get_main_api_key() Optional~str~ - +get_fallback_api_key() Optional~str~ - +get_config() Optional~Configuration~ - +is_configured() bool - +delete_api_keys() void - +clear() void - +keyring_available bool - +config_file_path Path - } - ConfigManager --> Configuration : loads/saves -``` - -### Load Flow - -```mermaid -sequenceDiagram - participant Caller - participant CM as ConfigManager - participant FS as "config.json" - participant KR as "System Keyring" - - Caller->>CM: load() - CM->>FS: check exists / read - alt file missing - FS-->>CM: not found - CM-->>Caller: False - else file present - FS-->>CM: JSON content - CM->>CM: Configuration.from_dict(data) - CM->>KR: get_password(cluster_api_key) - CM->>KR: get_password(main_api_key) - CM->>KR: get_password(fallback_api_key) - KR-->>CM: key values (or None) - CM-->>Caller: True - end -``` - -### Save Flow - -```mermaid -sequenceDiagram - participant Caller - participant CM as ConfigManager - participant FS as "config.json" - participant KR as "System Keyring" - - Caller->>CM: save(fields..., api_keys...) - CM->>FS: ensure_directory(CONFIG_DIR) - alt no in-memory config - CM->>CM: load() existing or create default Configuration - end - CM->>CM: apply provided field updates - CM->>CM: Configuration.validate() (if models set) - CM->>KR: set_password(...) for each provided API key - CM->>FS: write JSON (version + Configuration.to_dict()) - CM-->>Caller: done (or raises ConfigurationError) -``` - -Key behaviors: -- **Partial updates**: `save()` accepts every field as an optional keyword argument; only provided values overwrite the in-memory `Configuration`, allowing incremental configuration (e.g., `codewiki configure` sub-commands that set one field at a time). -- **Validation gate**: full validation (`Configuration.validate()`) only runs once `main_model` and `cluster_model` are both set, avoiding premature failures during multi-step setup. -- **Keyring failures surface as `ConfigurationError`**, with an actionable message when the OS keychain is unavailable or misconfigured. -- **`is_configured()`** combines two checks: all three API keys must be retrievable from keyring, and `Configuration.is_complete()` must be true (all three model names set). -- **`clear()`** performs a full reset — deleting API keys from keyring and removing `config.json` — used by commands like `codewiki configure --reset`. - -## Bridging to Runtime Execution: to_backend_config - -`Configuration.to_backend_config()` is the seam between this module's persistent settings and the runtime configuration consumed by the documentation-generation backend. It: - -1. Fetches any missing API keys from the keyring via a fresh `ConfigManager` instance (if not explicitly passed in). -2. Merges `runtime_instructions` (per-invocation `AgentInstructions`) over the persisted `agent_instructions`, with runtime values taking precedence field-by-field. -3. Constructs and returns a backend `Config` object (`Config.from_cli(...)`) populated with all model, token, temperature, and clustering settings, ready to drive a documentation job. - -```mermaid -flowchart LR - A["CLI Command"] --> B["ConfigManager.load()"] - B --> C["Configuration"] - C --> D["Configuration.to_backend_config()"] - D --> E["ConfigManager (keyring lookup for missing keys)"] - D --> F["Merge AgentInstructions (runtime over persisted)"] - D --> G["Config.from_cli(...)"] - G --> H["Backend Config"] -``` - -This bridging pattern keeps the CLI's persistent, user-facing settings model (`Configuration`) decoupled from the backend's execution-time `Config` model. The backend `Config` class itself is documented in the [Config Core](../config-core.md) module. - -## Relationship to Other CLI Modules - -- **[Job Models](../job_models/job_models.md)** — represents the runtime state of an in-progress documentation job (`DocumentationJob`, `JobStatus`, `LLMConfig`, `GenerationOptions`, `JobStatistics`). Where `Configuration` describes *persistent user preferences*, the job models describe the *live execution* of a single generation run, often derived from a `Configuration` via `to_backend_config()`. -- **[Generation](../generation/generation.md)** — the `CLIDocumentationGenerator` adapter consumes a resolved `Configuration`/backend `Config` to drive the documentation pipeline. -- **[Utils](../utils/utils.md)** — provides shared filesystem (`ensure_directory`, `safe_read`, `safe_write`) and error types (`ConfigurationError`, `FileSystemError`) used internally by `ConfigManager`. - -## Summary - -The Configuration module is the trust boundary for user secrets and the single source of truth for persistent CLI preferences. `Configuration` and `AgentInstructions` define what can be configured; `ConfigManager` defines how those settings are safely loaded, validated, updated, and translated into a form the backend documentation pipeline can execute against. diff --git a/docs/reference/architecture/cli-core/generation.md b/docs/reference/architecture/cli-core/generation.md deleted file mode 100644 index fb1d79d9..00000000 --- a/docs/reference/architecture/cli-core/generation.md +++ /dev/null @@ -1,167 +0,0 @@ -# Generation - -The Generation module is the central orchestration layer of the CodeWiki CLI. It bridges the command-line interface with the backend documentation engine, coordinating the full end-to-end pipeline that turns a source code repository into a structured, hierarchical documentation set — including dependency analysis, LLM-driven module clustering, documentation generation, optional Mermaid diagram extraction, and optional HTML output. - -Its single core component, `CLIDocumentationGenerator`, acts as an **adapter**: it wraps the backend's `DocumentationGenerator` (from [Backend Core](../backend-core.md)) with CLI-specific concerns such as progress reporting, colored logging, verbose diagnostics, and job lifecycle tracking via the [Job Models](../job_models/job_models.md) module. - -## Purpose and Role in the System - -As a child of [CLI Core](../cli-core.md), the Generation module is invoked whenever a user runs a documentation generation command. It does not implement dependency analysis, clustering, or documentation writing itself — those responsibilities belong to `backend-core`. Instead, Generation is responsible for: - -1. **Translating CLI configuration** into the backend's `Config` object (model selection, API keys, base URLs, token limits, clustering depth, agent instructions, additional source paths). -2. **Orchestrating the multi-stage pipeline** (dependency analysis → module clustering → documentation generation → optional HTML generation → finalization) with a five-stage `ProgressTracker`. -3. **Tracking job state** using a `DocumentationJob` model, recording statistics, generated files, and success/failure outcomes. -4. **Handling synthetic module fallback** to avoid context-window overflows when the LLM clustering step returns an empty module tree. -5. **Optionally generating an HTML viewer** via [HTML Generation](../html_generation/html_generation.md) and extracting Mermaid diagrams into a dedicated directory. - -## Architecture Overview - -```mermaid -flowchart TD - CLI["CLI Command Entry Point"] --> CDG["CLIDocumentationGenerator"] - CDG --> PT["ProgressTracker (utils)"] - CDG --> Job["DocumentationJob (job_models)"] - CDG --> BC["BackendConfig (config-core)"] - CDG --> DG["DocumentationGenerator (backend-core)"] - DG --> GB["DependencyGraphBuilder"] - DG --> AO["AgentOrchestrator"] - CDG --> CM["cluster_modules (backend-core)"] - CDG --> HG["HTMLGenerator (html_generation)"] - CDG --> LOG["ColoredFormatter (backend-core logging)"] -``` - -## Core Component - -### CLIDocumentationGenerator - -`CLIDocumentationGenerator` (`codewiki/cli/adapters/doc_generator.py`) is instantiated with: - -- `repo_path` — path to the repository being documented -- `output_dir` — target directory for generated docs -- `config` — a dictionary of LLM/model configuration (models, API keys, base URLs, token limits, temperature settings, clustering parameters, agent instructions, additional source paths) -- `verbose` — whether to emit detailed diagnostic output -- `generate_html` — whether to produce an HTML viewer (`index.html`) -- `diagrams_dir` — optional separate output directory for extracted Mermaid diagrams - -On construction, it: -- Creates a `ProgressTracker` configured for 5 pipeline stages. -- Creates a `DocumentationJob` and populates its metadata (repository path/name, output directory, `LLMConfig`). -- Configures backend logging by attaching a `ColoredFormatter`-based handler to the `codewiki.src.be` logger namespace, switching verbosity between `INFO` (verbose) and `WARNING` (quiet) levels, and disabling propagation to avoid duplicate log lines. - -#### Key Responsibilities - -| Responsibility | Method | -|---|---| -| Build backend configuration from CLI config dict | `generate()` | -| Run the full async generation pipeline | `_run_backend_generation()` | -| Generate the HTML viewer | `_run_html_generation()` | -| Ensure job metadata file exists | `_finalize_job()` | -| Configure colored backend logging | `_configure_backend_logging()` | - -## Generation Pipeline - -The `generate()` method is the public entry point. It is synchronous from the caller's perspective but internally drives an `asyncio`-based backend pipeline. It returns a completed `DocumentationJob` (see [Job Models](../job_models/job_models.md)) or raises an `APIError` on failure. - -```mermaid -sequenceDiagram - participant Caller as "CLI Command" - participant CDG as "CLIDocumentationGenerator" - participant BC as "BackendConfig" - participant DG as "DocumentationGenerator" - participant CM as "cluster_modules" - participant HG as "HTMLGenerator" - participant Job as "DocumentationJob" - - Caller->>CDG: generate() - CDG->>Job: start() - CDG->>BC: Config.from_cli(...) - CDG->>DG: _run_backend_generation(backend_config) - DG->>DG: graph_builder.build_dependency_graph() - DG->>CM: cluster_modules(leaf_nodes, components, config) - Note over DG,CM: Synthetic module fallback if tree is empty - DG->>DG: generate_module_documentation(components, leaf_nodes) - opt diagrams_dir configured - DG->>DG: extract_and_save_mermaid_diagrams() - end - opt generate_html is true - CDG->>HG: generate(output_path, ...) - end - CDG->>CDG: _finalize_job() - CDG->>Job: complete() - CDG-->>Caller: DocumentationJob -``` - -### Stage 1 — Dependency Analysis - -Instantiates the backend `DocumentationGenerator` from [Documentation Generator](../documentation-generator/documentation-generator.md) (which internally wires up `DependencyGraphBuilder` and `AgentOrchestrator` from [Backend Core](../backend-core.md)) and calls `doc_generator.graph_builder.build_dependency_graph()`. The result is a map of `components` and a list of `leaf_nodes`. Statistics (`total_files_analyzed`, `leaf_nodes`) are recorded on the `DocumentationJob`. Failures are wrapped as `APIError("Dependency analysis failed: ...")`. - -### Stage 2 — Module Clustering - -Loads a cached `first_module_tree.json` if present, otherwise calls `cluster_modules(leaf_nodes, components, backend_config)` to invoke the LLM clustering model. The result is cached to disk (`first_module_tree_path`) and then persisted as the working module tree (`module_tree_path`). - -**Synthetic Module Patch**: If the module tree ends up empty despite having leaf nodes (to prevent an LLM "whole-repo" fallback that could exceed API context limits), the generator batches leaf nodes into synthetic modules of a configurable size (`CODEWIKI_MAX_FILES_PER_MODULE`, default `5`) and re-persists the tree. This safeguard applies even when loading from cache, closing a previously identified cache-bypass gap. - -The final module count is stored on the `DocumentationJob.module_count` field. - -### Stage 3 — Documentation Generation - -Calls `doc_generator.generate_module_documentation(components, leaf_nodes)`, which performs the topologically-ordered (leaf-first) generation of module documentation via the `AgentOrchestrator`. After generation: - -- `doc_generator.create_documentation_metadata(...)` writes `metadata.json`. -- Generated `.md` and `.json` files in the output directory are collected into `DocumentationJob.files_generated`. -- If `diagrams_dir` was configured, Mermaid diagrams embedded in generated markdown are extracted via `extract_and_save_mermaid_diagrams` and indexed with `create_diagrams_readme`. - -### Stage 4 — HTML Generation (Optional) - -If `generate_html=True`, `_run_html_generation()` uses the [HTML Generation](../html_generation/html_generation.md) module's `HTMLGenerator` to detect repository info (name, URL, GitHub Pages URL) and render `index.html`, auto-loading the module tree and metadata from the output directory. The generated file is appended to `DocumentationJob.files_generated`. - -### Stage 5 — Finalization - -`_finalize_job()` verifies that `metadata.json` exists in the output directory; if the backend did not already write it, the generator writes the job's own JSON representation (`DocumentationJob.to_json()`) as a fallback. - -## Configuration Mapping - -`generate()` builds the backend `Config` object (see [Config Core](../config-core.md)) via `Config.from_cli(...)`, translating the CLI's flat configuration dictionary into per-provider settings: - -```mermaid -flowchart LR - subgraph CLIConfig["CLI config dict"] - A1["main_model / cluster_model / fallback_model"] - A2["*_api_key"] - A3["*_base_url"] - A4["*_api_version"] - A5["*_max_tokens / *_temperature"] - A6["max_token_per_module / max_depth"] - A7["agent_instructions"] - A8["additional_paths"] - end - CLIConfig --> FromCLI["Config.from_cli()"] - FromCLI --> BackendConfig["Backend Config instance"] - BackendConfig --> DG2["DocumentationGenerator"] -``` - -Additional source paths supplied in the CLI config are normalized to absolute paths (resolved relative to `repo_path`) before being passed through as `additional_source_paths`. - -In verbose mode, the generator prints a detailed configuration summary (model names, base URLs, token limits, module settings, additional paths, and a preview of custom agent instructions) before kicking off Stage 1. - -## Progress Tracking and Logging - -The Generation module relies on utilities from the [Utils](../utils/utils.md) module: - -- **`ProgressTracker`**: manages a 5-stage weighted progress model (Dependency Analysis 40%, Module Clustering 20%, Documentation Generation 30%, HTML Generation 5%, Finalization 5%), providing `start_stage`, `update_stage`, `complete_stage`, elapsed-time formatting, and ETA estimation. -- **`CLILogger`**: complements verbose/non-verbose console output alongside `ProgressTracker`'s stage banners. - -Backend logs are captured by attaching a handler using `ColoredFormatter` (from `backend-core`'s Logging Config child module) directly to the `codewiki.src.be` logger, ensuring consistent colored output regardless of whether the CLI or backend emitted the log line, while preventing duplicate propagation to the root logger. - -## Error Handling - -All backend-facing calls (`build_dependency_graph`, `cluster_modules`, `generate_module_documentation`) are wrapped in `try`/`except` blocks that convert unexpected exceptions into `APIError` with contextual messages (e.g., `"Dependency analysis failed: ..."`). At the top level, `generate()` catches both `APIError` and generic `Exception`, marks the `DocumentationJob` as failed via `job.fail(str(e))`, and re-raises so the CLI layer can present the error to the user. - -## Relationship to Other Modules - -- **[CLI Core](../cli-core.md)**: parent module; Generation is one of its functional children alongside [Configuration](../configuration/configuration.md), [Job Models](../job_models/job_models.md), [Git Integration](../git_integration/git_integration.md), [HTML Generation](../html_generation/html_generation.md), and [Utils](../utils/utils.md). -- **[Job Models](../job_models/job_models.md)**: supplies `DocumentationJob`, `LLMConfig`, `JobStatus`, and `JobStatistics`, which Generation populates throughout the pipeline. -- **[HTML Generation](../html_generation/html_generation.md)**: invoked optionally at Stage 4 to render the documentation as a browsable HTML site. -- **[Utils](../utils/utils.md)**: provides `ProgressTracker` for stage-based progress reporting. -- **[Backend Core](../backend-core.md)**: supplies the actual documentation engine (`DocumentationGenerator`, `AgentOrchestrator`, `DependencyGraphBuilder`) that Generation orchestrates but does not reimplement. -- **[Config Core](../config-core.md)**: supplies the `Config` class used to translate CLI settings into the backend's runtime configuration. diff --git a/docs/reference/architecture/cli-core/git_integration.md b/docs/reference/architecture/cli-core/git_integration.md deleted file mode 100644 index ee7b3ac3..00000000 --- a/docs/reference/architecture/cli-core/git_integration.md +++ /dev/null @@ -1,147 +0,0 @@ -# Git Integration - -The Git Integration module provides the `GitManager` component, which encapsulates all git repository operations required by the CodeWiki CLI's documentation workflow. It is responsible for validating that a target directory is a git repository, inspecting working-directory cleanliness, creating dedicated documentation branches, committing generated documentation, and deriving useful metadata such as remote URLs, branch names, commit hashes, and GitHub pull-request links. - -This module is a child of the [CLI Core](../cli_core.md) module and is used primarily by the CLI's generation workflow (see [Generation](../generation/generation.md)) to safely manage version control state before, during, and after documentation is generated. - -## Purpose and Scope - -When the CLI is run with branch-creation options (e.g. `--create-branch`), it must: - -1. Confirm the target path is inside a valid git repository. -2. Ensure there are no uncommitted changes that could be silently mixed into a documentation commit (unless the user explicitly forces the operation). -3. Create a uniquely named, timestamped branch dedicated to documentation output. -4. Stage and commit the generated documentation files. -5. Surface repository metadata (current branch, commit hash, remote URL) back to the CLI so it can print helpful summaries. -6. Build a ready-to-open GitHub pull-request URL if the remote is hosted on GitHub. - -All of this logic is centralized in the `GitManager` class, keeping git operations isolated from the rest of the CLI so that other components (progress reporting, configuration, documentation generation) do not need to know about the underlying git library (`GitPython`). - -## Core Component: GitManager - -`GitManager` (`CodeWiki.codewiki.cli.git_manager::GitManager`) wraps a `git.Repo` instance (from the `GitPython` library) and exposes a small, purpose-built API surface consumed by the CLI layer. - -### Initialization - -On construction, `GitManager` resolves the given `repo_path` to an absolute path and attempts to open it as a git repository using `git.Repo(repo_path, search_parent_directories=True)`. The `search_parent_directories=True` flag allows the manager to locate a `.git` directory even if `repo_path` points to a subdirectory of the repository (similar to how the native `git` command walks upward from the current directory). - -If no valid repository is found, a `RepositoryError` is raised with actionable guidance (e.g., suggesting `git init`). `RepositoryError` is a CLI-specific exception type that carries a dedicated exit code (`EXIT_REPOSITORY_ERROR`), allowing the CLI's top-level error handler to report failures consistently and set correct process exit statuses. - -### Key Responsibilities - -#### 1. Working Directory Cleanliness Check - -`check_clean_working_directory()` inspects the repository for uncommitted modifications or untracked files using `repo.is_dirty(untracked_files=True)`. When dirty, it builds a human-readable summary listing up to three modified and three untracked files (with a "... and N more" suffix for longer lists), returning a `(is_clean, status_message)` tuple. This check underpins the safety guarantee that documentation generation will not accidentally commit unrelated in-progress changes. - -#### 2. Documentation Branch Creation - -`create_documentation_branch(force: bool = False)` creates and checks out a new branch named `docs/codewiki-` (format `%Y%m%d-%H%M%S`). Before creating the branch it re-runs the cleanliness check unless `force=True`, raising a detailed `RepositoryError` with copy-pastable remediation commands (`git status`, `git add -A && git commit`, `git stash`) if the working directory is dirty. - -To guard against timestamp collisions, it checks existing branch names and appends an incrementing numeric suffix (`-1`, `-2`, ...) if a collision would otherwise occur. Branch creation and checkout are performed via `repo.create_head(...)` followed by `.checkout()`; any underlying `GitCommandError` is translated into a `RepositoryError`. - -#### 3. Committing Documentation - -`commit_documentation(docs_path: Path, message: Optional[str] = None)` stages the documentation output directory via `repo.index.add([str(docs_path)])` and commits it with either a caller-supplied message or the default `"Add generated documentation\n\nGenerated by CodeWiki CLI"`. It returns the resulting commit's `hexsha`. Failures during add/commit are wrapped in `RepositoryError`. - -#### 4. Repository Metadata Accessors - -- `get_remote_url(remote_name="origin")` — returns the URL of the named remote, or `None` if it does not exist. -- `get_current_branch()` — returns the active branch name, or the literal string `"HEAD"` when in a detached-HEAD state (caught via `TypeError` from `GitPython`). -- `get_commit_hash()` — returns the current `HEAD` commit's hex SHA. -- `branch_exists(branch_name)` — returns whether a branch with the given name already exists locally. - -#### 5. GitHub Pull-Request URL Construction - -`get_github_pr_url(branch_name)` derives a ready-to-use GitHub "compare" URL (`/compare/`) from the `origin` remote, but only when that remote points to `github.com`. It normalizes the remote URL by: -- Stripping a trailing slash and any `.git` suffix. -- Converting SSH-style remotes (`git@github.com:org/repo`) into HTTPS form (`https://github.com/org/repo`). - -If there is no remote or it is not a GitHub remote, it returns `None`, allowing the CLI to gracefully skip printing a PR link for non-GitHub repositories. - -## Error Handling - -All git-related failures surface as `RepositoryError` (defined outside this module in the CLI's shared error utilities), rather than leaking raw `GitCommandError` or `InvalidGitRepositoryError` exceptions from `GitPython`. This gives the CLI a single, well-known exception type to catch at the top level and translate into a clean error message and a dedicated process exit code, keeping git-library implementation details out of user-facing error handling. - -## Architecture - -```mermaid -classDiagram - class GitManager { - +Path repo_path - +Repo repo - +__init__(repo_path) - +check_clean_working_directory() Tuple - +create_documentation_branch(force) str - +commit_documentation(docs_path, message) str - +get_remote_url(remote_name) str - +get_current_branch() str - +get_commit_hash() str - +branch_exists(branch_name) bool - +get_github_pr_url(branch_name) str - } - class RepositoryError { - +__init__(message) - } - GitManager ..> RepositoryError : raises -``` - -## Interaction with the CLI Generation Workflow - -`GitManager` is typically instantiated by the CLI's generation flow when branch-based documentation workflows are requested. The sequence below illustrates a typical end-to-end interaction: verifying repository state, creating a branch, running documentation generation (handled by `CLIDocumentationGenerator` in the [Generation](../generation/generation.md) module), committing the results, and reporting a PR link. - -```mermaid -sequenceDiagram - participant CLI as "CLI Entry Point" - participant GM as "GitManager" - participant DocGen as "CLIDocumentationGenerator" - participant Repo as "git.Repo" - - CLI->>GM: "GitManager(repo_path)" - GM->>Repo: "git.Repo(repo_path)" - CLI->>GM: "check_clean_working_directory()" - GM-->>CLI: "(is_clean, status_message)" - alt "Working directory dirty and not forced" - GM-->>CLI: "raise RepositoryError" - else "Clean or forced" - CLI->>GM: "create_documentation_branch(force)" - GM->>Repo: "create_head(branch_name)" - GM->>Repo: "checkout()" - GM-->>CLI: "branch_name" - CLI->>DocGen: "generate()" - DocGen-->>CLI: "DocumentationJob" - CLI->>GM: "commit_documentation(docs_path, message)" - GM->>Repo: "index.add([docs_path])" - GM->>Repo: "index.commit(message)" - GM-->>CLI: "commit_hash" - CLI->>GM: "get_github_pr_url(branch_name)" - GM-->>CLI: "pr_url or None" - end -``` - -## Process Flow: Documentation Branch Creation - -```mermaid -flowchart TD - Start["create_documentation_branch(force)"] --> CheckForce{{force?}} - CheckForce -->|"false"| CheckClean["check_clean_working_directory()"] - CheckForce -->|"true"| GenName["Generate timestamped branch name"] - CheckClean --> IsClean{{"is_clean?"}} - IsClean -->|"false"| RaiseErr["raise RepositoryError with remediation steps"] - IsClean -->|"true"| GenName - GenName --> CheckExists{{"branch name exists?"}} - CheckExists -->|"true"| AppendCounter["Append incrementing counter suffix"] - AppendCounter --> CheckExists - CheckExists -->|"false"| CreateBranch["repo.create_head(branch_name)"] - CreateBranch --> Checkout["new_branch.checkout()"] - Checkout --> ReturnName["return branch_name"] -``` - -## Relationship to Other Modules - -- **[CLI Core](../cli_core.md)** — the parent module; `GitManager` is one of the core services orchestrated alongside configuration management, job models, and progress tracking to deliver the full `codewiki generate` experience. -- **[Generation](../generation/generation.md)** — `CLIDocumentationGenerator` performs the actual documentation generation (dependency analysis, clustering, LLM-driven writing, optional HTML output). The CLI entry point coordinates `GitManager` and `CLIDocumentationGenerator` together: creating a branch before generation, and committing the resulting files afterward. -- **[Job Models](../job_models/job_models.md)** — while `GitManager` itself does not depend on `DocumentationJob`, the surrounding CLI workflow often records git-derived metadata (branch name, commit hash) alongside job statistics for reporting purposes. - -## Dependencies - -`GitManager` depends on the third-party [`GitPython`](https://gitpython.readthedocs.io/) library (imported as `git`) for all low-level repository interactions, including `git.Repo`, `git.InvalidGitRepositoryError`, and `git.exc.GitCommandError`. It also depends on the CLI's shared `RepositoryError` exception type for consistent error reporting across the CLI. diff --git a/docs/reference/architecture/cli-core/html_generation.md b/docs/reference/architecture/cli-core/html_generation.md deleted file mode 100644 index 40f0eb9d..00000000 --- a/docs/reference/architecture/cli-core/html_generation.md +++ /dev/null @@ -1,131 +0,0 @@ -# Html Generation - -The Html Generation module produces a self-contained, static HTML documentation viewer suitable for GitHub Pages (or any static file host). It is the final, optional presentation layer of the CodeWiki CLI pipeline: after the [Generation](../generation/generation.md) stage has produced Markdown documentation files, a module tree, and metadata, the Html Generation module packages that output into a single browsable `index.html` file with embedded configuration, styles, and client-side rendering logic. - -## Purpose and Scope - -The `HTMLGenerator` class is the sole core component of this module. Its responsibilities are: - -- **Template loading** — reads a static HTML template (`viewer_template.html`) shipped with the package. -- **Data discovery** — auto-loads `module_tree.json` and `metadata.json` from a documentation output directory when explicit data is not supplied. -- **Placeholder substitution** — injects title, repository link, embedded JSON data (module tree, metadata, config), and an "info panel" HTML fragment into the template. -- **Repository introspection** — inspects a local git repository to derive a display name, remote URL, and a predicted GitHub Pages URL. -- **Atomic file output** — writes the final HTML file safely using the shared filesystem utilities. - -This module has no knowledge of *how* documentation content was produced; it only consumes the artifacts (`module_tree.json`, `metadata.json`, generated Markdown files) that the [Generation](../generation/generation.md) module and the backend documentation pipeline produce. This keeps Html Generation a pure "rendering/packaging" concern, decoupled from LLM orchestration, dependency analysis, and job tracking. - -## Core Component - -### HTMLGenerator - -`HTMLGenerator` (`codewiki/cli/html_generator.py`) encapsulates all HTML viewer generation logic. - -| Method | Responsibility | -|---|---| -| `__init__(template_dir)` | Resolves the template directory, defaulting to the package's `templates/github_pages` folder. | -| `load_module_tree(docs_dir)` | Reads `module_tree.json` from the docs directory; falls back to a minimal single-node structure if the file is missing. | -| `load_metadata(docs_dir)` | Reads `metadata.json`; returns `None` (non-critical) if missing or unparsable. | -| `generate(...)` | Orchestrates the full generation: auto-loads data, builds the info panel, computes paths/links, serializes JSON, performs placeholder substitution, and writes the output file. | -| `_build_info_content(metadata)` | Builds an HTML fragment (model name, generation timestamp, commit hash, component count, max depth) displayed in the viewer's info panel. | -| `_escape_html(text)` | Escapes HTML-sensitive characters to prevent malformed markup when embedding user/repo-derived strings (e.g., title). | -| `detect_repository_info(repo_path)` | Uses `GitPython` to read the repository name, normalize the remote URL (including `git@github.com:` SSH URLs), and compute the expected `https://.github.io//` Pages URL. | - -## Architecture - -```mermaid -flowchart TD - Docs["Documentation Output Directory"] --> ModuleTreeJson["module_tree.json"] - Docs --> MetadataJson["metadata.json"] - Template["viewer_template.html"] --> Generator["HTMLGenerator"] - ModuleTreeJson --> Generator - MetadataJson --> Generator - RepoPath["Repository Path"] --> DetectInfo["detect_repository_info()"] - DetectInfo --> Generator - Generator -->|"safe_write()"| IndexHtml["index.html"] - Generator -->|"on failure"| FSError["FileSystemError"] -``` - -### Dependencies - -- **`codewiki.cli.utils.fs`** — `safe_read` / `safe_write` provide atomic, encoding-safe file I/O used to load the template/JSON files and write the final HTML output. See the [Utils](../utils/utils.md) module for other shared CLI utilities such as logging and progress tracking. -- **`codewiki.cli.utils.errors::FileSystemError`** — raised when the template file is missing or when reading/writing fails, allowing the CLI layer to surface a consistent error type. -- **`git` (GitPython)** — used only within `detect_repository_info` to introspect the local repository; failures are caught and silently ignored, degrading gracefully to a viewer without repository links. - -## Integration with the CLI Pipeline - -Html Generation is invoked as an optional, final stage by [`CLIDocumentationGenerator`](../generation/generation.md) (from the [Generation](../generation/generation.md) module), which drives the overall CLI workflow: dependency analysis → module clustering → documentation generation → **HTML generation** → job finalization. - -```mermaid -sequenceDiagram - participant CLIGen as "CLIDocumentationGenerator" - participant HTMLGen as "HTMLGenerator" - participant FS as "safe_read / safe_write" - participant Git as "GitPython" - - CLIGen->>HTMLGen: HTMLGenerator() - CLIGen->>HTMLGen: detect_repository_info(repo_path) - HTMLGen->>Git: Repo(repo_path) - Git-->>HTMLGen: remote URL, name - HTMLGen-->>CLIGen: name, url, github_pages_url - CLIGen->>HTMLGen: generate(output_path, title, repository_url, docs_dir) - HTMLGen->>FS: safe_read(module_tree.json) - HTMLGen->>FS: safe_read(metadata.json) - HTMLGen->>FS: safe_read(viewer_template.html) - HTMLGen->>HTMLGen: _build_info_content(metadata) - HTMLGen->>HTMLGen: substitute placeholders - HTMLGen->>FS: safe_write(index.html) - HTMLGen-->>CLIGen: index.html written -``` - -This corresponds to the `_run_html_generation` step inside `CLIDocumentationGenerator.generate()`: it is only executed when the CLI was invoked with `generate_html=True`, after the backend has produced Markdown files, `module_tree.json`, and `metadata.json` in the output directory. On success, `"index.html"` is appended to `DocumentationJob.files_generated` (see the [Job Models](../job_models/job_models.md) module). - -## Generation Flow - -```mermaid -flowchart TD - Start["generate() called"] --> CheckDocsDir{{"docs_dir provided?"}} - CheckDocsDir -->|"yes"| AutoLoadTree["load_module_tree(docs_dir)"] - CheckDocsDir -->|"yes"| AutoLoadMeta["load_metadata(docs_dir)"] - CheckDocsDir -->|"no"| UseProvided["use provided module_tree/metadata"] - AutoLoadTree --> Defaults - AutoLoadMeta --> Defaults - UseProvided --> Defaults - Defaults["Apply defaults for module_tree/config"] --> LoadTemplate["Load viewer_template.html"] - LoadTemplate -->|"missing"| RaiseErr["raise FileSystemError"] - LoadTemplate -->|"found"| BuildInfo["_build_info_content(metadata)"] - BuildInfo --> BuildRepoLink["Build repository link HTML"] - BuildRepoLink --> ComputeBasePath["Compute docs_base_path"] - ComputeBasePath --> SerializeJson["Serialize config/module_tree/metadata to JSON"] - SerializeJson --> Replace["Replace template placeholders"] - Replace --> WriteOut["safe_write(output_path)"] - WriteOut --> Done["index.html generated"] -``` - -### Template Placeholders - -The generator performs a straightforward string substitution over the template file. The following placeholders are populated by `generate()`: - -| Placeholder | Source | -|---|---| -| `{{TITLE}}` | Escaped `title` argument (e.g., repository name) | -| `{{REPO_LINK}}` | HTML anchor to `repository_url`, empty if not provided | -| `{{SHOW_INFO}}` | `"block"` or `"none"` depending on whether info content was built | -| `{{INFO_CONTENT}}` | HTML fragment from `_build_info_content` (model, timestamp, commit, stats) | -| `{{CONFIG_JSON}}` | JSON-serialized `config` dictionary | -| `{{MODULE_TREE_JSON}}` | JSON-serialized module tree structure | -| `{{METADATA_JSON}}` | JSON-serialized metadata, or the literal `null` | -| `{{DOCS_BASE_PATH}}` | Relative path from the output file to the docs directory | - -## Error Handling - -- Missing `viewer_template.html` raises `FileSystemError`, propagated up to the CLI layer. -- Missing or malformed `module_tree.json` is either substituted with a minimal fallback structure (`load_module_tree`) or, on unexpected read/parse errors, raises `FileSystemError`. -- Missing or malformed `metadata.json` is treated as non-critical: `load_metadata` swallows exceptions and returns `None`, resulting in the info panel being hidden (`{{SHOW_INFO}} = "none"`). -- Git introspection failures in `detect_repository_info` are caught broadly, so the generator always returns a usable (if partially empty) info dictionary rather than failing the whole pipeline. - -## Relationship to Other Modules - -- **[Generation](../generation/generation.md)** — the `CLIDocumentationGenerator` orchestrates when and how `HTMLGenerator` is invoked as part of the overall documentation generation job. -- **[Job Models](../job_models/job_models.md)** — the generated `index.html` file path is recorded in the `DocumentationJob.files_generated` list. -- **[Utils](../utils/utils.md)** — shares low-level filesystem helpers (`safe_read`/`safe_write`) and error types used throughout the CLI. -- **[cli-core](../cli-core.md)** — the parent module that aggregates Html Generation alongside generation, configuration, job models, git integration, and shared utilities into the full CLI toolchain. diff --git a/docs/reference/architecture/cli-core/job_models.md b/docs/reference/architecture/cli-core/job_models.md deleted file mode 100644 index 7b3e33bd..00000000 --- a/docs/reference/architecture/cli-core/job_models.md +++ /dev/null @@ -1,244 +0,0 @@ -# Job Models - -## Introduction - -The Job Models module defines the core data structures used to represent, track, and persist documentation generation jobs within the CodeWiki CLI. It is a foundational, dependency-free module that provides typed dataclasses and enums for job state, configuration, and statistics — enabling consistent serialization, deserialization, and status tracking across the CLI documentation pipeline. - -This module is a child of the [CLI Core](../cli_core.md) module and is consumed by sibling modules such as [Generation](../generation/generation.md), [Configuration](../configuration/configuration.md), [Utils](../utils/utils.md), and other CLI orchestration components that need to create, update, or persist job records. - -## Purpose and Scope - -The Job Models module has a single, focused responsibility: **define the shape of a documentation job and its lifecycle**. It does not perform any I/O, orchestration, or business logic beyond simple state transitions and serialization helpers. This keeps the module lightweight, easily testable, and safe to import from anywhere in the CLI codebase without introducing circular dependencies. - -Key responsibilities: -- Define the `JobStatus` enum representing the lifecycle states of a job. -- Define `GenerationOptions` for user-configurable generation behavior (branching, GitHub Pages, caching, output paths). -- Define `LLMConfig` for capturing which language models and endpoint were used for a run. -- Define `JobStatistics` for capturing quantitative results of a run (files analyzed, tree depth, tokens consumed). -- Define `DocumentationJob`, the aggregate root that ties together identity, timing, status, configuration, and results for a single documentation generation run. -- Provide robust `to_dict()` / `to_json()` / `from_dict()` methods for safe persistence and rehydration, including defensive coercion of malformed or partial data. - -## Core Components - -### JobStatus - -`JobStatus` is a `str`-backed `Enum` representing the possible lifecycle states of a documentation job: - -| Value | Meaning | -|---|---| -| `pending` | Job has been created but not yet started | -| `running` | Job is actively executing | -| `completed` | Job finished successfully | -| `failed` | Job terminated with an error | - -Because it inherits from both `str` and `Enum`, instances serialize naturally to plain strings in JSON output while still supporting type-safe comparisons in code (e.g., `job.status == JobStatus.RUNNING`). - -### GenerationOptions - -`GenerationOptions` is a dataclass capturing user-facing flags that control how a documentation run behaves: - -- `create_branch` (`bool`): Whether to create a dedicated git branch for the generated docs. Consumed by the [Git Integration](../git_integration/git_integration.md) module. -- `github_pages` (`bool`): Whether to prepare output for GitHub Pages publishing. -- `no_cache` (`bool`): Whether to bypass any caching layer during generation. -- `custom_output` (`Optional[str]`): An optional override for the output directory. - -### LLMConfig - -`LLMConfig` records which language models and endpoint were used to produce a job's output: - -- `main_model` (`str`): The primary model used for documentation generation. -- `cluster_model` (`str`): The model used for clustering/module grouping decisions. -- `base_url` (`str`): The API base URL for the LLM provider. - -This is a plain, required-field dataclass (no defaults), reflecting that when an `LLMConfig` is attached to a job, all three fields are expected to be known. - -### JobStatistics - -`JobStatistics` aggregates quantitative metrics collected during a documentation run: - -- `total_files_analyzed` (`int`): Count of source files processed. -- `leaf_nodes` (`int`): Count of leaf-level components/modules identified. -- `max_depth` (`int`): Maximum depth of the module hierarchy produced. -- `total_tokens_used` (`int`): Total LLM tokens consumed for the run. - -All fields default to `0`, so a freshly created job has a valid, zeroed-out statistics object even before execution begins. - -### DocumentationJob - -`DocumentationJob` is the central aggregate of this module. It represents a single documentation generation run end-to-end, combining identity, repository context, git metadata, timing, status, and nested configuration/statistics objects. - -**Fields:** - -| Field | Type | Description | -|---|---|---| -| `job_id` | `str` | UUID4 identifier, auto-generated if not supplied | -| `repository_path` | `str` | Absolute filesystem path to the repository being documented | -| `repository_name` | `str` | Human-readable repository name | -| `output_directory` | `str` | Destination directory for generated docs | -| `commit_hash` | `str` | Git commit SHA the job was run against | -| `branch_name` | `Optional[str]` | Git branch name, if applicable | -| `timestamp_start` | `str` | ISO-format start timestamp, auto-populated | -| `timestamp_end` | `Optional[str]` | ISO-format end timestamp, set on completion/failure | -| `status` | `JobStatus` | Current lifecycle status | -| `error_message` | `Optional[str]` | Populated when the job fails | -| `files_generated` | `List[str]` | Paths of documentation files produced | -| `module_count` | `int` | Number of modules documented | -| `generation_options` | `GenerationOptions` | Options selected for this run | -| `llm_config` | `Optional[LLMConfig]` | LLM configuration used, if known | -| `statistics` | `JobStatistics` | Quantitative results of the run | - -**Lifecycle methods:** - -- `start()` — Transitions status to `RUNNING` and refreshes `timestamp_start`. -- `complete()` — Transitions status to `COMPLETED` and sets `timestamp_end`. -- `fail(error_message)` — Transitions status to `FAILED`, records the error, and sets `timestamp_end`. - -**Serialization methods:** - -- `to_dict()` — Produces a JSON-serializable `dict` representation, flattening nested dataclasses (`GenerationOptions`, `LLMConfig`, `JobStatistics`) into plain dicts and converting `JobStatus` to its string value. -- `to_json()` — Convenience wrapper around `to_dict()` that returns a pretty-printed JSON string (2-space indent). -- `from_dict(data)` — Classmethod that reconstructs a `DocumentationJob` from a raw dictionary (e.g., loaded from a JSON file), using defensive coercion helpers to tolerate missing, malformed, or partial fields. - -### Defensive Coercion Helpers - -The module defines several private module-level helper functions used exclusively by `DocumentationJob.from_dict()` to safely rebuild nested objects from untrusted or partial data: - -- `_coerce_job_status(value, default)` — Converts a raw value to a valid `JobStatus`, falling back to `JobStatus.PENDING` (or a supplied default) if the value is `None` or not a recognized status string. -- `_coerce_int(value, default)` — Safely converts a value to `int`, returning a default (`0`) on `None`, `TypeError`, or `ValueError`. -- `_coerce_generation_options(value)` — Rebuilds a `GenerationOptions` from a dict, defaulting missing keys. -- `_coerce_llm_config(value)` — Rebuilds an `LLMConfig` from a dict, or returns `None` if no data is present. -- `_coerce_statistics(value)` — Rebuilds a `JobStatistics` from a dict, defaulting missing/invalid numeric fields to `0`. - -This coercion layer makes `DocumentationJob.from_dict()` resilient to schema drift (e.g., loading job files written by an older version of the CLI) without raising exceptions during deserialization. - -## Architecture - -### Component Structure - -```mermaid -classDiagram - class JobStatus { - <> - PENDING - RUNNING - COMPLETED - FAILED - } - class GenerationOptions { - +bool create_branch - +bool github_pages - +bool no_cache - +str custom_output - } - class LLMConfig { - +str main_model - +str cluster_model - +str base_url - } - class JobStatistics { - +int total_files_analyzed - +int leaf_nodes - +int max_depth - +int total_tokens_used - } - class DocumentationJob { - +str job_id - +str repository_path - +str repository_name - +str output_directory - +str commit_hash - +str branch_name - +str timestamp_start - +str timestamp_end - +JobStatus status - +str error_message - +List files_generated - +int module_count - +GenerationOptions generation_options - +LLMConfig llm_config - +JobStatistics statistics - +start() - +complete() - +fail(error_message) - +to_dict() dict - +to_json() str - +from_dict(data) DocumentationJob - } - DocumentationJob --> JobStatus : "status" - DocumentationJob --> GenerationOptions : "generation_options" - DocumentationJob --> LLMConfig : "llm_config (optional)" - DocumentationJob --> JobStatistics : "statistics" -``` - -### Job Lifecycle State Machine - -```mermaid -stateDiagram-v2 - [*] --> PENDING : "DocumentationJob() created" - PENDING --> RUNNING : "start()" - RUNNING --> COMPLETED : "complete()" - RUNNING --> FAILED : "fail(error_message)" - COMPLETED --> [*] - FAILED --> [*] -``` - -### Serialization / Deserialization Flow - -```mermaid -sequenceDiagram - participant Caller as "CLI Component" - participant Job as "DocumentationJob" - participant Coerce as "Coercion Helpers" - - Caller->>Job: "DocumentationJob(...)" - Job-->>Caller: "job instance (status=PENDING)" - Caller->>Job: "job.start()" - Caller->>Job: "job.to_dict() / job.to_json()" - Job-->>Caller: "dict / JSON string" - Note over Caller: "Persisted to disk or transmitted" - - Caller->>Job: "DocumentationJob.from_dict(raw_data)" - Job->>Coerce: "_coerce_job_status(raw_data.status)" - Job->>Coerce: "_coerce_int(raw_data.module_count)" - Job->>Coerce: "_coerce_generation_options(raw_data.generation_options)" - Job->>Coerce: "_coerce_llm_config(raw_data.llm_config)" - Job->>Coerce: "_coerce_statistics(raw_data.statistics)" - Coerce-->>Job: "validated nested objects" - Job-->>Caller: "reconstructed DocumentationJob" -``` - -## Integration with Other Modules - -The Job Models module is intentionally dependency-free (it only imports from the Python standard library), which allows it to be safely imported by any layer of the CLI without risk of circular imports: - -- **[CLI Core](../cli_core.md)** (parent module): Exposes `DocumentationJob`, `GenerationOptions`, `JobStatistics`, `JobStatus`, and `LLMConfig` as part of its public component surface, making them available to all CLI subsystems. -- **[Generation](../generation/generation.md)**: The `CLIDocumentationGenerator` creates and drives `DocumentationJob` instances through their lifecycle (`start()` → `complete()`/`fail()`), populating `files_generated`, `module_count`, and `statistics` as generation proceeds. -- **[Configuration](../configuration/configuration.md)**: `ConfigManager` and the `Configuration`/`AgentInstructions` models supply values (such as model names and base URLs) that are captured into a job's `LLMConfig`. -- **[Git Integration](../git_integration/git_integration.md)**: `GitManager` supplies `commit_hash` and `branch_name` values and consults `GenerationOptions.create_branch` to decide whether to create a dedicated branch for generated documentation. -- **[Utils](../utils/utils.md)**: `CLILogger` and `ProgressTracker`/`ModuleProgressBar` report on job progress using the status and statistics captured on `DocumentationJob` instances. -- **[HTML Generation](../html_generation/html_generation.md)**: `HTMLGenerator` may reference `files_generated` and `output_directory` from a completed job when producing rendered output. - -```mermaid -flowchart TD - JobModels["Job Models"] - CliCore["CLI Core (parent)"] - Generation["Generation"] - Configuration["Configuration"] - GitIntegration["Git Integration"] - HtmlGeneration["HTML Generation"] - Utils["Utils"] - - CliCore --> JobModels - Generation -->|"creates & drives"| JobModels - Configuration -->|"populates LLMConfig"| JobModels - GitIntegration -->|"populates commit_hash / branch_name"| JobModels - HtmlGeneration -->|"reads files_generated"| JobModels - Utils -->|"reports on status/statistics"| JobModels -``` - -## Design Notes - -- **Immutability of shape, mutability of state**: While the dataclasses themselves are mutable (no `frozen=True`), the module encourages controlled mutation through explicit lifecycle methods (`start()`, `complete()`, `fail()`) rather than direct field assignment, keeping status transitions predictable and auditable. -- **Safe round-tripping**: The combination of `to_dict()`/`to_json()` and `from_dict()` with defensive coercion helpers ensures that job records can be persisted to disk (e.g., as job history or resumable state) and reloaded reliably even if the on-disk schema is incomplete or slightly out of date. -- **String-backed enum**: `JobStatus` extending `str` means job status values serialize directly as readable strings in JSON without requiring custom encoders, simplifying integration with logging, CLI output, and external tooling. -- **Zero external dependencies**: The module relies only on `dataclasses`, `datetime`, `typing`, `enum`, `uuid`, and `json` from the standard library, making it trivially testable and reusable across CLI contexts. diff --git a/docs/reference/architecture/config-core/config-core.md b/docs/reference/architecture/config-core/config-core.md deleted file mode 100644 index 8b26d412..00000000 --- a/docs/reference/architecture/config-core/config-core.md +++ /dev/null @@ -1,220 +0,0 @@ -# Config Core - -The Config Core module defines the central `Config` dataclass that CodeWiki uses to drive every stage of the documentation-generation pipeline. It is the single source of truth for repository paths, output directories, LLM provider settings (models, API keys, base URLs, temperatures, and token limits), and agent-instruction customization such as include/exclude file patterns, focus modules, documentation type, and free-form custom instructions. - -Config Core is intentionally small and dependency-light: it contains one component, `Config`, but it is consumed heavily by [Backend Core](backend-core.md) (which executes the actual analysis and generation pipeline), [CLI Core](cli-core.md) (which builds a `Config` from a locally persisted configuration file), and [Frontend Core](frontend-core.md) (which builds a `Config` for each web-submitted documentation job). Because nearly every other module depends on `Config`, understanding its shape and construction paths is a prerequisite for understanding the rest of the system. - -## Purpose and Responsibilities - -`Config` serves three distinct responsibilities: - -1. **Data container** — holds all repository, output, and per-provider LLM settings (cluster/main/fallback models, keys, base URLs, API versions, max tokens, temperatures, and the field name used to send max-tokens to the provider). -2. **Construction factory** — offers four classmethod entry points (`from_args`, `from_web_job`, `from_cli`, `from_config_manager`) that adapt different callers' inputs (argparse namespaces, web job parameters, explicit CLI parameters, or a `ConfigManager`) into a fully validated `Config` instance. -3. **Derived-behavior provider** — exposes properties and helper methods (`include_patterns`, `exclude_patterns`, `focus_modules`, `doc_type`, `custom_instructions`, `all_source_paths`, `is_multi_path_mode`, `get_prompt_addition`, `validate_source_paths`) that downstream consumers use instead of re-implementing agent-instruction parsing or multi-path logic themselves. - -## Core Component - -### Config (`codewiki/src/config.py`) - -`Config` is a `@dataclass` with required fields (`repo_path`, `output_dir`, `dependency_graph_dir`, `docs_dir`, `max_depth`, `main_model`, `cluster_model`, `fallback_model`, `cluster_api_key`, `main_api_key`, `fallback_api_key`) and a large set of optional fields with sensible defaults for base URLs, API versions, max-tokens, temperatures, and multi-path/diagram directory support. - -```mermaid -classDiagram - class Config { - +str repo_path - +str output_dir - +str dependency_graph_dir - +str docs_dir - +int max_depth - +str main_model - +str cluster_model - +str fallback_model - +str cluster_api_key - +str main_api_key - +str fallback_api_key - +Optional~str~ cluster_base_url - +Optional~str~ main_base_url - +Optional~str~ fallback_base_url - +int cluster_max_tokens - +int main_max_tokens - +int fallback_max_tokens - +int max_token_per_module - +int max_token_per_leaf_module - +float cluster_temperature - +float main_temperature - +float fallback_temperature - +Optional~Dict~ agent_instructions - +Optional~str~ diagrams_dir - +Optional~List~str~~ additional_source_paths - +to_dict(include_secrets) Dict - +from_dict(data)$ Config - +include_patterns() Optional~List~ - +exclude_patterns() Optional~List~ - +focus_modules() Optional~List~ - +doc_type() Optional~str~ - +custom_instructions() Optional~str~ - +all_source_paths() List~str~ - +validate_source_paths() void - +is_multi_path_mode() bool - +get_prompt_addition() str - +from_args(args)$ Config - +from_web_job(repo_path, docs_dir)$ Config - +from_cli(...)$ Config - +from_config_manager(manager, repo_path, output_dir)$ Config - } -``` - -## Field Groups - -| Group | Fields | Purpose | -|---|---|---| -| Paths | `repo_path`, `output_dir`, `dependency_graph_dir`, `docs_dir`, `diagrams_dir` | Where source code is read from and where analysis artifacts and docs are written | -| Multi-path support | `additional_source_paths` | Allows analyzing multiple source directories as one unified documentation set | -| Model selection | `main_model`, `cluster_model`, `fallback_model` | Which LLM is used for generation, clustering, and fallback | -| Per-provider credentials | `cluster_api_key`, `main_api_key`, `fallback_api_key` | Required, runtime-only secrets — never serialized by `to_dict()` unless explicitly requested | -| Per-provider connectivity | `*_base_url`, `*_api_version` | Endpoint configuration for each provider | -| Per-provider limits | `*_max_tokens`, `*_max_token_field`, `max_token_per_module`, `max_token_per_leaf_module` | Token budgeting for generation and clustering | -| Per-provider sampling | `*_temperature`, `*_temperature_supported` | Controls determinism/creativity per provider, with a flag for providers that reject custom temperatures | -| Agent customization | `agent_instructions` | Dict-based include/exclude patterns, focus modules, doc type, and custom instructions consumed via properties | - -## Secret Handling: `to_dict` / `from_dict` - -`Config` maintains a frozen set of runtime-only secret field names (`cluster_api_key`, `main_api_key`, `fallback_api_key`). `to_dict()` strips these by default so that serialized configuration (e.g., cached to disk or logged) never leaks API keys; callers must pass `include_secrets=True` to get a fully round-trippable dict. `from_dict()` reconstructs a `Config` by filtering the input to only known dataclass fields, so extra keys are safely ignored, but if secrets were stripped they must be supplied separately or construction raises a `TypeError` (missing required field). - -```mermaid -flowchart TD - A["Config instance"] --> B["to_dict(include_secrets=False)"] - B --> C["Plain dict, no API keys"] - A --> D["to_dict(include_secrets=True)"] - D --> E["Plain dict, includes API keys"] - C --> F["from_dict(data)"] - E --> F - F -->|"missing secrets"| G["TypeError: missing required field"] - F -->|"secrets present"| H["Reconstructed Config"] -``` - -## Agent Instructions Properties - -`agent_instructions` is an optional dict (or an object exposing `to_dict()`, e.g. `AgentInstructions` from [CLI Core](cli-core.md)). Five read-only properties expose its contents without requiring callers to know its internal shape: - -- `include_patterns` — file glob patterns to include in analysis -- `exclude_patterns` — file glob patterns to exclude -- `focus_modules` — module names that should receive more detailed documentation -- `doc_type` — one of `api`, `architecture`, `user-guide`, `developer`, or a free-form string -- `custom_instructions` — free-form additional guidance text - -`get_prompt_addition()` combines `doc_type`, `focus_modules`, and `custom_instructions` into a single prompt-ready string, escaping curly braces in `custom_instructions` (via `escape_format_braces`) so JSON-like content does not break downstream `.format()` calls in the generation pipeline consumed by [Backend Core](backend-core.md). - -```mermaid -flowchart TD - AI["agent_instructions dict"] --> IP["include_patterns"] - AI --> EP["exclude_patterns"] - AI --> FM["focus_modules"] - AI --> DT["doc_type"] - AI --> CI["custom_instructions"] - DT --> GPA["get_prompt_addition()"] - FM --> GPA - CI -->|"escape_format_braces"| GPA - GPA --> Prompt["Combined prompt-addition string"] -``` - -## Multi-Path Source Support - -`additional_source_paths` enables analyzing more than one directory as a single logical repository. `all_source_paths` always returns `repo_path` as the first absolute path, followed by any additional paths. `validate_source_paths()` raises `ValueError`/`OSError` if any path is missing, not a directory, or unreadable. `is_multi_path_mode()` is a simple boolean check used by the analysis pipeline in [Backend Core](backend-core.md) to decide whether to merge multiple source trees. - -```mermaid -flowchart TD - Start["Config.validate_source_paths()"] --> CheckPrimary{{"repo_path exists and is dir?"}} - CheckPrimary -->|"no"| Err1["raise ValueError"] - CheckPrimary -->|"yes"| HasAdditional{{"additional_source_paths set?"}} - HasAdditional -->|"no"| Done["Validation OK single-path mode"] - HasAdditional -->|"yes"| Loop["For each additional path"] - Loop --> CheckExists{{"path exists and is dir?"}} - CheckExists -->|"no"| Err2["raise ValueError"] - CheckExists -->|"yes"| CheckRead{{"path readable?"}} - CheckRead -->|"no"| Err3["raise OSError"] - CheckRead -->|"yes"| Loop - Loop --> Done2["Validation OK multi-path mode"] -``` - -## Construction Paths - -`Config` provides four classmethods for building an instance, each tailored to a different caller in the system. - -```mermaid -flowchart TD - subgraph CLIFlow["CLI Core entry point"] - CM["ConfigManager: persisted JSON plus keyring"] - CM -->|"from_config_manager"| FC1["Config.from_cli(...)"] - end - subgraph WebFlow["Frontend Core entry point"] - BW["BackgroundWorker for web job"] - BW -->|"from_web_job"| FA1["Config.from_args wrapping Namespace"] - end - subgraph EnvFlow["Environment-driven CLI entry point"] - ArgParse["argparse.Namespace from CLI arguments"] - ArgParse -->|"from_args"| FA2["Reads MAIN_MODEL, FALLBACK_MODEL, CLUSTER_API_KEY, MAIN_API_KEY, FALLBACK_API_KEY from environment"] - end - subgraph DirectFlow["Direct parameter entry point"] - Caller["Any caller with explicit parameters"] - Caller -->|"from_cli"| Validate["Validation block: keys, urls, types, ranges, max_token_field enum"] - Validate --> VSP["validate_source_paths()"] - VSP --> Instance["Config instance"] - end - FA1 --> FA2 - FC1 --> Validate -``` - -### `from_args(args)` - -Builds a `Config` purely from environment variables (`MAIN_MODEL`, `CLUSTER_MODEL`, `LLM_BASE_URL`, `FALLBACK_MODEL`, `CLUSTER_API_KEY`, `MAIN_API_KEY`, `FALLBACK_API_KEY`) plus the `repo_path` supplied on `args`. It computes a sanitized repo name for the docs output directory and raises `ValueError` if `FALLBACK_MODEL` or any per-provider API key is missing. This is the lowest-level, environment-driven constructor. - -### `from_web_job(repo_path, docs_dir)` - -A thin wrapper used by [Frontend Core](frontend-core.md)'s background worker. It delegates to `from_args` (wrapping `repo_path` in a synthetic `argparse.Namespace`) and then overrides `docs_dir` with the job-specific output directory, avoiding the need to fabricate a fake CLI namespace at the call site. - -### `from_cli(...)` - -The most comprehensive constructor, accepting every field explicitly (models, keys, base URLs, API versions, token limits, temperatures, max-token field names, `agent_instructions`, `diagrams_dir`, `additional_source_paths`). It performs an extensive validation block before construction: - -- Required, non-empty API keys and base URLs for all three providers -- Type coercion and validation for token limits (`int`) and temperatures (`float`) -- Range checks: token limits must be positive, temperatures must be within `0.0`–`2.0` -- Enum checks: `*_max_token_field` must be `max_tokens` or `max_completion_tokens` - -After construction, it calls `validate_source_paths()` to ensure the repository and any additional paths actually exist and are accessible before returning the instance. - -### `from_config_manager(manager, repo_path, output_dir)` - -Used by [CLI Core](cli-core.md) to bridge its persisted `ConfigManager`/`Configuration` model into a runtime `Config`. It pulls the loaded `Configuration` object and per-provider API keys from the `ConfigManager`, validates that models and keys are present (raising actionable `ValueError`s referencing the `codewiki config set` command), extracts `additional_source_paths` from `agent_instructions` if present, and finally delegates to `from_cli(...)` with all fields populated from the manager. - -```mermaid -sequenceDiagram - participant CLI as "CLI Core (ConfigManager)" - participant Config as "Config.from_config_manager" - participant FromCli as "Config.from_cli" - participant Validate as "validate_source_paths" - - CLI->>Config: from_config_manager(manager, repo_path, output_dir) - Config->>CLI: manager.get_config() - Config->>CLI: get_cluster_api_key / get_main_api_key / get_fallback_api_key - Config->>Config: check models and keys are present - Config->>FromCli: from_cli(repo_path, output_dir, models, keys, urls, tokens, temps) - FromCli->>FromCli: validation block keys urls types ranges enums - FromCli->>Validate: validate_source_paths() - Validate-->>FromCli: OK or raises ValueError or OSError - FromCli-->>Config: Config instance - Config-->>CLI: Config instance -``` - -## Integration with Other Modules - -- **[Backend Core](backend-core.md)** — the `AgentOrchestrator`, `DocumentationGenerator`, and dependency-analyzer components consume a fully constructed `Config` for repo paths, model selection, token limits, and prompt additions (`get_prompt_addition()`), and use `all_source_paths()`/`is_multi_path_mode()` to drive multi-path analysis. -- **[CLI Core](cli-core.md)** — `ConfigManager` and the `Configuration`/`AgentInstructions` models persist user settings to disk; `Config.from_config_manager` bridges that persisted state into the runtime `Config` used for a documentation run. -- **[Frontend Core](frontend-core.md)** — `BackgroundWorker` builds a `Config` per submitted job via `Config.from_web_job`, using job-specific `repo_path` and `docs_dir` values while relying on environment-configured models and keys. - -## Design Rationale - -- **Secrets never leak by default.** The `_RUNTIME_ONLY_SECRET_FIELDS` frozenset and the `include_secrets` flag on `to_dict()` ensure API keys are excluded from any dict representation used for caching, logging, or persistence, unless a caller explicitly opts in for same-process reconstruction. -- **Fail fast, fail clearly.** `from_cli` performs exhaustive validation (presence, type, range, enum) before constructing the object, and `validate_source_paths()` checks filesystem accessibility immediately after — so configuration errors surface with actionable messages before any expensive analysis or LLM calls begin. -- **One shape, many origins.** Regardless of whether a `Config` originates from CLI environment variables, a persisted `ConfigManager` configuration, or a web job submission, all paths converge on the same validated dataclass shape, so the rest of the pipeline ([Backend Core](backend-core.md)) never needs to know which caller produced it. diff --git a/docs/reference/architecture/frontend-core/frontend-core.md b/docs/reference/architecture/frontend-core/frontend-core.md deleted file mode 100644 index 3bb2ed7e..00000000 --- a/docs/reference/architecture/frontend-core/frontend-core.md +++ /dev/null @@ -1,257 +0,0 @@ -# Frontend Core - -## Introduction - -The Frontend Core module implements CodeWiki's **web application layer** — a FastAPI-based service that lets users submit GitHub repository URLs, tracks documentation-generation jobs asynchronously, caches completed results, and serves the generated documentation back to the browser. - -It acts as the bridge between end users (submitting repositories through a web form) and the heavier documentation-generation machinery implemented in [Backend Core](backend-core.md) (specifically `DocumentationGenerator`) and the shared runtime settings in [Config Core](config-core.md) (`Config`). - -Unlike `cli-core` and `backend-core`, this module is not further decomposed into child sub-modules in the module tree — it is a compact, single-layer module. This document therefore covers all of its components directly, without separate sub-module pages. - -## Responsibilities - -- Validate and normalize submitted GitHub repository URLs -- Queue documentation-generation jobs and process them on a background thread -- Cache generated documentation by repository URL to avoid redundant regeneration -- Persist job status and cache metadata to disk so state survives restarts -- Render the web UI (submission form, job list, generated docs viewer) via Jinja2 templates -- Expose HTTP endpoints (via FastAPI route handlers) for submission, status polling, and documentation viewing - -## Architecture Overview - -The module is organized around a simple pipeline: a web request creates or looks up a `JobStatus`, which is queued to the `BackgroundWorker`. The worker clones the repository, invokes the documentation generator, and stores results through the `CacheManager`. All directories, timeouts, and queue sizing are centralized in `WebAppConfig`. - -```mermaid -flowchart TD - User["Browser / API Client"] -->|"submit repo_url"| Routes["WebRoutes"] - Routes -->|"validate URL"| GitProc["GitHubRepoProcessor"] - Routes -->|"check cache"| Cache["CacheManager"] - Routes -->|"enqueue job"| Worker["BackgroundWorker"] - Routes -->|"render HTML"| Templates["StringTemplateLoader / render_template"] - - Worker -->|"clone repository"| GitProc - Worker -->|"build Config.from_web_job"| ConfigCore["Config (config-core)"] - Worker -->|"generate docs"| DocGen["DocumentationGenerator (backend-core)"] - Worker -->|"store result path"| Cache - Worker -->|"persist status"| JobsFile[("jobs.json")] - - Cache -->|"persist index"| CacheFile[("cache_index.json")] - - subgraph models_group["Data Models"] - JobStatus["JobStatus"] - CacheEntry["CacheEntry"] - RepositorySubmission["RepositorySubmission"] - JobStatusResponse["JobStatusResponse"] - end - - Routes --> models_group - Worker --> models_group - Cache --> models_group -``` - -**Cross-module dependencies:** -- [Backend Core](backend-core.md) — `DocumentationGenerator` performs the actual dependency analysis and LLM-driven documentation generation invoked by `BackgroundWorker`. -- [Config Core](config-core.md) — `Config.from_web_job()` builds the runtime configuration (models, API keys, directories) used for each documentation job. - -## Core Components - -### WebAppConfig — Central Settings - -`WebAppConfig` (in `config.py`) is a plain class holding static configuration constants used across the whole module: - -- **Directories**: `CACHE_DIR`, `TEMP_DIR`, `OUTPUT_DIR` -- **Queue settings**: `QUEUE_SIZE` -- **Cache settings**: `CACHE_EXPIRY_DAYS` -- **Job cleanup**: `JOB_CLEANUP_HOURS`, `RETRY_COOLDOWN_MINUTES` -- **Server defaults**: `DEFAULT_HOST`, `DEFAULT_PORT` -- **Git clone settings**: `CLONE_TIMEOUT`, `CLONE_DEPTH` - -It also provides `ensure_directories()` (creates cache/temp/output folders) and `get_absolute_path()`. Every other component in this module reads its defaults from `WebAppConfig` unless overridden by an explicit constructor argument. - -### GitHubRepoProcessor — Repository Validation & Cloning - -`GitHubRepoProcessor` is a stateless utility class (all static methods) responsible for: - -- `is_valid_github_url(url)` — ensures the URL points to `github.com`/`www.github.com` with a valid `owner/repo` path -- `get_repo_info(url)` — extracts `owner`, `repo`, `full_name`, and a normalized `clone_url` -- `clone_repository(clone_url, target_dir, commit_id=None)` — clones via `git clone` (shallow, depth-limited by `WebAppConfig.CLONE_DEPTH`, unless a specific `commit_id` is requested, in which case a full clone + `git checkout` is performed); cleans up the target directory on failure - -This component has no dependency on any other module — it only shells out to `git` and reads settings from `WebAppConfig`. - -### CacheManager — Documentation Cache - -`CacheManager` maintains an on-disk index (`cache_index.json`) mapping a SHA-256 hash of the repository URL (`get_repo_hash`) to a `CacheEntry` describing where the generated docs live and when they were created/last accessed. - -Key behaviors: -- `get_cached_docs(repo_url)` — returns the cached docs path if the entry exists and has not expired (`CACHE_EXPIRY_DAYS`); expired entries are automatically removed -- `add_to_cache(repo_url, docs_path)` — creates/updates a `CacheEntry` and persists the index -- `remove_from_cache(repo_url)` / `cleanup_expired_cache()` — cache invalidation utilities -- Corrupted index files are detected and backed up rather than crashing the app - -### BackgroundWorker — Asynchronous Job Processing - -`BackgroundWorker` is the core orchestration engine of the module. It owns: - -- A bounded `Queue` (`processing_queue`, sized by `WebAppConfig.QUEUE_SIZE`) of job IDs waiting to be processed -- An in-memory `job_status: Dict[str, JobStatus]` map -- A `jobs.json` file for persisting completed job state across restarts - -Lifecycle: -1. `start()` launches a daemon thread running `_worker_loop()`, which polls the queue and dispatches jobs to `_process_job()`. -2. `add_job(job_id, job)` registers a new `JobStatus` and enqueues its ID. -3. `_process_job(job_id)`: - - Checks `CacheManager` first — if valid cached docs exist, marks the job `completed` immediately. - - Otherwise resolves repo info via `GitHubRepoProcessor.get_repo_info()`, clones the repository into a per-job temp directory (`GitHubRepoProcessor.clone_repository`, optionally checking out a specific `commit_id`). - - Builds a `Config` via `Config.from_web_job(repo_path, docs_dir)` (see [Config Core](config-core.md)). - - Instantiates `DocumentationGenerator(config, job.commit_id)` from [Backend Core](backend-core.md) and runs its async `run()` method in a dedicated event loop. - - On success, registers the output path with `CacheManager.add_to_cache()` and marks the job `completed`; on failure, marks it `failed` with an `error_message`. - - Always cleans up the temporary cloned repository directory. - -`load_job_statuses()` / `save_job_statuses()` persist only `completed` jobs to `jobs.json`. If no jobs file exists yet, `_reconstruct_jobs_from_cache()` rebuilds job entries directly from the `CacheManager`'s index for backward compatibility (older deployments that only had cache data). - -```mermaid -sequenceDiagram - participant Browser - participant Routes as "WebRoutes" - participant Worker as "BackgroundWorker" - participant Cache as "CacheManager" - participant Git as "GitHubRepoProcessor" - participant DocGen as "DocumentationGenerator" - - Browser->>Routes: POST / (repo_url, commit_id) - Routes->>Git: is_valid_github_url / get_repo_info - Routes->>Cache: get_cached_docs(repo_url) - alt Cache hit - Cache-->>Routes: docs_path - Routes-->>Browser: Render success message - else Cache miss - Routes->>Worker: add_job(job_id, JobStatus) - Worker-->>Routes: queued - Routes-->>Browser: Render "queued" message - Worker->>Git: clone_repository(clone_url, temp_dir, commit_id) - Worker->>DocGen: run() (async documentation generation) - DocGen-->>Worker: docs_dir populated - Worker->>Cache: add_to_cache(repo_url, docs_path) - Worker->>Worker: save_job_statuses() - end -``` - -### Data Models - -Defined in `models.py`, these dataclasses and Pydantic models flow between the components above: - -| Model | Kind | Purpose | -|---|---|---| -| `RepositorySubmission` | Pydantic `BaseModel` | Validates form input containing a `repo_url: HttpUrl` | -| `JobStatusResponse` | Pydantic `BaseModel` | Shape of the `/api/jobs/{job_id}` JSON response (status, timestamps, error, docs path, model used, commit) | -| `JobStatus` | `dataclass` | In-memory/on-disk representation of a job's lifecycle: `queued` → `processing` → `completed`/`failed` | -| `CacheEntry` | `dataclass` | Cache index entry: repo URL, its hash, docs path, creation and last-access timestamps | - -`JobStatus` and `CacheEntry` are pure data holders serialized manually (via `dataclasses.asdict`/manual dict construction) by `BackgroundWorker` and `CacheManager` respectively — they carry no business logic themselves. - -### WebRoutes — HTTP Route Handlers - -`WebRoutes` implements the FastAPI-facing handlers, wired to a `BackgroundWorker` and `CacheManager` instance: - -- `index_get(request)` — renders the main submission form plus the 100 most recent jobs -- `index_post(request, repo_url, commit_id)` — the primary submission flow: - 1. Cleans up expired jobs (`cleanup_old_jobs`) - 2. Validates the URL via `GitHubRepoProcessor` - 3. Normalizes the URL and derives a URL-safe `job_id` (`owner--repo`) - 4. Checks for an existing in-flight or recently-failed job (respecting `WebAppConfig.RETRY_COOLDOWN_MINUTES`) to prevent duplicate work - 5. Checks the cache; if a hit, synthesizes a `completed` `JobStatus` for immediate display - 6. Otherwise creates a new `queued` `JobStatus` and calls `BackgroundWorker.add_job()` -- `get_job_status(job_id)` — JSON API returning a `JobStatusResponse` -- `view_docs(job_id)` — redirects to the static documentation viewer for a completed job -- `serve_generated_docs(job_id, filename)` — resolves and renders a specific generated Markdown file (with path-traversal protection), falling back to reconstructing job state from the cache if no in-memory job exists; loads `module_tree.json`/`metadata.json` for navigation and converts Markdown to HTML for display -- Helper methods `_normalize_github_url`, `_repo_full_name_to_job_id`, `_job_id_to_repo_full_name`, `cleanup_old_jobs` support the above flows - -### Template Rendering - -`template_utils.py` provides a thin Jinja2 integration layer: - -- `StringTemplateLoader` — a custom `jinja2.BaseLoader` that serves a template directly from a Python string (no filesystem template directory needed), enabling templates to be defined inline as Python string constants elsewhere in the application -- `render_template(template, context)` — configures a `Jinja2` `Environment` (autoescaping HTML/XML, `trim_blocks`/`lstrip_blocks` enabled) and renders the given template string against a context dict -- `render_navigation(module_tree, current_page)` — renders sidebar navigation HTML from a documentation module tree structure -- `render_job_list(jobs)` — renders the recent-jobs list HTML fragment - -`WebRoutes` uses `render_template` to produce every `HTMLResponse` it returns. - -## Component Relationships - -```mermaid -classDiagram - class WebAppConfig { - +CACHE_DIR - +TEMP_DIR - +QUEUE_SIZE - +CACHE_EXPIRY_DAYS - +CLONE_TIMEOUT - +ensure_directories() - } - class GitHubRepoProcessor { - +is_valid_github_url(url) - +get_repo_info(url) - +clone_repository(clone_url, target_dir, commit_id) - } - class CacheManager { - +cache_index - +get_cached_docs(repo_url) - +add_to_cache(repo_url, docs_path) - +remove_from_cache(repo_url) - } - class BackgroundWorker { - +job_status - +processing_queue - +start() - +add_job(job_id, job) - +get_job_status(job_id) - } - class WebRoutes { - +index_get(request) - +index_post(request, repo_url, commit_id) - +get_job_status(job_id) - +serve_generated_docs(job_id, filename) - } - class JobStatus - class CacheEntry - class RepositorySubmission - class JobStatusResponse - class StringTemplateLoader - - WebRoutes --> BackgroundWorker - WebRoutes --> CacheManager - WebRoutes --> GitHubRepoProcessor - WebRoutes --> StringTemplateLoader - BackgroundWorker --> CacheManager - BackgroundWorker --> GitHubRepoProcessor - BackgroundWorker --> JobStatus - CacheManager --> CacheEntry - WebRoutes --> JobStatusResponse - GitHubRepoProcessor --> WebAppConfig - CacheManager --> WebAppConfig - BackgroundWorker --> WebAppConfig -``` - -## Job Lifecycle State Machine - -```mermaid -stateDiagram-v2 - [*] --> queued: "add_job()" - queued --> processing: "worker picks up job" - processing --> completed: "cache hit OR generation succeeds" - processing --> failed: "clone or generation error" - failed --> queued: "resubmission after cooldown" - completed --> [*] - failed --> [*] -``` - -## Integration with the Rest of CodeWiki - -- **Documentation generation**: `BackgroundWorker._process_job()` delegates the actual analysis and Markdown/diagram generation to `DocumentationGenerator` from [Backend Core](backend-core.md). Frontend Core does not implement any dependency analysis itself — it only manages the job lifecycle, caching, and presentation around that generator. -- **Configuration**: Every job builds a fresh runtime `Config` via `Config.from_web_job(repo_path, docs_dir)`, defined in [Config Core](config-core.md). This keeps LLM model selection, API keys, and output directories consistent with the rest of the pipeline while allowing the web app to supply job-specific paths. -- **Independent of CLI**: Unlike [CLI Core](cli-core.md), which drives documentation generation from the command line with its own `ConfigManager` and job models, Frontend Core is a self-contained HTTP-facing alternative entry point that shares the same downstream `DocumentationGenerator` and `Config` but has its own job-tracking (`JobStatus`) and caching (`CacheManager`) implementations tailored for a multi-user web environment (queueing, retry cooldowns, cache expiry). - -## Summary - -Frontend Core provides the web-facing shell around CodeWiki's documentation engine: validating and queueing repository submissions, running generation jobs on a background thread, caching results to avoid repeat work, and rendering both the submission UI and the generated documentation itself. Its five main building blocks — `WebAppConfig`, `GitHubRepoProcessor`, `CacheManager`, `BackgroundWorker`, and `WebRoutes` — form a straightforward pipeline, with `BackgroundWorker` acting as the connective tissue to the heavier [Backend Core](backend-core.md) documentation generator and [Config Core](config-core.md) configuration model. diff --git a/docs/reference/architecture/test-clustering/test-clustering.md b/docs/reference/architecture/test-clustering/test-clustering.md deleted file mode 100644 index 849c62b6..00000000 --- a/docs/reference/architecture/test-clustering/test-clustering.md +++ /dev/null @@ -1,208 +0,0 @@ -# Test Clustering - -## Purpose - -The Test Clustering module is a collection of standalone diagnostic and validation scripts used to exercise the LLM-driven **module clustering** functionality that lives inside the backend's documentation-generation pipeline (`cluster_modules`). Unlike a conventional `pytest` suite, these scripts are executable Python programs (`python3 script.py`) that print human-readable pass/fail reports to the console and exit with a non-zero status code on failure, making them suitable for quick manual runs, CI smoke checks, or debugging sessions when the clustering behavior of the underlying LLM changes. - -Clustering is the step in the CodeWiki pipeline where a flat list of code components (functions, classes, files) discovered by the [dependency analyzer](backend-core.md) is grouped by an LLM into a hierarchical module tree (e.g. "Auth Module", "API Module") that later becomes the basis for the generated documentation structure. Because this step depends on free-form LLM output, it is especially prone to format drift (e.g., the LLM returning quoted strings or class names instead of integer IDs). The scripts in this module were written to reproduce, isolate, and regression-test these failure modes. - -## Scope and Relationship to Other Modules - -This module does not define new production functionality; instead, it directly imports and drives components from other parts of the system: - -- **`cluster_modules`, `create_component_id_map`, `normalize_component_ids_by_lookup`** — the LLM clustering functions under test, part of the backend's documentation-generation pipeline (see [Backend Core](backend-core.md)). -- **`Node`** — the dependency-graph node model representing a single code component, documented as part of [Backend Core](backend-core/dependency-analyzer-models/dependency-analyzer-models.md). -- **`Config`** — the pipeline configuration object (model names, API keys, base URLs, token thresholds), documented in [Config Core](config-core.md). - -Because the clustering step is LLM-backed, most scripts in this module require valid API credentials (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, or the `MAIN_API_KEY` / `CLUSTER_API_KEY` / `FALLBACK_API_KEY` overrides) to be set in the environment or a local `.env.local` file before they can make live calls. The two validation-focused scripts (`test_clustering_validation.py` and `test_id_based_clustering.py`) are pure logic simulations and do **not** require any network access or API keys. - -## Architecture Overview - -```mermaid -flowchart TD - subgraph scripts["Test Clustering Scripts"] - Debug["test_clustering_debug.py"] - Forced["test_clustering_forced.py"] - Local["test_clustering_local.py"] - Integration["test_clustering_integration.py"] - Validation["test_clustering_validation.py"] - IdBased["test_id_based_clustering.py"] - end - - subgraph backend["Backend Clustering Pipeline"] - ClusterFn["cluster_modules()"] - IdMap["create_component_id_map()"] - Normalize["normalize_component_ids_by_lookup()"] - LLMClient["LLM Client"] - end - - ConfigMod["Config"] - NodeMod["Node"] - - Debug -->|"invokes"| ClusterFn - Forced -->|"invokes"| ClusterFn - Local -->|"invokes"| ClusterFn - ClusterFn -->|"calls"| LLMClient - - Integration -->|"invokes directly"| IdMap - Integration -->|"invokes directly"| Normalize - - Validation -->|"simulates validation logic of"| ClusterFn - IdBased -->|"simulates parsing/validation logic of"| ClusterFn - - Debug -->|"constructs"| NodeMod - Forced -->|"constructs"| NodeMod - Local -->|"constructs"| NodeMod - Debug -->|"constructs"| ConfigMod - Forced -->|"constructs"| ConfigMod - Local -->|"constructs"| ConfigMod -``` - -## The Common `TestResults` Pattern - -Nearly every script in this module defines its own local `TestResults` class rather than importing a shared one. This is a deliberate consequence of these being independent, copy-paste-friendly diagnostic scripts rather than a shared test library. Each implementation follows the same basic contract: - -1. Accumulate `(name, passed, details)` tuples via an `add_test(...)` method. -2. Print a formatted summary via `print_summary()`. -3. Return or expose an overall boolean/exit-code indicating whether all tests passed. - -| Script | `TestResults` Behavior | Exit Code Semantics | -|---|---|---| -| `test_clustering_debug.py` | Simple `add_test(name, passed, details)`; summary prints ✅/❌ per test | `sys.exit(0 if success else 1)` | -| `test_clustering_forced.py` | Same shape, message-only details | `sys.exit(0 if passed else 1)` | -| `test_clustering_local.py` | Same shape; returns `passed_count == total` | `sys.exit(0 if success else 1)` | -| `test_clustering_integration.py` | Tracks explicit `passed` / `failed` counters plus a list of dict-based test records | `main()` returns `0` or `1`, used as process exit code | -| `test_clustering_validation.py` | Tracks `passed` / `failed` counters and a `failures` list; exposes a `success` property | `exit(0 if success else 1)` | -| `test_id_based_clustering.py` | Minimal `(name, passed)` tuple list; `print_summary()` returns `all_passed` | `sys.exit(0)` / `sys.exit(1)` | - -```mermaid -classDiagram - class TestResultsBase { - +add_test(name, passed, details) - +print_summary() bool - } - class DebugResults - class ForcedResults - class LocalResults - class IntegrationResults { - +passed int - +failed int - } - class ValidationResults { - +success bool - } - class IdBasedResults - - TestResultsBase <|-- DebugResults - TestResultsBase <|-- ForcedResults - TestResultsBase <|-- LocalResults - TestResultsBase <|-- IntegrationResults - TestResultsBase <|-- ValidationResults - TestResultsBase <|-- IdBasedResults -``` - -Note: `TestResultsBase` above is a conceptual grouping for documentation purposes only — each script defines its own independent class with no shared base class or import relationship in the actual source code. - -## Script Reference - -### `test_clustering_debug.py` - -Runs a full, live invocation of `cluster_modules` against four hand-crafted `Node` components (`AuthController`, `AuthService`, `UserController`, `UserService`) and monkey-patches the LLM client factory (`create_llm_client`) so that the raw LLM response text can be captured and printed. This is the go-to script when the clustering output looks wrong and you need to see exactly what the LLM returned (including whether it emitted the expected `` tag). - -Key characteristics: -- Loads credentials from `.env.local` via `python-dotenv`. -- Builds a `Config` object with `repo_path` overridable via the `CODEWIKI_TEST_REPO` environment variable (defaults to a local `fixtures/sample_repo` directory). -- Prints up to 2000 characters of the captured LLM response for inspection. -- Reports failure with diagnostic detail when the module tree comes back empty (including whether the `` tag was present in the response). - -### `test_clustering_forced.py` - -Similar in structure to the debug script, but deliberately sets `max_token_per_module=50` on the `Config` to force the clustering logic to invoke the LLM (clustering is normally skipped when the component set is small enough to fit under the token threshold). It also lowers `cluster_max_tokens` to keep the forced call cheap. Ten synthetic `Component{i}` nodes are generated to guarantee the threshold is exceeded. - -This script is useful for confirming that: -- The token-based trigger for invoking the LLM clustering path actually fires. -- The LLM's response still respects the expected `` tag format under a forced, low-budget scenario. - -### `test_clustering_local.py` - -A more guarded, "safe to run locally" variant that requires an existing on-disk repository (`CODEWIKI_TEST_REPO`) and specific Java source files inside it (`AuthController.java`, `AuthService.java`, `UserController.java`, `UserService.java`, under an OpenFrame-style API project layout). If the repo or API keys are missing, the script exits early with a clear message rather than attempting an LLM call. It wraps the `cluster_modules` invocation in a `try`/`except` block and prints a full traceback on unexpected exceptions, in addition to the standard `TestResults` summary. - -### `test_clustering_integration.py` - -The most comprehensive script in this module. Rather than calling `cluster_modules` end-to-end, it exercises two lower-level helper functions directly: -- `create_component_id_map(components)` — builds the `id -> FQDN` lookup and human-readable ID descriptions used to keep LLM prompts compact. -- `normalize_component_ids_by_lookup(module_tree, id_to_fqdn)` — converts LLM-returned component IDs back into fully-qualified component names, filtering out anything invalid. - -It defines a lightweight `MockNode` class (with `fqdn`, `name`, `file_path` attributes) to avoid depending on the real `Node` model, and a `capture_log_warnings` decorator that redirects the `codewiki.src.be.cluster_modules` logger into an in-memory buffer so tests can assert on specific warning strings (e.g., `"Non-integer ID"`, `"Invalid ID 999"`, `"Valid range"`). - -Ten focused test functions cover the following normalization edge cases: - -| Test | Input | Expected Outcome | -|---|---|---| -| `test_component_id_map_creation` | 5 sample components | Sequential integer IDs `0..4` map 1:1 to FQDNs | -| `test_valid_integer_ids` | `[0, 1, 2]` | All normalize cleanly, no warnings | -| `test_invalid_quoted_integers` | `["0", "1", "2"]` | Accepted via `int()` conversion (resilient behavior), no warnings | -| `test_invalid_class_names` | `["AuthService", "CountedGenericQueryResult"]` | All rejected with `"Non-integer ID"` warnings | -| `test_mixed_invalid_ids` | `[0, "1", "AuthService", 999]` | Only `0` and `"1"` (→ 1) normalize; the rest are rejected | -| `test_out_of_range_ids` | `[0, 1, 999]` | `999` rejected with `"Invalid ID 999"` / `"Valid range"` warning | -| `test_json_loads_normalization` | `json.loads("[0, 1, 2]")` | Normalizes cleanly, confirming safe JSON parsing works | -| `test_empty_list` | `[]` | No components, no warnings | -| `test_duplicate_ids` | `[0, 1, 1, 2]` | Duplicates preserved (4 components, FQDN for `1` appears twice) | -| `test_negative_ids` | `[-1, 0, 1]` | `-1` rejected with `"Invalid ID -1"` warning | - -### `test_clustering_validation.py` - -A pure-simulation script that re-implements the validation logic from `cluster_modules.py` (referenced as lines 338–369 in the source comments) inside a local `simulate_validation(response_content, max_id)` function. It parses a raw JSON string with `json.loads` (explicitly avoiding `eval()` for safety) and checks that every component ID in every module is an integer within `[0, max_id]`. It is run against a table of ten hard-coded test cases covering valid bare integers, quoted integers, string class names, mixed types, out-of-range and negative IDs, malformed/trailing-comma JSON, empty component lists, and modules with no `components` key at all — asserting that the pass/fail outcome matches the expected `should_pass` flag for each case. - -### `test_id_based_clustering.py` - -The most isolated script — it has **no dependency on the `codewiki` package at all** and instead re-implements small snippets of the clustering logic inline to document and verify three specific historical bug fixes: - -1. **`json.loads()` instead of `eval()`** for parsing LLM responses (`test_json_parsing`) — demonstrates that quoted string IDs parse but would subsequently fail type validation. -2. **Correct 4-tuple unpacking** of a mocked `format_potential_core_components()` function (`test_return_types`) — guards against a regression where the last tuple element (a `Dict` of ID descriptions) was mistakenly treated as the string passed to a token counter. -3. **Integer ID range validation** (`test_id_validation`) and **ID-to-FQDN normalization** (`test_normalization`) — smaller, self-contained versions of the same checks performed by `test_clustering_integration.py`, useful as a minimal reproduction when debugging without the full backend installed. - -## Running the Scripts - -All scripts are plain Python entry points and can be run directly, for example: - -```bash -python3 test_clustering_debug.py -python3 test_clustering_forced.py -python3 test_clustering_local.py -python3 test_clustering_integration.py -python3 test_clustering_validation.py -python3 test_id_based_clustering.py -``` - -Scripts that perform live LLM calls (`test_clustering_debug.py`, `test_clustering_forced.py`, `test_clustering_local.py`) read credentials and model overrides from environment variables such as `MAIN_MODEL`, `CLUSTER_MODEL`, `FALLBACK_MODEL`, `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `MAIN_API_KEY`, `CLUSTER_API_KEY`, and `FALLBACK_API_KEY`, and optionally load a `.env.local` file via `python-dotenv`. The purely logical scripts (`test_clustering_validation.py`, `test_id_based_clustering.py`) run without any environment configuration. - -## Typical Debugging Flow - -```mermaid -sequenceDiagram - participant Dev as Developer - participant Debug as "test_clustering_debug.py" - participant Patch as "Monkey-patched LLM Client" - participant Cluster as "cluster_modules()" - - Dev->>Debug: python3 test_clustering_debug.py - Debug->>Patch: patch create_llm_client() - Debug->>Cluster: cluster_modules(leaf_nodes, components, config) - Cluster->>Patch: client.call(prompt) - Patch-->>Cluster: raw LLM response text - Patch-->>Debug: capture response (global variable) - Cluster-->>Debug: module_tree dict - Debug->>Dev: print captured response + module_tree summary - Debug->>Dev: TestResults.print_summary() -``` - -## Summary - -The Test Clustering module provides a layered set of confidence checks for the LLM-based clustering step of the CodeWiki pipeline: - -- **End-to-end scripts** (`test_clustering_debug.py`, `test_clustering_forced.py`, `test_clustering_local.py`) exercise the real `cluster_modules` function against live LLM calls, differing mainly in how they trigger the LLM path and how much diagnostic output they surface. -- **Targeted unit-style scripts** (`test_clustering_integration.py`) validate the ID-mapping and normalization helper functions in isolation, covering a wide range of malformed-input edge cases without requiring any LLM call. -- **Pure simulation scripts** (`test_clustering_validation.py`, `test_id_based_clustering.py`) re-implement small pieces of the validation logic to document and regression-test specific historical bugs (unsafe `eval()` usage, tuple-unpacking mistakes, ID range checks) without any dependency on the live backend or network access. - -For details on the underlying components these scripts depend on, see [Backend Core](backend-core.md) (clustering pipeline and dependency graph models) and [Config Core](config-core.md) (pipeline configuration). diff --git a/docs/reference/architecture/test-multi-path/sample_fixtures.md b/docs/reference/architecture/test-multi-path/sample_fixtures.md deleted file mode 100644 index 5f477219..00000000 --- a/docs/reference/architecture/test-multi-path/sample_fixtures.md +++ /dev/null @@ -1,125 +0,0 @@ -# Sample Fixtures - -The Sample Fixtures module provides a small, self-contained set of illustrative components used to exercise the multi-path analysis capabilities of the dependency analyzer. It models a minimal API-driven application composed of a request controller, a core service, and a pluggable data-processing interface. These fixtures are not part of the production CodeWiki system itself — they exist as test data to validate that the [Backend Core](backend-core.md) dependency analysis and call-graph tooling correctly resolve cross-file and cross-package relationships (controller → service → plugin) across multiple source roots. - -## Purpose and Scope - -Sample Fixtures serves as a synthetic fixture package with three cooperating pieces: - -- **`APIController`** — the entry point that receives API-style requests and routes them to the service layer. -- **`MainService`** — the core business logic component that validates and processes request data, and depends on helper utilities from an external `deps` package. -- **`PluginInterface` / `DataPlugin`** — an extensible plugin abstraction, defined in a separate `external` package, representing a third-party/pluggable extension point in the dependency graph. - -Because the components live in distinct sub-packages (`main.controller`, `main.service`, `external.plugin`), this module is specifically useful for verifying that multi-path/multi-root dependency resolution works correctly — hence its parent module name, [Test Multi Path](test-multi-path.md). - -## Module Position - -Sample Fixtures is a child module of [Test Multi Path](test-multi-path.md), alongside its sibling module [Test Suites](test_suites.md), which contains the test runners (`IntegrationTestRunner`, `TestResults`, `Colors`) that exercise these fixtures. - -```mermaid -graph TD - Parent["Test Multi Path"] --> SampleFixtures["Sample Fixtures"] - Parent --> TestSuites["Test Suites"] - SampleFixtures -.->|"exercised by"| TestSuites -``` - -## Component Architecture - -The three components form a simple layered call chain: a controller delegates to a service, and the service (conceptually) could delegate further to a plugin-based extension point. `DataPlugin` implements the `PluginInterface` contract, demonstrating polymorphic extension. - -```mermaid -classDiagram - class APIController { - +service: MainService - +request_count: int - +__init__() - +handle_request(endpoint, data) dict - } - class MainService { - +name: str - +active: bool - +__init__(name) - +start(config) bool - +process_request(data) dict - +stop() - } - class PluginInterface { - +initialize() bool - +execute(context) dict - } - class DataPlugin { - +name: str - +initialized: bool - +__init__(name) - +initialize() bool - +execute(context) dict - } - APIController --> MainService : "uses" - DataPlugin --|> PluginInterface : "implements" -``` - -### APIController - -`APIController` (defined in `test-multi-path/main/controller.py`) is the request-routing entry point of the fixture application. On construction it instantiates a `MainService` named `"api-service"` and tracks a running `request_count`. - -Its single public method, `handle_request(endpoint, data)`, performs simple string-based routing: - -- `"/process"` — delegates to `MainService.process_request(data)`. -- `"/health"` — returns a status payload including the current `request_count`. -- Any other endpoint — returns an `{"error": "Unknown endpoint"}` response. - -```mermaid -sequenceDiagram - participant Client - participant Controller as APIController - participant Service as MainService - - Client->>Controller: handle_request("/process", data) - Controller->>Controller: request_count += 1 - Controller->>Service: process_request(data) - Service->>Service: check active flag - Service-->>Controller: {"status": "success", "result": processed} - Controller-->>Client: response dict -``` - -### MainService - -`MainService` (defined in `test-multi-path/main/service.py`) encapsulates the core business logic of the fixture application. It is constructed with a `name` and starts in an inactive state (`active = False`). - -Key behaviors: - -- **`start(config)`** — validates the supplied configuration using `validate_input` (imported from an external `deps.helper` module, outside this module's scope) and, if valid, flips `active` to `True`. -- **`process_request(data)`** — raises a `RuntimeError` if the service has not been started; otherwise calls `process_data` (also from `deps.helper`) and wraps the result in a structured response dict containing `status`, `service`, and `result`. -- **`stop()`** — resets `active` to `False`. - -This component demonstrates a cross-package dependency: `MainService` lives in `main.service` but relies on helper functions declared in a separate `deps` package, which is a key scenario the dependency analyzer's multi-path resolution is designed to detect correctly. - -### PluginInterface and DataPlugin - -Defined in `test-multi-path/external/plugin.py`, these two classes model a pluggable extension mechanism located in a distinct `external` package: - -- **`PluginInterface`** is an abstract base defining two contract methods, `initialize()` and `execute(context)`, both of which raise `NotImplementedError` in the base class. -- **`DataPlugin`** is a concrete implementation. It tracks an `initialized` flag, requires `initialize()` to be called before `execute()` can run (otherwise raising `RuntimeError`), and its `execute(context)` method extracts a `"data"` key from the passed `context` dict and returns a formatted result string wrapped in a response dict. - -```mermaid -stateDiagram-v2 - [*] --> Uninitialized: "DataPlugin(name)" - Uninitialized --> Initialized: "initialize()" - Initialized --> Initialized: "execute(context)" - Uninitialized --> Error: "execute() before initialize()" -``` - -## Cross-Module Dependency Flow - -The fixture's package layout — `main.controller`, `main.service`, and `external.plugin` — intentionally spans multiple directories/roots so that the dependency analysis pipeline described in [Backend Core](backend-core.md) can be validated against realistic multi-path import resolution scenarios (e.g., resolving `from service import MainService` and `from deps.helper import process_data, validate_input` across different source roots). - -```mermaid -graph LR - Controller["main.controller.APIController"] -->|"imports"| Service["main.service.MainService"] - Service -->|"imports"| Deps["deps.helper (external)"] - Plugin["external.plugin.DataPlugin"] -->|"implements"| Iface["external.plugin.PluginInterface"] -``` - -## Usage in Testing - -Sample Fixtures components are consumed by the sibling [Test Suites](test_suites.md) module, whose `IntegrationTestRunner` exercises the controller-to-service call chain and validates that analysis output (call graphs, dependency edges) matches expectations for this intentionally cross-package layout. diff --git a/docs/reference/architecture/test-multi-path/test-multi-path.md b/docs/reference/architecture/test-multi-path/test-multi-path.md deleted file mode 100644 index 72143a16..00000000 --- a/docs/reference/architecture/test-multi-path/test-multi-path.md +++ /dev/null @@ -1,87 +0,0 @@ -# Test Multi Path - -## Purpose - -The Test Multi Path module is a self-contained **test fixture and validation suite** used to verify the multi-path source analysis capability of CodeWiki's dependency analysis pipeline. It does not implement production application logic; instead, it provides: - -1. **Sample application code** (a small "service + controller + plugin" codebase spread across multiple simulated source roots: `main/`, `deps/`, `external/`, and `vendor/`) that mimics a real repository with cross-directory imports. -2. **Test runner scripts** that configure the [Config](config-core.md) object with `additional_source_paths`, invoke `DependencyGraphBuilder` from the [Backend Core](backend-core.md) module, and assert that components are correctly discovered, namespaced, and (where applicable) linked across path boundaries. - -This module exists to answer a specific engineering question for CodeWiki itself: *"When a repository's source code is split across multiple root directories (e.g., a monorepo with `main`, `deps`, and `vendor` folders), can the dependency analyzer correctly parse, namespace, and graph components from all of them without ID collisions or missed dependencies?"* - -## Architecture Overview - -The module has two cooperating halves: **fixtures** (sample code to be analyzed) and **suites** (scripts that drive the analysis and assert on results). - -```mermaid -flowchart TD - subgraph fixtures["Sample Application Fixtures"] - Service["MainService"] - Controller["APIController"] - Plugin["DataPlugin / PluginInterface"] - end - - subgraph suites["Test Suites and Runners"] - MultiPathSuite["test_multi_path.py suite"] - IntegrationRunner["IntegrationTestRunner"] - end - - subgraph external_deps["External Dependencies"] - ConfigCls["Config"] - Builder["DependencyGraphBuilder"] - end - - MultiPathSuite -->|"constructs"| ConfigCls - IntegrationRunner -->|"constructs"| ConfigCls - ConfigCls -->|"declares additional_source_paths"| Builder - Builder -->|"parses"| Service - Builder -->|"parses"| Controller - Builder -->|"parses"| Plugin - MultiPathSuite -->|"asserts on"| Builder - IntegrationRunner -->|"asserts on"| Builder -``` - -- **Config** and **DependencyGraphBuilder** are core components of the [Config Core](config-core.md) and [Backend Core](backend-core.md) modules respectively; Test Multi Path exercises them but does not own them. -- The fixtures (`main/service.py`, `main/controller.py`, `external/plugin.py`) are ordinary Python source files that stand in for a target repository being documented. -- The suites (`test_multi_path.py`, `integration_test.py`) are executable scripts (not pytest-based) that print colorized pass/fail output and return process exit codes. - -## Sub-modules - -### [Sample Fixtures](sample_fixtures.md) - -Contains the sample application code used as analysis input: `MainService`, `APIController`, `DataPlugin`, and `PluginInterface`. These classes simulate a layered application (controller → service → helper) spread across the `main/` and `external/` source roots so that the analyzer has realistic cross-file and cross-path relationships to discover. - -### [Test Suites](test_suites.md) - -Contains the executable validation logic: the `test_multi_path.py` suite (`Colors`, `TestResults`) which runs eight discrete scenario checks (single path, multiple paths, namespacing, cross-path dependencies, invalid paths, empty paths, relative vs. absolute paths), and `integration_test.py` (`IntegrationTestRunner`, `TestResults`) which runs a single comprehensive end-to-end scenario against a freshly generated temporary repository with `main/`, `deps/`, and `vendor/` roots. - -## How the Suites Use the Analysis Pipeline - -Both test scripts follow the same general pattern when invoking the dependency analysis pipeline documented in [Backend Core](backend-core.md): - -```mermaid -sequenceDiagram - participant Suite as "Test Script" - participant Cfg as "Config" - participant Builder as "DependencyGraphBuilder" - - Suite->>Cfg: Create Config with repo_path and additional_source_paths - Suite->>Cfg: validate_source_paths() (optional, for invalid-path test) - Suite->>Builder: DependencyGraphBuilder(config) - Suite->>Builder: build_dependency_graph() - Builder-->>Suite: Returns (components, leaf_nodes) - Suite->>Suite: Assert namespaces, counts, and cross-path edges - Suite->>Suite: Print summary and exit code -``` - -Key points validated by these scripts: - -- **Namespacing:** component IDs are prefixed by their top-level source root (for example `main.service.MainService`) so that same-named classes in different roots (e.g., `main/service.py` vs. `deps/helper.py`) never collide. -- **Multi-path detection:** `Config.is_multi_path_mode()` correctly reports whether more than one source root was configured. -- **Path validation:** `Config.validate_source_paths()` raises an error when an additional path does not exist on disk. -- **Cross-path dependency behavior:** at the time these tests were written, import statements crossing source-root boundaries are parsed but not yet resolved into graph edges — the suites document this as expected/known behavior rather than a failure. - -## Relationship to Other Modules - -- **[Config Core](config-core.md):** Test Multi Path constructs `Config` instances (including the `additional_source_paths` and `validate_source_paths` behavior) to drive the scenarios under test. -- **[Backend Core](backend-core.md):** Test Multi Path exercises `DependencyGraphBuilder`, which in turn relies on the dependency-analyzer pipeline (AST parsing, node/graph models) documented in that module. diff --git a/docs/reference/architecture/test-multi-path/test_suites.md b/docs/reference/architecture/test-multi-path/test_suites.md deleted file mode 100644 index a2bcae90..00000000 --- a/docs/reference/architecture/test-multi-path/test_suites.md +++ /dev/null @@ -1,261 +0,0 @@ -# Test Suites - -## Introduction - -The Test Suites module contains the executable test scripts that validate CodeWiki's **multi-path dependency analysis** feature — the ability to analyze source code spread across several independent root directories (e.g. `main/`, `deps/`, `vendor/`) as a single logical repository while keeping component identifiers correctly namespaced and dependency edges correctly resolved. - -This module provides two complementary test scripts: - -- **`test_multi_path.py`** — a scenario-based smoke-test suite that exercises the `DependencyGraphBuilder` (see [Backend Core](backend-core.md)) with a series of independent, self-contained checks (single path, multiple paths, namespacing, cross-path dependencies, invalid paths, empty paths, relative vs. absolute paths). -- **`integration_test.py`** — a full end-to-end integration test that programmatically builds a temporary multi-directory sample repository, runs the complete analysis pipeline against it, and asserts on namespace counts, cross-namespace dependency detection, and overall component totals. - -Both scripts are executable as standalone CLI programs (`python test_multi_path.py`, `python integration_test.py`) and return a process exit code suitable for use in CI pipelines. - -This module is a child of the [Test Multi Path](test-multi-path.md) module and is a sibling of the [Sample Fixtures](sample_fixtures.md) module, which supplies some of the on-disk fixture files (`main/controller.py`, `main/service.py`, `external/plugin.py`) referenced by `test_multi_path.py`. - -## Module Purpose in the System - -Test Suites does not implement product functionality — it is a verification harness for the dependency analysis subsystem documented in [Backend Core](backend-core.md), specifically its multi-path graph construction logic. It exercises: - -- **`Config`** (from the [Config Core](config-core.md) module) — used to declare a primary `repo_path` plus a list of `additional_source_paths`. -- **`DependencyGraphBuilder`** (part of the dependency analyzer pipeline in [Backend Core](backend-core.md)) — the component under test, responsible for parsing all configured paths and producing a namespaced component graph. - -Because these scripts drive real analysis runs, they act as a living specification for how multi-path namespacing and cross-path dependency resolution are expected to behave, including documenting known limitations (e.g. cross-namespace dependency resolution is not yet implemented at the AST-parsing level). - -## Core Components - -| Component | File | Responsibility | -|---|---|---| -| `Colors` | `test_multi_path.py` | ANSI color code constants used for terminal output formatting | -| `TestResults` (scenario suite) | `test_multi_path.py` | Accumulates pass/fail results for the seven scenario tests and prints a summary | -| `IntegrationTestRunner` | `integration_test.py` | Orchestrates the full end-to-end integration test: environment setup, config creation, pipeline execution, and multi-stage validation | -| `TestResults` (integration suite) | `integration_test.py` | Accumulates detailed assertion results (with optional `details` text) for the integration run and prints a validation summary | - -> **Note:** Both scripts define a class named `TestResults`, but they are distinct, independent implementations local to each file — there is no shared base class or import relationship between them. - -### Colors - -A simple constants class holding ANSI escape codes (`GREEN`, `RED`, `YELLOW`, `BLUE`, `BOLD`, `END`) used by the module-level print helpers (`print_header`, `print_success`, `print_error`, `print_warning`, `print_info`) in `test_multi_path.py` to produce readable, color-coded console output. - -### TestResults (test_multi_path.py) - -```text -class TestResults: - def __init__(self): - self.results = {} # test_name -> bool - - def add_test(self, name, passed) -> None - def print_summary(self) -> int # returns 0 (all passed) or 1 (failures) -``` - -Used by `run_all_tests()` as a simple dictionary-backed accumulator. Each of the seven scenario test functions returns a boolean, which is recorded under a descriptive test name. `print_summary()` renders a pass/fail report and computes the suite's overall exit code. - -### IntegrationTestRunner - -`IntegrationTestRunner` is a stateful orchestrator class that owns the entire lifecycle of the integration test: - -```text -class IntegrationTestRunner: - def __init__(self): - self.test_dir # temp directory root - self.main_path # main/ source path - self.deps_path # deps/ source path - self.vendor_path # vendor/ source path - self.config # Config instance - self.builder # DependencyGraphBuilder instance - self.components # Dict[str, Node] result - self.leaf_nodes # leaf node result - self.results # TestResults instance -``` - -Its `run()` method executes the following ordered steps, each implemented as a dedicated method: - -1. `setup_test_environment()` — creates a temporary directory with three sub-paths (`main/`, `deps/`, `vendor/`) and populates each with hand-written Python fixture files (`service.py`, `api.py`, `models.py`, `controller.py`, `utils.py` in `main/`; `helper.py`, `validator.py`, `cache.py` in `deps/`; `logger.py`, `metrics.py` in `vendor/`) that intentionally cross-import between paths. -2. `create_config()` — builds a `Config` (see [Config Core](config-core.md)) with `repo_path` set to `main/` and `additional_source_paths` set to `[deps/, vendor/]`. -3. `validate_paths()` — confirms the root and all additional paths exist on disk. -4. `execute_dependency_parser()` — instantiates `DependencyGraphBuilder` (see [Backend Core](backend-core.md)), verifies `config.is_multi_path_mode()` returns `True`, and calls `build_dependency_graph()`. -5. `verify_namespaces()` — groups discovered component IDs by their top-level namespace (`main`, `deps`, `vendor`) and checks expected per-namespace component counts. -6. `verify_cross_path_dependencies()` — inspects component `dependencies` for edges that cross namespace boundaries. -7. `verify_no_warnings()` — checks the builder for any recorded warnings. -8. `verify_file_counts()` — checks the total component count against the expected total (14). -9. `print_detailed_output()` — dumps all component IDs, cross-namespace dependencies, and per-namespace counts for manual inspection. -10. `cleanup()` — removes the temporary test directory (always executed via `finally`). - -### TestResults (integration_test.py) - -A richer accumulator than the scenario-suite version — each recorded test entry carries an optional `details` string, allowing informational warnings to be attached to otherwise-passing assertions: - -```text -class TestResults: - def __init__(self): - self.tests = [] # list of {"name", "passed", "details"} - - def add_test(self, name, passed, details="") -> None - def print_summary(self) -> None # prints PASS/FAIL + warnings block - def all_passed(self) -> bool -``` - -`IntegrationTestRunner.run()` uses `all_passed()` to determine the final process exit code (`0` on success, `1` on any failed assertion, `2` if the test crashes with an unhandled exception). - -## Architecture - -### Class Relationships - -```mermaid -classDiagram - class Colors { - +GREEN - +RED - +YELLOW - +BLUE - +BOLD - +END - } - class ScenarioTestResults { - -results dict - +add_test(name, passed) - +print_summary() int - } - class IntegrationTestRunner { - -test_dir - -main_path - -deps_path - -vendor_path - -config - -builder - -components - -leaf_nodes - -results - +setup_test_environment() - +create_config() - +validate_paths() - +execute_dependency_parser() - +verify_namespaces() - +verify_cross_path_dependencies() - +verify_no_warnings() - +verify_file_counts() - +print_detailed_output() - +cleanup() - +run() int - } - class IntegrationTestResults { - -tests list - +add_test(name, passed, details) - +print_summary() - +all_passed() bool - } - IntegrationTestRunner --> IntegrationTestResults : records into - IntegrationTestRunner --> Config : creates - IntegrationTestRunner --> DependencyGraphBuilder : invokes - ScenarioTestResults ..> Colors : uses for output -``` - -### Scenario Test Suite Flow (`test_multi_path.py`) - -```mermaid -flowchart TD - Start["run_all_tests()"] --> T1["test_single_path()"] - Start --> T2["test_multiple_paths()"] - Start --> T3["test_component_namespacing()"] - Start --> T4["test_cross_path_dependencies()"] - Start --> T5["test_invalid_path_handling()"] - Start --> T6["test_empty_additional_paths()"] - Start --> T7["test_relative_vs_absolute_paths()"] - T1 --> Builder["DependencyGraphBuilder.build_dependency_graph()"] - T2 --> Builder - T3 --> Builder - T4 --> Builder - T6 --> Builder - T7 --> Builder - T5 --> Validate["Config.validate_source_paths()"] - Builder --> Results["TestResults.add_test()"] - Validate --> Results - Results --> Summary["TestResults.print_summary()"] - Summary --> ExitCode["Process exit code (0 or 1)"] -``` - -Each scenario test constructs its own `Config` (see [Config Core](config-core.md)) via the local `create_test_config()` helper, pointing `repo_path` at the `main/` fixture directory and, where relevant, supplying `additional_source_paths` pointing at `deps/` and/or `external/` fixture directories (see [Sample Fixtures](sample_fixtures.md)). - -### Integration Test Pipeline (`integration_test.py`) - -```mermaid -sequenceDiagram - participant Main as "main()" - participant Runner as "IntegrationTestRunner" - participant Cfg as "Config" - participant Builder as "DependencyGraphBuilder" - participant Results as "IntegrationTestResults" - - Main->>Runner: run() - Runner->>Runner: setup_test_environment() - Note over Runner: Creates main/, deps/, vendor/ with cross-importing fixture files - Runner->>Cfg: create_config() - Cfg-->>Runner: Config(repo_path, additional_source_paths) - Runner->>Runner: validate_paths() - Runner->>Builder: execute_dependency_parser() - Builder->>Builder: build_dependency_graph() - Builder-->>Runner: components, leaf_nodes - Runner->>Results: verify_namespaces() - Runner->>Results: verify_cross_path_dependencies() - Runner->>Results: verify_no_warnings() - Runner->>Results: verify_file_counts() - Runner->>Runner: print_detailed_output() - Runner->>Results: print_summary() - Results-->>Runner: all_passed() - Runner->>Runner: cleanup() - Runner-->>Main: exit code 0, 1, or 2 -``` - -### Dependency on Backend Core - -```mermaid -flowchart LR - subgraph TS["Test Suites"] - SM["Scenario Test Suite (test_multi_path.py)"] - IT["Integration Test Runner (integration_test.py)"] - end - subgraph CC["Config Core"] - Config["Config"] - end - subgraph DA["Dependency Analyzer Core"] - DGB["DependencyGraphBuilder"] - end - SM --> Config - IT --> Config - SM --> DGB - IT --> DGB -``` - -Both scripts treat `DependencyGraphBuilder` (see [Backend Core](backend-core.md)) as the system under test and `Config` (see [Config Core](config-core.md)) as the primary input mechanism for expressing multi-path analysis intent via `additional_source_paths` and `is_multi_path_mode()`. - -## Test Coverage Summary - -### Scenario Test Suite (`test_multi_path.py`) - -| # | Test | Validates | -|---|---|---| -| 1 | Single Path (Backward Compatibility) | Analyzing one `repo_path` with no additional paths still discovers expected components | -| 2 | Multiple Paths With Unique Components | Components from `main/`, `deps/`, and `external/` are all discovered when passed as `additional_source_paths` | -| 3 | Component ID Namespacing | No duplicate component IDs occur across paths; IDs reflect their source path | -| 4 | Cross-Path Dependencies | Dependency edges are built between components across different source paths | -| 5 | Invalid Path Handling | `Config.validate_source_paths()` raises `ValueError`/`OSError` for a nonexistent additional path | -| 6 | Empty Additional Paths | An empty `additional_source_paths` list behaves identically to single-path mode | -| 7 | Relative vs. Absolute Paths | Absolute additional paths resolve correctly (relative-path resolution is noted as needing further logic) | - -### Integration Test (`integration_test.py`) - -| Step | Validates | -|---|---| -| Path validation | Root and all additional paths exist on disk before analysis begins | -| Multi-path mode detection | `config.is_multi_path_mode()` returns `True` when additional paths are configured | -| Namespace presence | All expected top-level namespaces (`main`, `deps`, `vendor`) appear in the component graph | -| Namespace component counts | Each namespace yields the expected number of class/function-level components (`main`: 7, `deps`: 5, `vendor`: 2 — 14 total) | -| Cross-namespace dependencies | Dependency edges spanning namespaces are detected where present; documents current limitation that import-level cross-path resolution is not yet fully implemented | -| Warning tracking | No unexpected warnings are recorded by the builder during analysis | - -## Related Modules - -- [Test Multi Path](test-multi-path.md) — parent module; defines the overall multi-path test harness this module belongs to. -- [Sample Fixtures](sample_fixtures.md) — sibling module providing on-disk sample components (`MainService`, `APIController`, `DataPlugin`, `PluginInterface`) referenced by the scenario test suite's `main/` and `external/` fixture directories. -- [Config Core](config-core.md) — supplies the `Config` class used to declare `repo_path` and `additional_source_paths` for every test scenario. -- [Backend Core](backend-core.md) — houses the dependency analyzer pipeline, including the `DependencyGraphBuilder` under test. From 581ac8efd3ff917fe2a941b2ea0fd149613f5ee3 Mon Sep 17 00:00:00 2001 From: "github-actions[bot]" Date: Fri, 18 Sep 2026 01:04:39 +0000 Subject: [PATCH 3/4] chore(docs): Remove 1 orphaned inline files [skip ci] --- .flamingo-ai-technical-writer-status.md | 5 ----- 1 file changed, 5 deletions(-) delete mode 100644 .flamingo-ai-technical-writer-status.md diff --git a/.flamingo-ai-technical-writer-status.md b/.flamingo-ai-technical-writer-status.md deleted file mode 100644 index 132db0fa..00000000 --- a/.flamingo-ai-technical-writer-status.md +++ /dev/null @@ -1,5 +0,0 @@ -# 🦩 Flamingo Code Documentation: Started - -Run ID: doc-orchestrator-1789693445469 -Status: In Progress -Started: 2026-09-18 01:04:22 UTC From 9146af8beac00b179a5f27eba04bb3ef1430db6a Mon Sep 17 00:00:00 2001 From: "github-actions[bot]" Date: Fri, 18 Sep 2026 01:06:17 +0000 Subject: [PATCH 4/4] chore(docs): restore documentation no stage regenerated [skip ci] --- docs/development/.gitignore | 7 + docs/development/README.md | 23 ++ docs/development/architecture/README.md | 251 ++++++++++++++++ docs/development/setup/environment.md | 58 ++++ docs/development/setup/local-development.md | 90 ++++++ docs/diagrams/architecture/.gitignore | 8 + docs/diagrams/architecture/README-2.mmd | 17 ++ docs/diagrams/architecture/README-3.mmd | 18 ++ docs/diagrams/architecture/README.mmd | 29 ++ .../architecture/agent-tools-core-2.mmd | 17 ++ .../architecture/agent-tools-core-3.mmd | 13 + .../architecture/agent-tools-core.mmd | 12 + .../architecture/analysis_pipeline-2.mmd | 67 +++++ .../architecture/analysis_pipeline-3.mmd | 27 ++ .../architecture/analysis_pipeline-4.mmd | 10 + .../architecture/analysis_pipeline-5.mmd | 11 + .../architecture/analysis_pipeline-6.mmd | 24 ++ .../architecture/analysis_pipeline-7.mmd | 9 + .../architecture/analysis_pipeline.mmd | 12 + docs/diagrams/architecture/backend-core-2.mmd | 18 ++ docs/diagrams/architecture/backend-core-3.mmd | 10 + docs/diagrams/architecture/backend-core.mmd | 20 ++ .../architecture/c_family_analyzers-2.mmd | 17 ++ .../architecture/c_family_analyzers-3.mmd | 50 ++++ .../architecture/c_family_analyzers-4.mmd | 13 + .../architecture/c_family_analyzers.mmd | 8 + docs/diagrams/architecture/cli-core.mmd | 27 ++ docs/diagrams/architecture/config-core-2.mmd | 9 + docs/diagrams/architecture/config-core-3.mmd | 10 + docs/diagrams/architecture/config-core-4.mmd | 12 + docs/diagrams/architecture/config-core-5.mmd | 21 ++ docs/diagrams/architecture/config-core-6.mmd | 16 + docs/diagrams/architecture/config-core.mmd | 43 +++ .../diagrams/architecture/configuration-2.mmd | 42 +++ .../diagrams/architecture/configuration-3.mmd | 20 ++ .../diagrams/architecture/configuration-4.mmd | 20 ++ .../diagrams/architecture/configuration-5.mmd | 16 + .../diagrams/architecture/configuration-6.mmd | 8 + docs/diagrams/architecture/configuration.mmd | 16 + .../dependency-analyzer-core-2.mmd | 22 ++ .../architecture/dependency-analyzer-core.mmd | 21 ++ .../dependency-analyzer-models-2.mmd | 60 ++++ .../dependency-analyzer-models.mmd | 14 + .../documentation-generator-2.mmd | 43 +++ .../documentation-generator-3.mmd | 13 + .../architecture/documentation-generator.mmd | 11 + .../diagrams/architecture/frontend-core-2.mmd | 24 ++ .../diagrams/architecture/frontend-core-3.mmd | 51 ++++ .../diagrams/architecture/frontend-core-4.mmd | 8 + docs/diagrams/architecture/frontend-core.mmd | 25 ++ docs/diagrams/architecture/generation-2.mmd | 26 ++ docs/diagrams/architecture/generation-3.mmd | 14 + docs/diagrams/architecture/generation.mmd | 11 + .../architecture/git_integration-2.mmd | 26 ++ .../architecture/git_integration-3.mmd | 13 + .../diagrams/architecture/git_integration.mmd | 18 ++ .../architecture/graph_construction-2.mmd | 19 ++ .../architecture/graph_construction-3.mmd | 15 + .../architecture/graph_construction-4.mmd | 8 + .../architecture/graph_construction-5.mmd | 22 ++ .../architecture/graph_construction.mmd | 11 + .../architecture/html_generation-2.mmd | 19 ++ .../architecture/html_generation-3.mmd | 17 ++ .../diagrams/architecture/html_generation.mmd | 10 + .../diagrams/architecture/java_analyzer-2.mmd | 16 + .../diagrams/architecture/java_analyzer-3.mmd | 6 + docs/diagrams/architecture/java_analyzer.mmd | 19 ++ .../javascript_typescript_analyzers-2.mmd | 45 +++ .../javascript_typescript_analyzers-3.mmd | 19 ++ .../javascript_typescript_analyzers-4.mmd | 17 ++ .../javascript_typescript_analyzers-5.mmd | 19 ++ .../javascript_typescript_analyzers.mmd | 10 + docs/diagrams/architecture/job_models-2.mmd | 7 + docs/diagrams/architecture/job_models-3.mmd | 20 ++ docs/diagrams/architecture/job_models-4.mmd | 15 + docs/diagrams/architecture/job_models.mmd | 52 ++++ docs/diagrams/architecture/llm-services-2.mmd | 20 ++ docs/diagrams/architecture/llm-services-3.mmd | 22 ++ docs/diagrams/architecture/llm-services-4.mmd | 16 + docs/diagrams/architecture/llm-services-5.mmd | 24 ++ docs/diagrams/architecture/llm-services.mmd | 24 ++ .../architecture/logging-config-2.mmd | 21 ++ .../architecture/logging-config-3.mmd | 13 + .../architecture/logging-config-4.mmd | 12 + docs/diagrams/architecture/logging-config.mmd | 14 + docs/diagrams/architecture/php_analyzer-2.mmd | 41 +++ docs/diagrams/architecture/php_analyzer-3.mmd | 20 ++ docs/diagrams/architecture/php_analyzer-4.mmd | 7 + docs/diagrams/architecture/php_analyzer.mmd | 9 + .../architecture/python_analyzer-2.mmd | 4 + .../architecture/python_analyzer-3.mmd | 17 ++ .../architecture/python_analyzer-4.mmd | 6 + .../diagrams/architecture/python_analyzer.mmd | 11 + .../architecture/sample_fixtures-2.mmd | 28 ++ .../architecture/sample_fixtures-3.mmd | 11 + .../architecture/sample_fixtures-4.mmd | 5 + .../architecture/sample_fixtures-5.mmd | 4 + .../diagrams/architecture/sample_fixtures.mmd | 4 + .../architecture/test-clustering-2.mmd | 23 ++ .../architecture/test-clustering-3.mmd | 15 + .../diagrams/architecture/test-clustering.mmd | 37 +++ .../architecture/test-multi-path-2.mmd | 12 + .../diagrams/architecture/test-multi-path.mmd | 25 ++ docs/diagrams/architecture/test_suites-2.mmd | 19 ++ docs/diagrams/architecture/test_suites-3.mmd | 25 ++ docs/diagrams/architecture/test_suites-4.mmd | 15 + docs/diagrams/architecture/test_suites.mmd | 46 +++ .../architecture/tree-sitter-analyzers.mmd | 22 ++ docs/diagrams/architecture/utils-2.mmd | 22 ++ docs/diagrams/architecture/utils-3.mmd | 11 + docs/diagrams/architecture/utils.mmd | 36 +++ docs/getting-started/.gitignore | 7 + docs/getting-started/first-steps.md | 95 ++++++ docs/getting-started/introduction.md | 85 ++++++ docs/getting-started/prerequisites.md | 71 +++++ docs/getting-started/quick-start.md | 114 ++++++++ docs/reference/architecture/.gitignore | 8 + docs/reference/architecture/README.md | 123 ++++++++ .../agent-tools-core/agent-tools-core.md | 123 ++++++++ .../architecture/backend-core/backend-core.md | 98 +++++++ .../analysis_pipeline.md | 274 ++++++++++++++++++ .../dependency-analyzer-core.md | 124 ++++++++ .../graph_construction.md | 218 ++++++++++++++ .../dependency-analyzer-models.md | 194 +++++++++++++ .../documentation-generator.md | 228 +++++++++++++++ .../backend-core/llm-services/llm-services.md | 224 ++++++++++++++ .../logging-config/logging-config.md | 181 ++++++++++++ .../c_family_analyzers.md | 260 +++++++++++++++++ .../tree-sitter-analyzers/java_analyzer.md | 178 ++++++++++++ .../javascript_typescript_analyzers.md | 229 +++++++++++++++ .../tree-sitter-analyzers/php_analyzer.md | 225 ++++++++++++++ .../tree-sitter-analyzers/python_analyzer.md | 196 +++++++++++++ .../tree-sitter-analyzers.md | 115 ++++++++ .../architecture/cli-core/cli-core.md | 103 +++++++ .../cli-core/configuration/configuration.md | 249 ++++++++++++++++ .../cli-core/generation/generation.md | 167 +++++++++++ .../git_integration/git_integration.md | 147 ++++++++++ .../html_generation/html_generation.md | 131 +++++++++ .../cli-core/job_models/job_models.md | 244 ++++++++++++++++ .../cli-core/cli-core/utils/utils.md | 168 +++++++++++ .../cli-core/cli_core/utils/utils.md | 168 +++++++++++ .../architecture/cli-core/configuration.md | 249 ++++++++++++++++ .../architecture/cli-core/generation.md | 167 +++++++++++ .../architecture/cli-core/git_integration.md | 147 ++++++++++ .../architecture/cli-core/html_generation.md | 131 +++++++++ .../architecture/cli-core/job_models.md | 244 ++++++++++++++++ .../architecture/config-core/config-core.md | 220 ++++++++++++++ .../frontend-core/frontend-core.md | 257 ++++++++++++++++ .../test-clustering/test-clustering.md | 208 +++++++++++++ .../test-multi-path/sample_fixtures.md | 125 ++++++++ .../test-multi-path/test-multi-path.md | 87 ++++++ .../test-multi-path/test_suites.md | 261 +++++++++++++++++ 152 files changed, 9349 insertions(+) create mode 100644 docs/development/.gitignore create mode 100644 docs/development/README.md create mode 100644 docs/development/architecture/README.md create mode 100644 docs/development/setup/environment.md create mode 100644 docs/development/setup/local-development.md create mode 100644 docs/diagrams/architecture/.gitignore create mode 100644 docs/diagrams/architecture/README-2.mmd create mode 100644 docs/diagrams/architecture/README-3.mmd create mode 100644 docs/diagrams/architecture/README.mmd create mode 100644 docs/diagrams/architecture/agent-tools-core-2.mmd create mode 100644 docs/diagrams/architecture/agent-tools-core-3.mmd create mode 100644 docs/diagrams/architecture/agent-tools-core.mmd create mode 100644 docs/diagrams/architecture/analysis_pipeline-2.mmd create mode 100644 docs/diagrams/architecture/analysis_pipeline-3.mmd create mode 100644 docs/diagrams/architecture/analysis_pipeline-4.mmd create mode 100644 docs/diagrams/architecture/analysis_pipeline-5.mmd create mode 100644 docs/diagrams/architecture/analysis_pipeline-6.mmd create mode 100644 docs/diagrams/architecture/analysis_pipeline-7.mmd create mode 100644 docs/diagrams/architecture/analysis_pipeline.mmd create mode 100644 docs/diagrams/architecture/backend-core-2.mmd create mode 100644 docs/diagrams/architecture/backend-core-3.mmd create mode 100644 docs/diagrams/architecture/backend-core.mmd create mode 100644 docs/diagrams/architecture/c_family_analyzers-2.mmd create mode 100644 docs/diagrams/architecture/c_family_analyzers-3.mmd create mode 100644 docs/diagrams/architecture/c_family_analyzers-4.mmd create mode 100644 docs/diagrams/architecture/c_family_analyzers.mmd create mode 100644 docs/diagrams/architecture/cli-core.mmd create mode 100644 docs/diagrams/architecture/config-core-2.mmd create mode 100644 docs/diagrams/architecture/config-core-3.mmd create mode 100644 docs/diagrams/architecture/config-core-4.mmd create mode 100644 docs/diagrams/architecture/config-core-5.mmd create mode 100644 docs/diagrams/architecture/config-core-6.mmd create mode 100644 docs/diagrams/architecture/config-core.mmd create mode 100644 docs/diagrams/architecture/configuration-2.mmd create mode 100644 docs/diagrams/architecture/configuration-3.mmd create mode 100644 docs/diagrams/architecture/configuration-4.mmd create mode 100644 docs/diagrams/architecture/configuration-5.mmd create mode 100644 docs/diagrams/architecture/configuration-6.mmd create mode 100644 docs/diagrams/architecture/configuration.mmd create mode 100644 docs/diagrams/architecture/dependency-analyzer-core-2.mmd create mode 100644 docs/diagrams/architecture/dependency-analyzer-core.mmd create mode 100644 docs/diagrams/architecture/dependency-analyzer-models-2.mmd create mode 100644 docs/diagrams/architecture/dependency-analyzer-models.mmd create mode 100644 docs/diagrams/architecture/documentation-generator-2.mmd create mode 100644 docs/diagrams/architecture/documentation-generator-3.mmd create mode 100644 docs/diagrams/architecture/documentation-generator.mmd create mode 100644 docs/diagrams/architecture/frontend-core-2.mmd create mode 100644 docs/diagrams/architecture/frontend-core-3.mmd create mode 100644 docs/diagrams/architecture/frontend-core-4.mmd create mode 100644 docs/diagrams/architecture/frontend-core.mmd create mode 100644 docs/diagrams/architecture/generation-2.mmd create mode 100644 docs/diagrams/architecture/generation-3.mmd create mode 100644 docs/diagrams/architecture/generation.mmd create mode 100644 docs/diagrams/architecture/git_integration-2.mmd create mode 100644 docs/diagrams/architecture/git_integration-3.mmd create mode 100644 docs/diagrams/architecture/git_integration.mmd create mode 100644 docs/diagrams/architecture/graph_construction-2.mmd create mode 100644 docs/diagrams/architecture/graph_construction-3.mmd create mode 100644 docs/diagrams/architecture/graph_construction-4.mmd create mode 100644 docs/diagrams/architecture/graph_construction-5.mmd create mode 100644 docs/diagrams/architecture/graph_construction.mmd create mode 100644 docs/diagrams/architecture/html_generation-2.mmd create mode 100644 docs/diagrams/architecture/html_generation-3.mmd create mode 100644 docs/diagrams/architecture/html_generation.mmd create mode 100644 docs/diagrams/architecture/java_analyzer-2.mmd create mode 100644 docs/diagrams/architecture/java_analyzer-3.mmd create mode 100644 docs/diagrams/architecture/java_analyzer.mmd create mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers-2.mmd create mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers-3.mmd create mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers-4.mmd create mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers-5.mmd create mode 100644 docs/diagrams/architecture/javascript_typescript_analyzers.mmd create mode 100644 docs/diagrams/architecture/job_models-2.mmd create mode 100644 docs/diagrams/architecture/job_models-3.mmd create mode 100644 docs/diagrams/architecture/job_models-4.mmd create mode 100644 docs/diagrams/architecture/job_models.mmd create mode 100644 docs/diagrams/architecture/llm-services-2.mmd create mode 100644 docs/diagrams/architecture/llm-services-3.mmd create mode 100644 docs/diagrams/architecture/llm-services-4.mmd create mode 100644 docs/diagrams/architecture/llm-services-5.mmd create mode 100644 docs/diagrams/architecture/llm-services.mmd create mode 100644 docs/diagrams/architecture/logging-config-2.mmd create mode 100644 docs/diagrams/architecture/logging-config-3.mmd create mode 100644 docs/diagrams/architecture/logging-config-4.mmd create mode 100644 docs/diagrams/architecture/logging-config.mmd create mode 100644 docs/diagrams/architecture/php_analyzer-2.mmd create mode 100644 docs/diagrams/architecture/php_analyzer-3.mmd create mode 100644 docs/diagrams/architecture/php_analyzer-4.mmd create mode 100644 docs/diagrams/architecture/php_analyzer.mmd create mode 100644 docs/diagrams/architecture/python_analyzer-2.mmd create mode 100644 docs/diagrams/architecture/python_analyzer-3.mmd create mode 100644 docs/diagrams/architecture/python_analyzer-4.mmd create mode 100644 docs/diagrams/architecture/python_analyzer.mmd create mode 100644 docs/diagrams/architecture/sample_fixtures-2.mmd create mode 100644 docs/diagrams/architecture/sample_fixtures-3.mmd create mode 100644 docs/diagrams/architecture/sample_fixtures-4.mmd create mode 100644 docs/diagrams/architecture/sample_fixtures-5.mmd create mode 100644 docs/diagrams/architecture/sample_fixtures.mmd create mode 100644 docs/diagrams/architecture/test-clustering-2.mmd create mode 100644 docs/diagrams/architecture/test-clustering-3.mmd create mode 100644 docs/diagrams/architecture/test-clustering.mmd create mode 100644 docs/diagrams/architecture/test-multi-path-2.mmd create mode 100644 docs/diagrams/architecture/test-multi-path.mmd create mode 100644 docs/diagrams/architecture/test_suites-2.mmd create mode 100644 docs/diagrams/architecture/test_suites-3.mmd create mode 100644 docs/diagrams/architecture/test_suites-4.mmd create mode 100644 docs/diagrams/architecture/test_suites.mmd create mode 100644 docs/diagrams/architecture/tree-sitter-analyzers.mmd create mode 100644 docs/diagrams/architecture/utils-2.mmd create mode 100644 docs/diagrams/architecture/utils-3.mmd create mode 100644 docs/diagrams/architecture/utils.mmd create mode 100644 docs/getting-started/.gitignore create mode 100644 docs/getting-started/first-steps.md create mode 100644 docs/getting-started/introduction.md create mode 100644 docs/getting-started/prerequisites.md create mode 100644 docs/getting-started/quick-start.md create mode 100644 docs/reference/architecture/.gitignore create mode 100644 docs/reference/architecture/README.md create mode 100644 docs/reference/architecture/backend-core/agent-tools-core/agent-tools-core.md create mode 100644 docs/reference/architecture/backend-core/backend-core.md create mode 100644 docs/reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md create mode 100644 docs/reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md create mode 100644 docs/reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md create mode 100644 docs/reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md create mode 100644 docs/reference/architecture/backend-core/documentation-generator/documentation-generator.md create mode 100644 docs/reference/architecture/backend-core/llm-services/llm-services.md create mode 100644 docs/reference/architecture/backend-core/logging-config/logging-config.md create mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md create mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md create mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md create mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md create mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md create mode 100644 docs/reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md create mode 100644 docs/reference/architecture/cli-core/cli-core.md create mode 100644 docs/reference/architecture/cli-core/cli-core/configuration/configuration.md create mode 100644 docs/reference/architecture/cli-core/cli-core/generation/generation.md create mode 100644 docs/reference/architecture/cli-core/cli-core/git_integration/git_integration.md create mode 100644 docs/reference/architecture/cli-core/cli-core/html_generation/html_generation.md create mode 100644 docs/reference/architecture/cli-core/cli-core/job_models/job_models.md create mode 100644 docs/reference/architecture/cli-core/cli-core/utils/utils.md create mode 100644 docs/reference/architecture/cli-core/cli_core/utils/utils.md create mode 100644 docs/reference/architecture/cli-core/configuration.md create mode 100644 docs/reference/architecture/cli-core/generation.md create mode 100644 docs/reference/architecture/cli-core/git_integration.md create mode 100644 docs/reference/architecture/cli-core/html_generation.md create mode 100644 docs/reference/architecture/cli-core/job_models.md create mode 100644 docs/reference/architecture/config-core/config-core.md create mode 100644 docs/reference/architecture/frontend-core/frontend-core.md create mode 100644 docs/reference/architecture/test-clustering/test-clustering.md create mode 100644 docs/reference/architecture/test-multi-path/sample_fixtures.md create mode 100644 docs/reference/architecture/test-multi-path/test-multi-path.md create mode 100644 docs/reference/architecture/test-multi-path/test_suites.md diff --git a/docs/development/.gitignore b/docs/development/.gitignore new file mode 100644 index 00000000..a5d01be5 --- /dev/null +++ b/docs/development/.gitignore @@ -0,0 +1,7 @@ +# VoltAgent temp files +temp/ + +# JSON intermediate files (except schema/config) +*.json +!*-schema.json +!*-config.json diff --git a/docs/development/README.md b/docs/development/README.md new file mode 100644 index 00000000..74ab77c4 --- /dev/null +++ b/docs/development/README.md @@ -0,0 +1,23 @@ +# Development Documentation + +This section covers everything you need to develop, test, and contribute to **CodeWiki** itself — the AI-powered documentation generator, not a repository you're documenting *with* CodeWiki. + +CodeWiki is a Python package (`pyproject.toml`, Python `>=3.12`) organized around a CLI (`codewiki/cli`), a backend documentation pipeline (`codewiki/src/be`), a FastAPI web application (`codewiki/src/fe`), and a shared runtime configuration model (`codewiki/src/config.py`). + +## Quick Navigation + +| Guide | Description | +|---|---| +| [Environment Setup](setup/environment.md) | IDE recommendations, required tools, and development environment variables. | +| [Local Development](setup/local-development.md) | Cloning the repo, installing in editable mode, running the CLI and web app locally, and debugging. | +| [Architecture Overview](architecture/README.md) | High-level module architecture, core components, and data flow through the documentation pipeline. | +| [Security](security/README.md) | Credential storage, safe file access, input validation, and secure coding practices used in CodeWiki. | +| [Testing](testing/README.md) | Structure of the diagnostic/validation scripts, how to run them, and expectations for new tests. | +| [Contributing Guidelines](contributing/guidelines.md) | Code style, branch naming, commit format, and the review checklist for pull requests. | + +## Where to Start + +1. Read the [Architecture Overview](architecture/README.md) to understand how the CLI, backend, and frontend modules fit together. +2. Follow [Environment Setup](setup/environment.md) and [Local Development](setup/local-development.md) to get a working local copy of CodeWiki. +3. Review [Security](security/README.md) and [Testing](testing/README.md) before submitting changes. +4. Read [Contributing Guidelines](contributing/guidelines.md) before opening a pull request. diff --git a/docs/development/architecture/README.md b/docs/development/architecture/README.md new file mode 100644 index 00000000..db28893a --- /dev/null +++ b/docs/development/architecture/README.md @@ -0,0 +1,251 @@ +# Architecture Overview + +CodeWiki is a layered, AI-driven documentation engine. This document describes the high-level architecture, core components, data flow, and key design decisions. + +For detailed module-level documentation, see the [Reference Architecture](../../reference/architecture/README.md). + +--- + +## High-Level Architecture + +CodeWiki is organized into six core layers, each with clearly defined responsibilities: + +```mermaid +graph TD + User["User (CLI or Web)"] --> CLICore["CLI Core"] + User --> FECore["Frontend Core (FastAPI)"] + + CLICore --> DocGen["Documentation Generator"] + FECore --> DocGen + + DocGen --> DepAnalysis["Dependency Analysis"] + DocGen --> AgentOrch["Agent Orchestration"] + + AgentOrch --> LLMSvc["LLM Services"] + AgentOrch --> Tools["Agent Tools"] + + DepAnalysis --> GraphBuilder["Dependency Graph Builder"] + GraphBuilder --> DocGen + + DocGen --> Utils["Utils (FileManager)"] + DocGen --> Output["Generated Documentation"] + + LLMSvc --> Provider["External LLM Providers"] +``` + +--- + +## Core Components + +| Component | Location | Responsibility | +|-----------|----------|---------------| +| **CLI Core** | `codewiki/cli/` | Command-line interface, config management, Git integration, HTML generation | +| **Frontend Core** | `codewiki/src/fe/` | FastAPI web app, job queue, caching, Markdown rendering | +| **Documentation Generator** | `codewiki/src/be/documentation_generator.py` | Orchestrates the full generation pipeline | +| **Agent Orchestration** | `codewiki/src/be/agent_orchestrator.py` | AI agent lifecycle: creation, tooling, execution, persistence | +| **LLM Services** | `codewiki/src/be/llm_services.py` | LLM provider factory, fallback chaining, request counting | +| **Dependency Analysis** | `codewiki/src/be/dependency_analyzer/` | AST parsing, call graph resolution, dependency graph building | +| **Utils** | `codewiki/src/utils.py` | Stateless `FileManager` for all file and JSON I/O | + +--- + +## Generation Pipeline (Data Flow) + +The complete documentation generation pipeline runs in a deterministic sequence of stages: + +```mermaid +flowchart TD + Start["Start Job"] --> DepAnalysis["1. Dependency Analysis"] + DepAnalysis --> Cluster["2. Module Clustering (LLM)"] + Cluster --> LeafDocs["3. Generate Leaf Module Docs"] + LeafDocs --> ParentDocs["4. Generate Parent Module Docs"] + ParentDocs --> Overview["5. Generate Repository Overview"] + Overview --> Metadata["6. Write metadata.json"] + Metadata --> EndNode["Done"] +``` + +### Stage Descriptions + +| Stage | Weight | Description | +|-------|--------|-------------| +| **Dependency Analysis** | 40% | Parse all source files, extract components and relationships, build dependency graph | +| **Module Clustering** | 20% | Use LLM to group related components into logical modules | +| **Documentation Generation** | 30% | Generate Markdown documentation, leaf-first | +| **HTML Generation** | 5% | Build static `index.html` for GitHub Pages (optional) | +| **Finalization** | 5% | Write `metadata.json`, commit to Git branch (optional) | + +--- + +## Dependency Analysis Pipeline + +The Dependency Analysis layer transforms raw source code into a structured dependency graph: + +```mermaid +flowchart LR + Repo["Repository"] --> RepoAnalyzer["Repo Analyzer"] + RepoAnalyzer --> FileTree["File Tree"] + FileTree --> CallGraphAnalyzer["Call Graph Analyzer"] + CallGraphAnalyzer --> LangAnalyzers["Language Analyzers"] + LangAnalyzers --> Nodes["Node Models"] + LangAnalyzers --> Relationships["CallRelationship Models"] + Nodes --> DepParser["Dependency Parser"] + Relationships --> DepParser + DepParser --> GraphBuilder["Dependency Graph Builder"] + GraphBuilder --> AnalysisResult["AnalysisResult"] +``` + +**Supported languages:** + +| Language | Analyzer | +|---------|---------| +| Python | `PythonASTAnalyzer` (native `ast` module) | +| JavaScript | `TreeSitterJSAnalyzer` | +| TypeScript | `TreeSitterTSAnalyzer` | +| Java | `TreeSitterJavaAnalyzer` | +| C# | `TreeSitterCSharpAnalyzer` | +| C | `TreeSitterCAnalyzer` | +| C++ | `TreeSitterCppAnalyzer` | +| PHP | `TreeSitterPHPAnalyzer` | + +--- + +## Agent Orchestration + +The Agent Orchestration layer manages the AI agent lifecycle for each module: + +```mermaid +flowchart TD + Start["Process Module"] --> CheckComplex{"is_complex_module()?"} + CheckComplex -->|"Yes"| ComplexAgent["Create Complex Agent"] + CheckComplex -->|"No"| LeafAgent["Create Leaf Agent"] + + ComplexAgent --> Tools1["read_code_components + str_replace_editor + generate_sub_module_documentation"] + LeafAgent --> Tools2["read_code_components + str_replace_editor"] + + Tools1 --> Execute["Run Agent (async)"] + Tools2 --> Execute + Execute --> Verify["Verify Output File"] + Verify --> Persist["Save module_tree.json"] +``` + +**Idempotency:** Documentation generation is skipped if a module's output file already exists. This allows safe re-runs and incremental updates. + +--- + +## LLM Services Architecture + +The LLM Services layer abstracts provider differences and manages three model roles: + +```mermaid +flowchart TD + Request["Generation Request"] --> TryMain["Main Model"] + TryMain -->|"Success"| ReturnMain["Return Response"] + TryMain -->|"Failure"| TryFallback["Fallback Model"] + TryFallback --> ReturnFallback["Return Fallback Response"] + + ClusterRequest["Cluster Request"] --> ClusterModel["Cluster Model"] + ClusterModel --> ClusterResponse["Cluster Response"] +``` + +| Role | Default | Purpose | +|------|---------|---------| +| **Main model** | `claude-sonnet-4` | Primary documentation generation | +| **Cluster model** | Same as main | Module grouping and structural reasoning | +| **Fallback model** | Same as main | Backup if main model fails | + +--- + +## Frontend Core Architecture + +```mermaid +sequenceDiagram + participant Browser + participant Routes as "WebRoutes" + participant Worker as "BackgroundWorker" + participant Cache as "CacheManager" + participant GitHub as "GitHubRepoProcessor" + participant DocGen as "DocumentationGenerator" + + Browser->>Routes: POST / (repo_url) + Routes->>Cache: Check cache (SHA-256 hash) + Cache-->>Routes: Hit or miss + Routes->>Worker: Enqueue job + Worker->>GitHub: Clone repository + Worker->>DocGen: Run documentation generation + DocGen-->>Worker: Write docs to output directory + Worker->>Cache: Store result + Browser->>Routes: GET /docs/{job_id} + Routes-->>Browser: Rendered HTML +``` + +--- + +## Output Directory Structure + +All generated documentation follows a deterministic hierarchical structure: + +```text +docs/ +├── README.md ← Repository overview +├── module_tree.json ← Module hierarchy +├── metadata.json ← Generation audit log +└── / + ├── .md ← Module documentation + └── / + └── .md ← Sub-module documentation +``` + +This structure mirrors the logical module hierarchy discovered from the codebase — not the raw file system layout. + +--- + +## Key Design Decisions + +### 1. Leaf-First (Bottom-Up) Generation + +Parent modules are documented *after* their children. This ensures: + +- LLM context stays bounded — each call only receives one module's context +- Parent overviews accurately summarize child documentation +- Generation is parallelizable at the leaf level + +### 2. Synthetic Module Fallback + +If the clustering LLM fails to produce any modules, CodeWiki automatically creates synthetic modules grouped by top-level directory. This prevents the system from attempting to document an entire large repository in a single LLM call. + +### 3. OS Keychain for API Keys + +API keys are never written to disk in plaintext. The CLI uses the system keychain (`keyring` library), which maps to macOS Keychain, Windows Credential Manager, or Linux Secret Service depending on the platform. + +### 4. Provider-Agnostic LLM Layer + +The LLM Services layer uses the OpenAI SDK for all providers, with configurable base URLs. This allows CodeWiki to work with Anthropic, Azure, LiteLLM proxies, and any OpenAI-compatible endpoint without code changes. + +### 5. Idempotent Generation + +The `process_module()` method skips generation if the output file already exists. This makes re-runs safe and supports incremental documentation updates. + +### 6. Deterministic Namespacing + +All component identifiers follow the pattern: + +```text +... +``` + +This allows cross-language and cross-repository dependency resolution without ambiguity. + +--- + +## Reference Documentation + +For deep-dive module documentation, see: + +- [Reference Architecture Overview](../../reference/architecture/README.md) +- [CLI Core](../../reference/architecture/cli-core/cli-core.md) +- [Agent Orchestration](../../reference/architecture/agent-orchestration/agent-orchestration.md) +- [LLM Services](../../reference/architecture/llm-services/llm-services.md) +- [Dependency Analysis](../../reference/architecture/dependency-analysis/dependency-analysis.md) +- [Documentation Generation](../../reference/architecture/documentation-generation/documentation-generation.md) +- [Frontend Core](../../reference/architecture/frontend-core/frontend-core.md) +- [Utils](../../reference/architecture/utils/utils.md) diff --git a/docs/development/setup/environment.md b/docs/development/setup/environment.md new file mode 100644 index 00000000..1ab4a478 --- /dev/null +++ b/docs/development/setup/environment.md @@ -0,0 +1,58 @@ +# Development Environment Setup + +This guide covers the tools and settings recommended for developing CodeWiki itself. + +## Required Development Tools + +| Tool | Version | Notes | +|---|---|---| +| Python | `>=3.12` | Matches `requires-python` in `pyproject.toml`. | +| pip | Latest | Used to install both runtime and `dev` optional dependencies. | +| Git | Any recent version | CodeWiki's own CLI and web app both shell out to `git` / use GitPython. | +| Node.js | `>=14.0.0` | Required by `mermaid-py`, which validates Mermaid diagrams embedded in generated docs during tests and generation. | +| Docker & Docker Compose | Recent version | Optional, for testing the containerized web app (`docker/docker-compose.yml`, `docker/Dockerfile`). | + +## Installing Development Dependencies + +CodeWiki defines an optional `dev` dependency group in `pyproject.toml`: + +```bash +pip install -e ".[dev]" +``` + +This installs, in addition to the runtime dependencies: + +- `pytest`, `pytest-cov`, `pytest-asyncio` — testing +- `black` — code formatting (line length 100, target `py312`) +- `mypy` — static type checking (`python_version = "3.12"`) +- `ruff` — linting + +## IDE Recommendations + +Any editor with solid Python tooling works well. If using **VS Code**, the following extensions align with the project's tooling: + +- **Python** (Microsoft) — core language support, linting, debugging +- **Pylance** — type-checking assistance (complements `mypy`) +- **Black Formatter** — matches the project's `[tool.black]` configuration (line-length 100, `py312` target) +- **Ruff** — matches the project's linter configuration +- **Mermaid Preview** — useful when reviewing generated architecture diagrams in Markdown output + +If using **PyCharm**, enable Black as the external formatter and configure the line length to 100 to match `[tool.black]` in `pyproject.toml`. + +## Development Environment Variables + +CodeWiki's core CLI configuration lives in `~/.codewiki/config.json` and the OS keyring — it does not require environment variables for normal development use. However, a few areas of the codebase do read environment variables directly: + +| Variable | Used By | Purpose | +|---|---|---| +| `OPENAI_API_KEY` | Ad-hoc clustering test scripts (`test_clustering_*.py`) | API key for OpenAI-compatible providers during manual testing. | +| `ANTHROPIC_API_KEY` | Ad-hoc clustering test scripts | API key for Anthropic providers during manual testing. | +| `MAIN_API_KEY` / `CLUSTER_API_KEY` / `FALLBACK_API_KEY` | Ad-hoc clustering test scripts | Per-role overrides used when running the standalone clustering diagnostics. | +| `PYTHONPATH` | Docker image / `run_web_app.py` path setup | Set to `/app` inside the container; locally, `run_web_app.py` inserts `codewiki/src` onto `sys.path` itself. | +| `APP_PORT` | `docker/docker-compose.yml` | Host port mapping for the containerized web app (defaults to `8000`). | + +> **Tip:** For scripts that read API keys from the environment, consider using a local `.env.local` file with `python-dotenv` (already a project dependency) rather than exporting secrets into your shell history. + +## Next Steps + +Continue to [Local Development](local-development.md) to clone the repository, install it in editable mode, and run the CLI or web app locally. diff --git a/docs/development/setup/local-development.md b/docs/development/setup/local-development.md new file mode 100644 index 00000000..99763148 --- /dev/null +++ b/docs/development/setup/local-development.md @@ -0,0 +1,90 @@ +# Local Development + +This guide walks through cloning CodeWiki, installing it in editable mode, and running it locally against a test repository. + +## Clone and Install + +```bash +git clone https://github.com/flamingo-stack/CodeWiki.git +cd CodeWiki + +# Editable install with development dependencies +pip install -e ".[dev]" +``` + +Editable installs (`-e`) mean changes to `codewiki/` source files take effect immediately without reinstalling — ideal for iterating on the CLI, backend pipeline, or web app. + +Verify the CLI is on your `$PATH` and pointing at your local checkout: + +```bash +codewiki --version +``` + +## Running the CLI Locally + +Once installed, you can run `codewiki` against any repository (including CodeWiki's own repository, or one of the bundled test fixtures under `test-multi-path/`): + +```bash +codewiki config set \ + --cluster-api-key "sk-..." --main-api-key "sk-..." \ + --cluster-model "your-model" --main-model "your-model" \ + --cluster-base-url "https://api.your-provider.com/v1" \ + --main-base-url "https://api.your-provider.com/v1" + +cd /path/to/some/repo +codewiki generate --verbose +``` + +You can also invoke the CLI as a module without relying on the installed console script: + +```bash +python -m codewiki --help +python -m codewiki generate +``` + +## Running the Web Application Locally + +The FastAPI web app can be started directly with Python: + +```bash +python codewiki/run_web_app.py +``` + +This inserts `codewiki/src` onto `sys.path` and delegates to `fe.web_app.main()`. By default it listens on `127.0.0.1:8000` (see `WebAppConfig.DEFAULT_HOST` / `DEFAULT_PORT` in `codewiki/src/fe/config.py`). + +Alternatively, run it in a container using Docker Compose: + +```bash +cd docker +docker compose up --build +``` + +The container maps port `8000` (overridable with the `APP_PORT` environment variable) and mounts `../output` for persistent cache/output storage, plus your `~/.ssh` directory (read-only) for cloning private repositories over SSH. + +## Hot Reload / Iterating on Code + +- **CLI changes**: Because the package is installed with `pip install -e .`, edits to any file under `codewiki/cli/` or `codewiki/src/` are picked up the next time you invoke `codewiki` — no reinstall needed. +- **Web app changes**: `run_web_app.py` does not enable an auto-reloading development server by default. If you need hot reload while iterating on FastAPI routes, restart `python codewiki/run_web_app.py` after each change, or run the underlying ASGI app through `uvicorn` with `--reload` if you invoke it that way directly. + +## Debug Configuration + +For debugging the CLI in an IDE (e.g., VS Code's Python debugger), set the program to run as a module with arguments, for example: + +```bash +python -m codewiki generate --verbose +``` + +Configure your IDE's launch configuration to run `codewiki/cli/main.py` (or `python -m codewiki`) with the working directory set to the repository you want to analyze, and pass CLI arguments like `generate --verbose` for step-by-step output. + +For debugging the web app, set breakpoints inside `codewiki/src/fe/routes.py` or `codewiki/src/fe/background_worker.py` and launch `codewiki/run_web_app.py` directly through your IDE's debugger rather than the CLI. + +## Testing Your Changes Against Sample Fixtures + +The repository includes a self-contained multi-path test fixture at `test-multi-path/` (with `main/`, `deps/`, `external/` subdirectories) specifically designed to exercise multi-root dependency analysis. Use it to sanity-check changes to the dependency analyzer without needing a large external repository: + +```bash +python test-multi-path/test_multi_path.py +python test-multi-path/integration_test.py +``` + +See the [Testing](../testing/README.md) guide for more on how these and the clustering diagnostic scripts are organized. diff --git a/docs/diagrams/architecture/.gitignore b/docs/diagrams/architecture/.gitignore new file mode 100644 index 00000000..b4f4c7a6 --- /dev/null +++ b/docs/diagrams/architecture/.gitignore @@ -0,0 +1,8 @@ +# CodeWiki temp files (dependency graphs can be 7GB+) +temp/ +dependency_graphs/ + +# JSON intermediate files (except schema/config) +*.json +!*-schema.json +!*-config.json diff --git a/docs/diagrams/architecture/README-2.mmd b/docs/diagrams/architecture/README-2.mmd new file mode 100644 index 00000000..765cb17c --- /dev/null +++ b/docs/diagrams/architecture/README-2.mmd @@ -0,0 +1,17 @@ +sequenceDiagram + participant Caller + participant Config as "Config Core" + participant Generator as "DocumentationGenerator" + participant Analyzer as "Dependency Analyzer" + participant Agents as "AgentOrchestrator" + participant Output as "Documentation Output" + + Caller->>Config: Build validated runtime configuration + Caller->>Generator: Start documentation run + Generator->>Analyzer: Analyze repository files and calls + Analyzer-->>Generator: Dependency graph and leaf components + Generator->>Generator: Cluster components into modules + Generator->>Agents: Generate leaf module documentation + Agents-->>Generator: Generated module content + Generator->>Output: Write Markdown, module tree, and metadata + Generator-->>Caller: Generation complete diff --git a/docs/diagrams/architecture/README-3.mmd b/docs/diagrams/architecture/README-3.mmd new file mode 100644 index 00000000..a75c9430 --- /dev/null +++ b/docs/diagrams/architecture/README-3.mmd @@ -0,0 +1,18 @@ +flowchart LR + Config["Config Core"] --> CLI["CLI Core"] + Config --> Frontend["Frontend Core"] + Config --> Backend["Backend Core"] + + CLI --> Backend + Frontend --> Backend + + Backend --> AgentTools["Agent Tools Core"] + Backend --> Analyzer["Dependency Analyzer Core"] + Backend --> Language["Tree-sitter Analyzers"] + Backend --> LLM["LLM Services"] + + TestPaths["Test Multi Path"] --> Config + TestPaths --> Analyzer + + TestCluster["Test Clustering"] --> Config + TestCluster --> Backend diff --git a/docs/diagrams/architecture/README.mmd b/docs/diagrams/architecture/README.mmd new file mode 100644 index 00000000..3b6ad602 --- /dev/null +++ b/docs/diagrams/architecture/README.mmd @@ -0,0 +1,29 @@ +flowchart TD + User["Developer or Web User"] --> Entry{{"Choose entry point?"}} + Entry -->|CLI| CLI["CLI Core"] + Entry -->|Web| Frontend["Frontend Core"] + + CLI --> RuntimeConfig["Config Core"] + Frontend --> RuntimeConfig + + CLI --> GitOps["Git Integration"] + Frontend --> RepoProcessor["GitHub Repository Processor"] + + GitOps --> Source["Source Repository"] + RepoProcessor --> Source + + RuntimeConfig --> Generator["DocumentationGenerator"] + Source --> Generator + + Generator --> Analysis["Dependency Analysis"] + Analysis --> Parsers["Language Parsers"] + Parsers --> Graph["Dependency Graph"] + + Graph --> Clustering["Module Clustering"] + Clustering --> Agents["LLM Agent Orchestration"] + Agents --> Docs["Markdown Documentation"] + + Docs --> Metadata["Module Tree and Metadata"] + Metadata --> HTML["Optional HTML Viewer"] + Docs --> Output["Generated Documentation Output"] + HTML --> Output diff --git a/docs/diagrams/architecture/agent-tools-core-2.mmd b/docs/diagrams/architecture/agent-tools-core-2.mmd new file mode 100644 index 00000000..d8ffbd0e --- /dev/null +++ b/docs/diagrams/architecture/agent-tools-core-2.mmd @@ -0,0 +1,17 @@ +sequenceDiagram + participant Agent + participant Tool as "str_replace_editor" + participant Edit as "EditTool" + participant FS as "Filesystem" + participant Validator as "Mermaid Validator" + + Agent->>Tool: command="create", working_dir="docs", path="module.md" + Tool->>Tool: resolve absolute_path under docs root + Tool->>Edit: EditTool(registry, docs_path) + Edit->>Edit: validate_path(command, path) + Edit->>FS: create parent dirs + write_file + FS-->>Edit: written + Edit-->>Tool: success log + Tool->>Validator: validate_mermaid_diagrams(path) + Validator-->>Tool: validation result + Tool-->>Agent: combined result string diff --git a/docs/diagrams/architecture/agent-tools-core-3.mmd b/docs/diagrams/architecture/agent-tools-core-3.mmd new file mode 100644 index 00000000..3c61cd9f --- /dev/null +++ b/docs/diagrams/architecture/agent-tools-core-3.mmd @@ -0,0 +1,13 @@ +flowchart LR + Edit["Edit .md file"] --> Check{{"path ends with .md?"}} + Check -->|"no"| Done["Return result"] + Check -->|"yes"| Attempts["Read mermaid_attempts counter"] + Attempts --> Limit{{"attempts >= MAX?"}} + Limit -->|"yes"| Skip["Skip validation, warn agent"] + Limit -->|"no"| Validate["validate_mermaid_diagrams"] + Validate --> HasError{{"errors found?"}} + HasError -->|"yes"| Increment["Increment counter in registry"] + HasError -->|"no"| Reset["Reset counter to 0"] + Increment --> Done + Reset --> Done + Skip --> Done diff --git a/docs/diagrams/architecture/agent-tools-core.mmd b/docs/diagrams/architecture/agent-tools-core.mmd new file mode 100644 index 00000000..f7eb85cd --- /dev/null +++ b/docs/diagrams/architecture/agent-tools-core.mmd @@ -0,0 +1,12 @@ +flowchart TD + Orchestrator["AgentOrchestrator"] -->|"builds"| Deps["CodeWikiDeps"] + Orchestrator -->|"invokes tool with"| ToolCall["str_replace_editor Tool Call"] + ToolCall -->|"uses ctx.deps"| Deps + ToolCall -->|"instantiates"| EditTool["EditTool"] + EditTool -->|"large .py view"| Filemap["Filemap"] + EditTool -->|"view/edit windows"| WindowExpander["WindowExpander"] + EditTool -->|"reads"| RepoFS[("Repository Files")] + EditTool -->|"reads/writes"| DocsFS[("Generated Docs Files")] + ToolCall -->|"post-edit .md check"| MermaidValidator["validate_mermaid_diagrams"] + Deps -->|"references"| Config["Config"] + Deps -->|"references"| Node["Node (component registry)"] diff --git a/docs/diagrams/architecture/analysis_pipeline-2.mmd b/docs/diagrams/architecture/analysis_pipeline-2.mmd new file mode 100644 index 00000000..9488e900 --- /dev/null +++ b/docs/diagrams/architecture/analysis_pipeline-2.mmd @@ -0,0 +1,67 @@ +classDiagram + class AnalysisService { + +call_graph_analyzer: CallGraphAnalyzer + +analyze_local_repository(repo_path, max_files, languages) dict + +analyze_repository_full(github_url, include_patterns, exclude_patterns) AnalysisResult + +analyze_repository_structure_only(github_url, include_patterns, exclude_patterns) dict + +cleanup_all() + -_clone_repository(github_url) str + -_analyze_structure(repo_dir, include, exclude) dict + -_analyze_call_graph(file_tree, repo_dir) dict + -_read_readme_file(repo_dir) str + -_cleanup_repository(temp_dir) + } + class RepoAnalyzer { + +include_patterns: list + +exclude_patterns: list + +analyze_repository_structure(repo_dir) dict + -_build_file_tree(repo_dir) dict + -_should_exclude_path(path, filename) bool + -_should_include_file(path, filename) bool + -_count_files(tree) int + -_calculate_size(tree) float + } + class CallGraphAnalyzer { + +functions: dict~str, Node~ + +call_relationships: list~CallRelationship~ + +analyze_code_files(code_files, base_dir) dict + +extract_code_files(file_tree) list + -_analyze_code_file(repo_dir, file_info) + -_resolve_call_relationships() + -_deduplicate_relationships() + -_generate_visualization_data() dict + +generate_llm_format() dict + -_select_most_connected_nodes(target_count) + } + class AnalysisResult { + +repository: Repository + +functions: list + +relationships: list + +file_tree: dict + +summary: dict + +visualization: dict + +readme_content: str + } + class Node { + +id: str + +name: str + +node_type: str + +file_path: str + +component_id: str + +docstring: str + +parameters: list + } + class CallRelationship { + +caller: str + +callee: str + +call_line: int + +is_resolved: bool + } + + AnalysisService --> RepoAnalyzer : uses + AnalysisService --> CallGraphAnalyzer : uses + AnalysisService --> AnalysisResult : produces + CallGraphAnalyzer --> Node : produces + CallGraphAnalyzer --> CallRelationship : produces + AnalysisResult --> Node : contains + AnalysisResult --> CallRelationship : contains diff --git a/docs/diagrams/architecture/analysis_pipeline-3.mmd b/docs/diagrams/architecture/analysis_pipeline-3.mmd new file mode 100644 index 00000000..bb86a740 --- /dev/null +++ b/docs/diagrams/architecture/analysis_pipeline-3.mmd @@ -0,0 +1,27 @@ +sequenceDiagram + participant Caller + participant AS as AnalysisService + participant Clone as "Cloning Utility" + participant RA as RepoAnalyzer + participant CGA as CallGraphAnalyzer + participant FS as Filesystem + + Caller->>AS: analyze_repository_full(github_url) + AS->>Clone: clone_repository(github_url) + Clone-->>AS: temp_dir + AS->>AS: parse_github_url(github_url) + AS->>RA: analyze_repository_structure(temp_dir) + RA->>FS: walk directory tree + RA-->>AS: file_tree + summary + AS->>CGA: extract_code_files(file_tree) + CGA-->>AS: code_files + AS->>CGA: analyze_code_files(code_files, temp_dir) + CGA->>CGA: dispatch per language, parse each file + CGA->>CGA: resolve_call_relationships() + CGA->>CGA: deduplicate_relationships() + CGA->>CGA: generate_visualization_data() + CGA-->>AS: functions, relationships, call_graph, visualization + AS->>FS: read README file + AS->>AS: build AnalysisResult + AS->>Clone: cleanup_repository(temp_dir) + AS-->>Caller: AnalysisResult diff --git a/docs/diagrams/architecture/analysis_pipeline-4.mmd b/docs/diagrams/architecture/analysis_pipeline-4.mmd new file mode 100644 index 00000000..d720e0bf --- /dev/null +++ b/docs/diagrams/architecture/analysis_pipeline-4.mmd @@ -0,0 +1,10 @@ +flowchart LR + subgraph Full["analyze_repository_full"] + F1["Clone"] --> F2["Structure Analysis"] --> F3["Call Graph Analysis"] --> F4["README Read"] --> F5["AnalysisResult"] + end + subgraph StructureOnly["analyze_repository_structure_only"] + S1["Clone"] --> S2["Structure Analysis"] --> S3["Structure Dict"] + end + subgraph Local["analyze_local_repository"] + L1["Local Path"] --> L2["Structure Analysis"] --> L3["Filter by language / max_files"] --> L4["Call Graph Analysis"] --> L5["Simplified Dict"] + end diff --git a/docs/diagrams/architecture/analysis_pipeline-5.mmd b/docs/diagrams/architecture/analysis_pipeline-5.mmd new file mode 100644 index 00000000..c961c238 --- /dev/null +++ b/docs/diagrams/architecture/analysis_pipeline-5.mmd @@ -0,0 +1,11 @@ +flowchart TD + Input["repo_dir: str or list[str]"] --> Check{{"Single path?"}} + Check -->|"Yes"| Single["_build_file_tree(repo_dir)"] + Check -->|"No"| Multi["_analyze_multiple_repositories(repo_dirs)"] + Multi --> Loop["For each repo_dir: compute namespace, build tree"] + Loop --> Wrap["Wrap tree with namespace prefix"] + Wrap --> Merge["Merge into single root tree"] + Single --> Summary1["Compute total_files, total_size_kb"] + Merge --> Summary2["Aggregate totals across repositories"] + Summary1 --> Out["file_tree + summary"] + Summary2 --> Out diff --git a/docs/diagrams/architecture/analysis_pipeline-6.mmd b/docs/diagrams/architecture/analysis_pipeline-6.mmd new file mode 100644 index 00000000..6595ab6f --- /dev/null +++ b/docs/diagrams/architecture/analysis_pipeline-6.mmd @@ -0,0 +1,24 @@ +flowchart TD + FT["file_tree"] --> Extract["extract_code_files()"] + Extract --> CodeFiles["code_files list"] + CodeFiles --> Dispatch{{"Dispatch by language"}} + Dispatch -->|"python"| PyA["Python AST Analyzer"] + Dispatch -->|"javascript"| JsA["JS Tree-Sitter Analyzer"] + Dispatch -->|"typescript"| TsA["TS Tree-Sitter Analyzer"] + Dispatch -->|"java"| JavaA["Java Tree-Sitter Analyzer"] + Dispatch -->|"csharp"| CsA["C# Tree-Sitter Analyzer"] + Dispatch -->|"c"| CA["C Tree-Sitter Analyzer"] + Dispatch -->|"cpp"| CppA["C++ Tree-Sitter Analyzer"] + Dispatch -->|"php"| PhpA["PHP Tree-Sitter Analyzer"] + PyA --> Agg["functions dict, call_relationships list"] + JsA --> Agg + TsA --> Agg + JavaA --> Agg + CsA --> Agg + CA --> Agg + CppA --> Agg + PhpA --> Agg + Agg --> Resolve["_resolve_call_relationships()"] + Resolve --> Dedup["_deduplicate_relationships()"] + Dedup --> Viz["_generate_visualization_data()"] + Viz --> Result["functions, relationships, call_graph, visualization"] diff --git a/docs/diagrams/architecture/analysis_pipeline-7.mmd b/docs/diagrams/architecture/analysis_pipeline-7.mmd new file mode 100644 index 00000000..625b8c6d --- /dev/null +++ b/docs/diagrams/architecture/analysis_pipeline-7.mmd @@ -0,0 +1,9 @@ +flowchart TD + Start["callee: str"] --> Direct{{"callee in func_lookup?"}} + Direct -->|"Yes"| Resolved["Mark resolved, rewrite callee to func_id"] + Direct -->|"No"| HasDot{{"Contains '.'?"}} + HasDot -->|"No"| Unresolved["Leave unresolved"] + HasDot -->|"Yes"| MethodName["Extract trailing segment after last '.'"] + MethodName --> MethodLookup{{"method_name in func_lookup?"}} + MethodLookup -->|"Yes"| Resolved + MethodLookup -->|"No"| Unresolved diff --git a/docs/diagrams/architecture/analysis_pipeline.mmd b/docs/diagrams/architecture/analysis_pipeline.mmd new file mode 100644 index 00000000..3ee233ea --- /dev/null +++ b/docs/diagrams/architecture/analysis_pipeline.mmd @@ -0,0 +1,12 @@ +flowchart TD + Client["Caller (e.g. Documentation Generator)"] --> AS["AnalysisService"] + AS -->|"clone_repository()"| Clone["Repository Cloning (analysis.cloning)"] + AS -->|"analyze_repository_structure()"| RA["RepoAnalyzer"] + AS -->|"analyze_code_files()"| CGA["CallGraphAnalyzer"] + RA -->|"file_tree"| CGA + CGA -->|"extract_code_files()"| Extract["Code File Extraction"] + CGA -->|"dispatch by language"| Analyzers["Language Analyzers"] + Analyzers --> TS["Tree-Sitter Analyzers module"] + CGA -->|"Node / CallRelationship"| Models["Dependency Analyzer Models module"] + AS -->|"AnalysisResult"| Models + AS -->|"Repository"| Models diff --git a/docs/diagrams/architecture/backend-core-2.mmd b/docs/diagrams/architecture/backend-core-2.mmd new file mode 100644 index 00000000..6e055d43 --- /dev/null +++ b/docs/diagrams/architecture/backend-core-2.mmd @@ -0,0 +1,18 @@ +sequenceDiagram + participant Caller + participant Generator as "DocumentationGenerator" + participant Builder as "DependencyGraphBuilder" + participant Agent as "AgentOrchestrator" + participant Tools as "Agent Tools" + participant Output as "Documentation Files" + + Caller->>Generator: run() + Generator->>Builder: build_dependency_graph() + Builder-->>Generator: components and leaf nodes + Generator->>Generator: cluster modules and order leaves first + Generator->>Agent: process leaf module + Agent->>Tools: inspect repository and write docs + Tools-->>Agent: tool results + Agent-->>Generator: module documentation complete + Generator->>Output: write parent overviews and metadata + Generator-->>Caller: documentation complete diff --git a/docs/diagrams/architecture/backend-core-3.mmd b/docs/diagrams/architecture/backend-core-3.mmd new file mode 100644 index 00000000..cdcb1dfd --- /dev/null +++ b/docs/diagrams/architecture/backend-core-3.mmd @@ -0,0 +1,10 @@ +flowchart LR + Models["Dependency Analyzer Models"] --> Analyzer["Dependency Analyzer Core"] + Language["Tree-sitter Analyzers"] --> Analyzer + Logging["Logging Config"] -.-> Analyzer + + Analyzer --> Generator["Documentation Generator"] + LLM["LLM Services"] --> Generator + Generator --> Orchestrator["AgentOrchestrator"] + Orchestrator --> Tools["Agent Tools Core"] + Tools --> Output["Markdown Documentation"] diff --git a/docs/diagrams/architecture/backend-core.mmd b/docs/diagrams/architecture/backend-core.mmd new file mode 100644 index 00000000..41fed11d --- /dev/null +++ b/docs/diagrams/architecture/backend-core.mmd @@ -0,0 +1,20 @@ +flowchart TD + Input["Source Repository"] --> Generator["DocumentationGenerator"] + Generator --> GraphBuilder["DependencyGraphBuilder"] + GraphBuilder --> Parser["DependencyParser"] + Parser --> Analysis["AnalysisService"] + Analysis --> RepoAnalyzer["RepoAnalyzer"] + Analysis --> CallAnalyzer["CallGraphAnalyzer"] + CallAnalyzer --> LanguageAnalyzers["Tree-sitter and Python Analyzers"] + LanguageAnalyzers --> Models["Dependency Analyzer Models"] + + GraphBuilder --> Components["Component Dependency Graph"] + Components --> Generator + + Generator --> Orchestrator["AgentOrchestrator"] + Orchestrator --> AgentTools["Agent Tools Core"] + Orchestrator --> LLM["CountingFallbackModel"] + LLM --> Provider["Configured LLM Provider"] + + AgentTools --> Docs["Generated Markdown Documentation"] + Generator --> Docs diff --git a/docs/diagrams/architecture/c_family_analyzers-2.mmd b/docs/diagrams/architecture/c_family_analyzers-2.mmd new file mode 100644 index 00000000..2e0228b8 --- /dev/null +++ b/docs/diagrams/architecture/c_family_analyzers-2.mmd @@ -0,0 +1,17 @@ +sequenceDiagram + participant Caller as "analyze_cpp_file()" + participant Analyzer as "TreeSitterCppAnalyzer" + participant TS as "tree-sitter Parser" + participant Pass1 as "_extract_nodes()" + participant Pass2 as "_extract_relationships()" + + Caller->>Analyzer: "__init__(file_path, content, repo_path)" + Analyzer->>TS: "parse(content)" + TS-->>Analyzer: "syntax tree" + Analyzer->>Pass1: "traverse(root, top_level_nodes, lines)" + Pass1->>Pass1: "detect class/struct/function/namespace/variable" + Pass1-->>Analyzer: "populated top_level_nodes + nodes list" + Analyzer->>Pass2: "traverse(root, top_level_nodes)" + Pass2->>Pass2: "detect calls, inheritance, instantiation, usage" + Pass2-->>Analyzer: "call_relationships list" + Analyzer-->>Caller: "nodes, call_relationships" diff --git a/docs/diagrams/architecture/c_family_analyzers-3.mmd b/docs/diagrams/architecture/c_family_analyzers-3.mmd new file mode 100644 index 00000000..9d42be91 --- /dev/null +++ b/docs/diagrams/architecture/c_family_analyzers-3.mmd @@ -0,0 +1,50 @@ +classDiagram + class TreeSitterCAnalyzer { + +file_path + +content + +repo_path + +nodes + +call_relationships + -_analyze() + -_extract_nodes() + -_extract_relationships() + } + class TreeSitterCppAnalyzer { + +file_path + +content + +repo_path + +nodes + +call_relationships + -_analyze() + -_extract_nodes() + -_extract_relationships() + } + class TreeSitterCSharpAnalyzer { + +file_path + +content + +repo_path + +nodes + +call_relationships + -_analyze() + -_extract_nodes() + -_extract_relationships() + } + class Node { + +id + +name + +component_type + +file_path + +source_code + } + class CallRelationship { + +caller + +callee + +call_line + +is_resolved + } + TreeSitterCAnalyzer --> Node : produces + TreeSitterCAnalyzer --> CallRelationship : produces + TreeSitterCppAnalyzer --> Node : produces + TreeSitterCppAnalyzer --> CallRelationship : produces + TreeSitterCSharpAnalyzer --> Node : produces + TreeSitterCSharpAnalyzer --> CallRelationship : produces diff --git a/docs/diagrams/architecture/c_family_analyzers-4.mmd b/docs/diagrams/architecture/c_family_analyzers-4.mmd new file mode 100644 index 00000000..843a0ffd --- /dev/null +++ b/docs/diagrams/architecture/c_family_analyzers-4.mmd @@ -0,0 +1,13 @@ +flowchart LR + Repo["Repository Source Files"] --> Parser["Dependency Parser"] + Parser -->|".c / .h"| CAnalyzer["TreeSitterCAnalyzer"] + Parser -->|".cpp / .hpp"| CppAnalyzer["TreeSitterCppAnalyzer"] + Parser -->|".cs"| CSharpAnalyzer["TreeSitterCSharpAnalyzer"] + CAnalyzer --> RawNodes["Raw Nodes + Relationships"] + CppAnalyzer --> RawNodes + CSharpAnalyzer --> RawNodes + RawNodes --> GraphBuilder["Dependency Graph Builder"] + RawNodes --> CallGraph["Call Graph Analyzer"] + CallGraph --> Resolved["Resolved Cross-File Relationships"] + GraphBuilder --> FinalGraph["Repository Dependency Graph"] + Resolved --> FinalGraph diff --git a/docs/diagrams/architecture/c_family_analyzers.mmd b/docs/diagrams/architecture/c_family_analyzers.mmd new file mode 100644 index 00000000..e144dcff --- /dev/null +++ b/docs/diagrams/architecture/c_family_analyzers.mmd @@ -0,0 +1,8 @@ +flowchart TD + Source["Source File Content"] --> Parser["tree-sitter Parser"] + Parser --> Tree["Concrete Syntax Tree"] + Tree --> Extract["Node Extraction Pass"] + Extract --> TopLevel["Top-Level Node Registry"] + TopLevel --> Rel["Relationship Extraction Pass"] + Rel --> Nodes["List of Node objects"] + Rel --> Calls["List of CallRelationship objects"] diff --git a/docs/diagrams/architecture/cli-core.mmd b/docs/diagrams/architecture/cli-core.mmd new file mode 100644 index 00000000..9eda5120 --- /dev/null +++ b/docs/diagrams/architecture/cli-core.mmd @@ -0,0 +1,27 @@ +flowchart TD + User["CLI User"] -->|"codewiki generate"| ConfigMgr["ConfigManager"] + ConfigMgr -->|"loads/saves"| ConfigFile[("~/.codewiki/config.json")] + ConfigMgr -->|"stores API keys"| Keyring[("OS Keyring")] + ConfigMgr -->|"provides"| Config["Configuration + AgentInstructions"] + + Config -->|"to_backend_config()"| CLIDocGen["CLIDocumentationGenerator"] + GitMgr["GitManager"] -->|"branch/commit info"| CLIDocGen + + CLIDocGen -->|"drives"| BackendGen["Backend DocumentationGenerator"] + CLIDocGen -->|"tracks progress via"| Progress["ProgressTracker"] + CLIDocGen -->|"builds"| Job["DocumentationJob"] + CLIDocGen -->|"optional"| HTMLGen["HTMLGenerator"] + + HTMLGen -->|"reads"| ModuleTree[("module_tree.json")] + HTMLGen -->|"reads"| Metadata[("metadata.json")] + HTMLGen -->|"writes"| IndexHTML[("index.html")] + + Job -->|"serializes to"| Metadata + + Logger["CLILogger"] -.->|"used by"| CLIDocGen + Logger -.->|"used by"| ConfigMgr + Logger -.->|"used by"| GitMgr + + BackendGen -->|"belongs to"| BackendCore["Backend Core module"] + + style BackendCore fill:#eee,stroke:#999,stroke-dasharray: 5 5 diff --git a/docs/diagrams/architecture/config-core-2.mmd b/docs/diagrams/architecture/config-core-2.mmd new file mode 100644 index 00000000..5a7d4885 --- /dev/null +++ b/docs/diagrams/architecture/config-core-2.mmd @@ -0,0 +1,9 @@ +flowchart TD + A["Config instance"] --> B["to_dict(include_secrets=False)"] + B --> C["Plain dict, no API keys"] + A --> D["to_dict(include_secrets=True)"] + D --> E["Plain dict, includes API keys"] + C --> F["from_dict(data)"] + E --> F + F -->|"missing secrets"| G["TypeError: missing required field"] + F -->|"secrets present"| H["Reconstructed Config"] diff --git a/docs/diagrams/architecture/config-core-3.mmd b/docs/diagrams/architecture/config-core-3.mmd new file mode 100644 index 00000000..b1b39a47 --- /dev/null +++ b/docs/diagrams/architecture/config-core-3.mmd @@ -0,0 +1,10 @@ +flowchart TD + AI["agent_instructions dict"] --> IP["include_patterns"] + AI --> EP["exclude_patterns"] + AI --> FM["focus_modules"] + AI --> DT["doc_type"] + AI --> CI["custom_instructions"] + DT --> GPA["get_prompt_addition()"] + FM --> GPA + CI -->|"escape_format_braces"| GPA + GPA --> Prompt["Combined prompt-addition string"] diff --git a/docs/diagrams/architecture/config-core-4.mmd b/docs/diagrams/architecture/config-core-4.mmd new file mode 100644 index 00000000..5cf96da4 --- /dev/null +++ b/docs/diagrams/architecture/config-core-4.mmd @@ -0,0 +1,12 @@ +flowchart TD + Start["Config.validate_source_paths()"] --> CheckPrimary{{"repo_path exists and is dir?"}} + CheckPrimary -->|"no"| Err1["raise ValueError"] + CheckPrimary -->|"yes"| HasAdditional{{"additional_source_paths set?"}} + HasAdditional -->|"no"| Done["Validation OK single-path mode"] + HasAdditional -->|"yes"| Loop["For each additional path"] + Loop --> CheckExists{{"path exists and is dir?"}} + CheckExists -->|"no"| Err2["raise ValueError"] + CheckExists -->|"yes"| CheckRead{{"path readable?"}} + CheckRead -->|"no"| Err3["raise OSError"] + CheckRead -->|"yes"| Loop + Loop --> Done2["Validation OK multi-path mode"] diff --git a/docs/diagrams/architecture/config-core-5.mmd b/docs/diagrams/architecture/config-core-5.mmd new file mode 100644 index 00000000..584684bd --- /dev/null +++ b/docs/diagrams/architecture/config-core-5.mmd @@ -0,0 +1,21 @@ +flowchart TD + subgraph CLIFlow["CLI Core entry point"] + CM["ConfigManager: persisted JSON plus keyring"] + CM -->|"from_config_manager"| FC1["Config.from_cli(...)"] + end + subgraph WebFlow["Frontend Core entry point"] + BW["BackgroundWorker for web job"] + BW -->|"from_web_job"| FA1["Config.from_args wrapping Namespace"] + end + subgraph EnvFlow["Environment-driven CLI entry point"] + ArgParse["argparse.Namespace from CLI arguments"] + ArgParse -->|"from_args"| FA2["Reads MAIN_MODEL, FALLBACK_MODEL, CLUSTER_API_KEY, MAIN_API_KEY, FALLBACK_API_KEY from environment"] + end + subgraph DirectFlow["Direct parameter entry point"] + Caller["Any caller with explicit parameters"] + Caller -->|"from_cli"| Validate["Validation block: keys, urls, types, ranges, max_token_field enum"] + Validate --> VSP["validate_source_paths()"] + VSP --> Instance["Config instance"] + end + FA1 --> FA2 + FC1 --> Validate diff --git a/docs/diagrams/architecture/config-core-6.mmd b/docs/diagrams/architecture/config-core-6.mmd new file mode 100644 index 00000000..76efd71d --- /dev/null +++ b/docs/diagrams/architecture/config-core-6.mmd @@ -0,0 +1,16 @@ +sequenceDiagram + participant CLI as "CLI Core (ConfigManager)" + participant Config as "Config.from_config_manager" + participant FromCli as "Config.from_cli" + participant Validate as "validate_source_paths" + + CLI->>Config: from_config_manager(manager, repo_path, output_dir) + Config->>CLI: manager.get_config() + Config->>CLI: get_cluster_api_key / get_main_api_key / get_fallback_api_key + Config->>Config: check models and keys are present + Config->>FromCli: from_cli(repo_path, output_dir, models, keys, urls, tokens, temps) + FromCli->>FromCli: validation block keys urls types ranges enums + FromCli->>Validate: validate_source_paths() + Validate-->>FromCli: OK or raises ValueError or OSError + FromCli-->>Config: Config instance + Config-->>CLI: Config instance diff --git a/docs/diagrams/architecture/config-core.mmd b/docs/diagrams/architecture/config-core.mmd new file mode 100644 index 00000000..8a87bba8 --- /dev/null +++ b/docs/diagrams/architecture/config-core.mmd @@ -0,0 +1,43 @@ +classDiagram + class Config { + +str repo_path + +str output_dir + +str dependency_graph_dir + +str docs_dir + +int max_depth + +str main_model + +str cluster_model + +str fallback_model + +str cluster_api_key + +str main_api_key + +str fallback_api_key + +Optional~str~ cluster_base_url + +Optional~str~ main_base_url + +Optional~str~ fallback_base_url + +int cluster_max_tokens + +int main_max_tokens + +int fallback_max_tokens + +int max_token_per_module + +int max_token_per_leaf_module + +float cluster_temperature + +float main_temperature + +float fallback_temperature + +Optional~Dict~ agent_instructions + +Optional~str~ diagrams_dir + +Optional~List~str~~ additional_source_paths + +to_dict(include_secrets) Dict + +from_dict(data)$ Config + +include_patterns() Optional~List~ + +exclude_patterns() Optional~List~ + +focus_modules() Optional~List~ + +doc_type() Optional~str~ + +custom_instructions() Optional~str~ + +all_source_paths() List~str~ + +validate_source_paths() void + +is_multi_path_mode() bool + +get_prompt_addition() str + +from_args(args)$ Config + +from_web_job(repo_path, docs_dir)$ Config + +from_cli(...)$ Config + +from_config_manager(manager, repo_path, output_dir)$ Config + } diff --git a/docs/diagrams/architecture/configuration-2.mmd b/docs/diagrams/architecture/configuration-2.mmd new file mode 100644 index 00000000..07b6b7d6 --- /dev/null +++ b/docs/diagrams/architecture/configuration-2.mmd @@ -0,0 +1,42 @@ +classDiagram + class Configuration { + +str main_model + +str cluster_model + +str fallback_model + +str default_output + +str cluster_base_url + +str main_base_url + +str fallback_base_url + +int cluster_max_tokens + +int main_max_tokens + +int fallback_max_tokens + +float cluster_temperature + +float main_temperature + +float fallback_temperature + +bool cluster_temperature_supported + +bool main_temperature_supported + +bool fallback_temperature_supported + +int max_token_per_module + +int max_token_per_leaf_module + +int max_depth + +AgentInstructions agent_instructions + +validate() + +to_dict() dict + +from_dict(data) Configuration + +is_complete() bool + +to_backend_config(...) Config + } + + class AgentInstructions { + +List~str~ include_patterns + +List~str~ exclude_patterns + +List~str~ focus_modules + +str doc_type + +str custom_instructions + +to_dict() dict + +from_dict(data) AgentInstructions + +is_empty() bool + +get_prompt_addition() str + } + + Configuration "1" *-- "1" AgentInstructions : agent_instructions diff --git a/docs/diagrams/architecture/configuration-3.mmd b/docs/diagrams/architecture/configuration-3.mmd new file mode 100644 index 00000000..8b326c48 --- /dev/null +++ b/docs/diagrams/architecture/configuration-3.mmd @@ -0,0 +1,20 @@ +classDiagram + class ConfigManager { + -Optional~str~ _cluster_api_key + -Optional~str~ _main_api_key + -Optional~str~ _fallback_api_key + -Optional~Configuration~ _config + -bool _keyring_available + +load() bool + +save(...) void + +get_cluster_api_key() Optional~str~ + +get_main_api_key() Optional~str~ + +get_fallback_api_key() Optional~str~ + +get_config() Optional~Configuration~ + +is_configured() bool + +delete_api_keys() void + +clear() void + +keyring_available bool + +config_file_path Path + } + ConfigManager --> Configuration : loads/saves diff --git a/docs/diagrams/architecture/configuration-4.mmd b/docs/diagrams/architecture/configuration-4.mmd new file mode 100644 index 00000000..7ecb115e --- /dev/null +++ b/docs/diagrams/architecture/configuration-4.mmd @@ -0,0 +1,20 @@ +sequenceDiagram + participant Caller + participant CM as ConfigManager + participant FS as "config.json" + participant KR as "System Keyring" + + Caller->>CM: load() + CM->>FS: check exists / read + alt file missing + FS-->>CM: not found + CM-->>Caller: False + else file present + FS-->>CM: JSON content + CM->>CM: Configuration.from_dict(data) + CM->>KR: get_password(cluster_api_key) + CM->>KR: get_password(main_api_key) + CM->>KR: get_password(fallback_api_key) + KR-->>CM: key values (or None) + CM-->>Caller: True + end diff --git a/docs/diagrams/architecture/configuration-5.mmd b/docs/diagrams/architecture/configuration-5.mmd new file mode 100644 index 00000000..14c17fb6 --- /dev/null +++ b/docs/diagrams/architecture/configuration-5.mmd @@ -0,0 +1,16 @@ +sequenceDiagram + participant Caller + participant CM as ConfigManager + participant FS as "config.json" + participant KR as "System Keyring" + + Caller->>CM: save(fields..., api_keys...) + CM->>FS: ensure_directory(CONFIG_DIR) + alt no in-memory config + CM->>CM: load() existing or create default Configuration + end + CM->>CM: apply provided field updates + CM->>CM: Configuration.validate() (if models set) + CM->>KR: set_password(...) for each provided API key + CM->>FS: write JSON (version + Configuration.to_dict()) + CM-->>Caller: done (or raises ConfigurationError) diff --git a/docs/diagrams/architecture/configuration-6.mmd b/docs/diagrams/architecture/configuration-6.mmd new file mode 100644 index 00000000..05559834 --- /dev/null +++ b/docs/diagrams/architecture/configuration-6.mmd @@ -0,0 +1,8 @@ +flowchart LR + A["CLI Command"] --> B["ConfigManager.load()"] + B --> C["Configuration"] + C --> D["Configuration.to_backend_config()"] + D --> E["ConfigManager (keyring lookup for missing keys)"] + D --> F["Merge AgentInstructions (runtime over persisted)"] + D --> G["Config.from_cli(...)"] + G --> H["Backend Config"] diff --git a/docs/diagrams/architecture/configuration.mmd b/docs/diagrams/architecture/configuration.mmd new file mode 100644 index 00000000..aa110208 --- /dev/null +++ b/docs/diagrams/architecture/configuration.mmd @@ -0,0 +1,16 @@ +flowchart TD + subgraph ConfigModule["Configuration Module"] + CM["ConfigManager"] + Cfg["Configuration"] + AI["AgentInstructions"] + end + + FS["~/.codewiki/config.json"] + KR["System Keyring"] + Backend["Backend Config"] + + CM -->|"load() / save()"| FS + CM -->|"get/set API keys"| KR + CM -->|"holds"| Cfg + Cfg -->|"has one"| AI + Cfg -->|"to_backend_config()"| Backend diff --git a/docs/diagrams/architecture/dependency-analyzer-core-2.mmd b/docs/diagrams/architecture/dependency-analyzer-core-2.mmd new file mode 100644 index 00000000..cc2c308c --- /dev/null +++ b/docs/diagrams/architecture/dependency-analyzer-core-2.mmd @@ -0,0 +1,22 @@ +sequenceDiagram + participant Builder as "DependencyGraphBuilder" + participant Parser as "DependencyParser" + participant Service as "AnalysisService" + participant RepoA as "RepoAnalyzer" + participant CallA as "CallGraphAnalyzer" + participant TS as "Tree-sitter Analyzers" + + Builder->>Parser: parse_repository() + Parser->>Service: _analyze_structure(repo_path) + Service->>RepoA: analyze_repository_structure() + RepoA-->>Service: file_tree + summary + Parser->>Service: _analyze_call_graph(file_tree, repo_path) + Service->>CallA: extract_code_files() / analyze_code_files() + CallA->>TS: analyze_python_file / analyze_javascript_file_treesitter / ... + TS-->>CallA: functions, relationships + CallA-->>Service: functions, relationships, visualization + Service-->>Parser: call_graph_result + Parser->>Parser: build namespaced Node components + Parser-->>Builder: components (Dict[str, Node]) + Builder->>Builder: build_graph_from_components / validate / filter leaves + Builder-->>Builder: (components, leaf_nodes) diff --git a/docs/diagrams/architecture/dependency-analyzer-core.mmd b/docs/diagrams/architecture/dependency-analyzer-core.mmd new file mode 100644 index 00000000..706eba5f --- /dev/null +++ b/docs/diagrams/architecture/dependency-analyzer-core.mmd @@ -0,0 +1,21 @@ +flowchart TD + subgraph AnalysisPipeline["Analysis Pipeline"] + direction TB + RepoAnalyzer["RepoAnalyzer"] -->|"file_tree"| AnalysisService["AnalysisService"] + CallGraphAnalyzer["CallGraphAnalyzer"] -->|"functions + relationships"| AnalysisService + end + + subgraph GraphConstruction["Graph Construction"] + direction TB + DependencyParser["DependencyParser"] -->|"components (Node)"| DependencyGraphBuilder["DependencyGraphBuilder"] + end + + Config["Config"] -.->|"repo paths, patterns"| DependencyGraphBuilder + DependencyGraphBuilder -->|"drives"| DependencyParser + DependencyParser -->|"uses"| AnalysisService + CallGraphAnalyzer -->|"delegates per-language"| TreeSitterAnalyzers["Tree-sitter Analyzers"] + AnalysisService -->|"produces"| Models["Node / CallRelationship / AnalysisResult"] + DependencyGraphBuilder -->|"writes"| GraphFile[("dependency_graph.json")] + + click TreeSitterAnalyzers "../tree-sitter-analyzers/tree-sitter-analyzers.md" + click Models "../dependency-analyzer-models/dependency-analyzer-models.md" diff --git a/docs/diagrams/architecture/dependency-analyzer-models-2.mmd b/docs/diagrams/architecture/dependency-analyzer-models-2.mmd new file mode 100644 index 00000000..2cff6eb2 --- /dev/null +++ b/docs/diagrams/architecture/dependency-analyzer-models-2.mmd @@ -0,0 +1,60 @@ +classDiagram + class Node { + +str id + +str name + +str component_type + +str file_path + +str relative_path + +Set~str~ depends_on + +Optional~str~ source_code + +int start_line + +int end_line + +bool has_docstring + +str docstring + +Optional~List~str~~ parameters + +Optional~str~ node_type + +Optional~List~str~~ base_classes + +Optional~str~ class_name + +Optional~str~ display_name + +Optional~str~ component_id + +str short_id + +str namespace + +bool is_from_deps + +get_display_name() str + } + + class CallRelationship { + +str caller + +str callee + +Optional~int~ call_line + +bool is_resolved + } + + class Repository { + +str url + +str name + +str clone_path + +str analysis_id + } + + class AnalysisResult { + +Repository repository + +List~Node~ functions + +List~CallRelationship~ relationships + +Dict file_tree + +Dict summary + +Dict visualization + +Optional~str~ readme_content + } + + class NodeSelection { + +List~str~ selected_nodes + +bool include_relationships + +Dict~str,str~ custom_names + } + + AnalysisResult "1" --> "1" Repository : describes + AnalysisResult "1" --> "*" Node : contains + AnalysisResult "1" --> "*" CallRelationship : contains + CallRelationship "*" --> "1" Node : caller/callee reference (by id) + NodeSelection "1" --> "*" Node : references (by id) diff --git a/docs/diagrams/architecture/dependency-analyzer-models.mmd b/docs/diagrams/architecture/dependency-analyzer-models.mmd new file mode 100644 index 00000000..2aeb4b4e --- /dev/null +++ b/docs/diagrams/architecture/dependency-analyzer-models.mmd @@ -0,0 +1,14 @@ +flowchart TD + Analyzers["Tree-sitter Analyzers"] -->|"produce"| Node["Node"] + Analyzers -->|"produce"| CallRel["CallRelationship"] + Parser["Dependency Parser / Graph Builder"] -->|"assembles"| Node + Parser -->|"assembles"| CallRel + RepoModel["Repository"] -->|"describes source of"| Node + AnalysisSvc["Analysis Service / Repo Analyzer / Call Graph Analyzer"] -->|"aggregates"| Node + AnalysisSvc -->|"aggregates"| CallRel + AnalysisSvc -->|"aggregates"| RepoModel + AnalysisSvc -->|"produces"| Result["AnalysisResult"] + Selection["NodeSelection"] -->|"filters"| Result + Result -->|"consumed by"| DocGen["Documentation Generator"] + Result -->|"consumed by"| Agent["Agent Orchestrator"] + Result -->|"consumed by"| CLIGen["CLI Documentation Generator"] diff --git a/docs/diagrams/architecture/documentation-generator-2.mmd b/docs/diagrams/architecture/documentation-generator-2.mmd new file mode 100644 index 00000000..6fefa10b --- /dev/null +++ b/docs/diagrams/architecture/documentation-generator-2.mmd @@ -0,0 +1,43 @@ +sequenceDiagram + participant Caller as "CLI/Web Caller" + participant DG as "DocumentationGenerator" + participant GB as "DependencyGraphBuilder" + participant Cluster as "cluster_modules" + participant AO as "AgentOrchestrator" + participant LLM as "call_llm" + participant FS as "file_manager" + + Caller->>DG: run() + DG->>GB: build_dependency_graph() + GB-->>DG: components, leaf_nodes + alt "module tree cache exists" + DG->>FS: load_json(first_module_tree.json) + else "no cache" + DG->>Cluster: cluster_modules(leaf_nodes, components, config) + Cluster-->>DG: module_tree + DG->>FS: save_json(first_module_tree.json) + end + alt "module_tree empty but leaf_nodes exist" + DG->>DG: build synthetic modules by top-level directory + DG->>FS: save_json(synthetic module_tree) + end + DG->>FS: save_json(module_tree.json) + DG->>DG: generate_module_documentation(components, leaf_nodes) + loop "for each module in leaf-first order" + alt "is_leaf_module" + DG->>AO: process_module(name, components, ids, path, dir) + AO->>LLM: agent.run(user_prompt) + LLM-->>AO: markdown docs + AO->>FS: save module_name.md + else "parent module" + DG->>DG: build_overview_structure(...) + DG->>LLM: call_llm(MODULE_OVERVIEW_PROMPT) + LLM-->>DG: "..." + DG->>FS: save parent module_name.md + end + end + DG->>DG: generate_parent_module_docs([], working_dir) + DG->>LLM: call_llm(REPO_OVERVIEW_PROMPT) + DG->>FS: save overview.md + DG->>DG: create_documentation_metadata(working_dir, components, len(leaf_nodes)) + DG-->>Caller: "documentation complete" diff --git a/docs/diagrams/architecture/documentation-generator-3.mmd b/docs/diagrams/architecture/documentation-generator-3.mmd new file mode 100644 index 00000000..64ae9844 --- /dev/null +++ b/docs/diagrams/architecture/documentation-generator-3.mmd @@ -0,0 +1,13 @@ +flowchart LR + subgraph Tree["Module Tree Example"] + Root["Backend"] --> Auth["Authentication"] + Root --> API["API Layer"] + Auth --> JWT["JWT"] + Auth --> OAuth["OAuth"] + end + subgraph Order["Processing Order"] + O1["1: JWT (leaf)"] --> O2["2: OAuth (leaf)"] + O2 --> O3["3: Authentication (parent)"] + O3 --> O4["4: API Layer (leaf/parent)"] + O4 --> O5["5: Backend (parent)"] + end diff --git a/docs/diagrams/architecture/documentation-generator.mmd b/docs/diagrams/architecture/documentation-generator.mmd new file mode 100644 index 00000000..de7a1fdb --- /dev/null +++ b/docs/diagrams/architecture/documentation-generator.mmd @@ -0,0 +1,11 @@ +flowchart TD + CLI["CLI Core: CLIDocumentationGenerator"] --> DG["Documentation Generator: DocumentationGenerator"] + Web["Frontend Core: BackgroundWorker"] --> DG + DG --> GraphBuilder["Dependency Analyzer Core: DependencyGraphBuilder"] + DG --> Cluster["cluster_modules"] + DG --> Orchestrator["Backend Core: AgentOrchestrator"] + DG --> LLM["LLM Services: call_llm / CountingFallbackModel"] + GraphBuilder --> Analyzers["Tree-sitter Analyzers"] + GraphBuilder --> Models["Dependency Analyzer Models"] + Orchestrator --> AgentTools["Agent Tools Core"] + DG --> ConfigCore["Config Core: Config"] diff --git a/docs/diagrams/architecture/frontend-core-2.mmd b/docs/diagrams/architecture/frontend-core-2.mmd new file mode 100644 index 00000000..9aba09d9 --- /dev/null +++ b/docs/diagrams/architecture/frontend-core-2.mmd @@ -0,0 +1,24 @@ +sequenceDiagram + participant Browser + participant Routes as "WebRoutes" + participant Worker as "BackgroundWorker" + participant Cache as "CacheManager" + participant Git as "GitHubRepoProcessor" + participant DocGen as "DocumentationGenerator" + + Browser->>Routes: POST / (repo_url, commit_id) + Routes->>Git: is_valid_github_url / get_repo_info + Routes->>Cache: get_cached_docs(repo_url) + alt Cache hit + Cache-->>Routes: docs_path + Routes-->>Browser: Render success message + else Cache miss + Routes->>Worker: add_job(job_id, JobStatus) + Worker-->>Routes: queued + Routes-->>Browser: Render "queued" message + Worker->>Git: clone_repository(clone_url, temp_dir, commit_id) + Worker->>DocGen: run() (async documentation generation) + DocGen-->>Worker: docs_dir populated + Worker->>Cache: add_to_cache(repo_url, docs_path) + Worker->>Worker: save_job_statuses() + end diff --git a/docs/diagrams/architecture/frontend-core-3.mmd b/docs/diagrams/architecture/frontend-core-3.mmd new file mode 100644 index 00000000..e4ac9525 --- /dev/null +++ b/docs/diagrams/architecture/frontend-core-3.mmd @@ -0,0 +1,51 @@ +classDiagram + class WebAppConfig { + +CACHE_DIR + +TEMP_DIR + +QUEUE_SIZE + +CACHE_EXPIRY_DAYS + +CLONE_TIMEOUT + +ensure_directories() + } + class GitHubRepoProcessor { + +is_valid_github_url(url) + +get_repo_info(url) + +clone_repository(clone_url, target_dir, commit_id) + } + class CacheManager { + +cache_index + +get_cached_docs(repo_url) + +add_to_cache(repo_url, docs_path) + +remove_from_cache(repo_url) + } + class BackgroundWorker { + +job_status + +processing_queue + +start() + +add_job(job_id, job) + +get_job_status(job_id) + } + class WebRoutes { + +index_get(request) + +index_post(request, repo_url, commit_id) + +get_job_status(job_id) + +serve_generated_docs(job_id, filename) + } + class JobStatus + class CacheEntry + class RepositorySubmission + class JobStatusResponse + class StringTemplateLoader + + WebRoutes --> BackgroundWorker + WebRoutes --> CacheManager + WebRoutes --> GitHubRepoProcessor + WebRoutes --> StringTemplateLoader + BackgroundWorker --> CacheManager + BackgroundWorker --> GitHubRepoProcessor + BackgroundWorker --> JobStatus + CacheManager --> CacheEntry + WebRoutes --> JobStatusResponse + GitHubRepoProcessor --> WebAppConfig + CacheManager --> WebAppConfig + BackgroundWorker --> WebAppConfig diff --git a/docs/diagrams/architecture/frontend-core-4.mmd b/docs/diagrams/architecture/frontend-core-4.mmd new file mode 100644 index 00000000..3be8bfef --- /dev/null +++ b/docs/diagrams/architecture/frontend-core-4.mmd @@ -0,0 +1,8 @@ +stateDiagram-v2 + [*] --> queued: "add_job()" + queued --> processing: "worker picks up job" + processing --> completed: "cache hit OR generation succeeds" + processing --> failed: "clone or generation error" + failed --> queued: "resubmission after cooldown" + completed --> [*] + failed --> [*] diff --git a/docs/diagrams/architecture/frontend-core.mmd b/docs/diagrams/architecture/frontend-core.mmd new file mode 100644 index 00000000..e36980d6 --- /dev/null +++ b/docs/diagrams/architecture/frontend-core.mmd @@ -0,0 +1,25 @@ +flowchart TD + User["Browser / API Client"] -->|"submit repo_url"| Routes["WebRoutes"] + Routes -->|"validate URL"| GitProc["GitHubRepoProcessor"] + Routes -->|"check cache"| Cache["CacheManager"] + Routes -->|"enqueue job"| Worker["BackgroundWorker"] + Routes -->|"render HTML"| Templates["StringTemplateLoader / render_template"] + + Worker -->|"clone repository"| GitProc + Worker -->|"build Config.from_web_job"| ConfigCore["Config (config-core)"] + Worker -->|"generate docs"| DocGen["DocumentationGenerator (backend-core)"] + Worker -->|"store result path"| Cache + Worker -->|"persist status"| JobsFile[("jobs.json")] + + Cache -->|"persist index"| CacheFile[("cache_index.json")] + + subgraph models_group["Data Models"] + JobStatus["JobStatus"] + CacheEntry["CacheEntry"] + RepositorySubmission["RepositorySubmission"] + JobStatusResponse["JobStatusResponse"] + end + + Routes --> models_group + Worker --> models_group + Cache --> models_group diff --git a/docs/diagrams/architecture/generation-2.mmd b/docs/diagrams/architecture/generation-2.mmd new file mode 100644 index 00000000..670f2390 --- /dev/null +++ b/docs/diagrams/architecture/generation-2.mmd @@ -0,0 +1,26 @@ +sequenceDiagram + participant Caller as "CLI Command" + participant CDG as "CLIDocumentationGenerator" + participant BC as "BackendConfig" + participant DG as "DocumentationGenerator" + participant CM as "cluster_modules" + participant HG as "HTMLGenerator" + participant Job as "DocumentationJob" + + Caller->>CDG: generate() + CDG->>Job: start() + CDG->>BC: Config.from_cli(...) + CDG->>DG: _run_backend_generation(backend_config) + DG->>DG: graph_builder.build_dependency_graph() + DG->>CM: cluster_modules(leaf_nodes, components, config) + Note over DG,CM: Synthetic module fallback if tree is empty + DG->>DG: generate_module_documentation(components, leaf_nodes) + opt diagrams_dir configured + DG->>DG: extract_and_save_mermaid_diagrams() + end + opt generate_html is true + CDG->>HG: generate(output_path, ...) + end + CDG->>CDG: _finalize_job() + CDG->>Job: complete() + CDG-->>Caller: DocumentationJob diff --git a/docs/diagrams/architecture/generation-3.mmd b/docs/diagrams/architecture/generation-3.mmd new file mode 100644 index 00000000..28f2c2c3 --- /dev/null +++ b/docs/diagrams/architecture/generation-3.mmd @@ -0,0 +1,14 @@ +flowchart LR + subgraph CLIConfig["CLI config dict"] + A1["main_model / cluster_model / fallback_model"] + A2["*_api_key"] + A3["*_base_url"] + A4["*_api_version"] + A5["*_max_tokens / *_temperature"] + A6["max_token_per_module / max_depth"] + A7["agent_instructions"] + A8["additional_paths"] + end + CLIConfig --> FromCLI["Config.from_cli()"] + FromCLI --> BackendConfig["Backend Config instance"] + BackendConfig --> DG2["DocumentationGenerator"] diff --git a/docs/diagrams/architecture/generation.mmd b/docs/diagrams/architecture/generation.mmd new file mode 100644 index 00000000..352a888d --- /dev/null +++ b/docs/diagrams/architecture/generation.mmd @@ -0,0 +1,11 @@ +flowchart TD + CLI["CLI Command Entry Point"] --> CDG["CLIDocumentationGenerator"] + CDG --> PT["ProgressTracker (utils)"] + CDG --> Job["DocumentationJob (job_models)"] + CDG --> BC["BackendConfig (config-core)"] + CDG --> DG["DocumentationGenerator (backend-core)"] + DG --> GB["DependencyGraphBuilder"] + DG --> AO["AgentOrchestrator"] + CDG --> CM["cluster_modules (backend-core)"] + CDG --> HG["HTMLGenerator (html_generation)"] + CDG --> LOG["ColoredFormatter (backend-core logging)"] diff --git a/docs/diagrams/architecture/git_integration-2.mmd b/docs/diagrams/architecture/git_integration-2.mmd new file mode 100644 index 00000000..2339eeb2 --- /dev/null +++ b/docs/diagrams/architecture/git_integration-2.mmd @@ -0,0 +1,26 @@ +sequenceDiagram + participant CLI as "CLI Entry Point" + participant GM as "GitManager" + participant DocGen as "CLIDocumentationGenerator" + participant Repo as "git.Repo" + + CLI->>GM: "GitManager(repo_path)" + GM->>Repo: "git.Repo(repo_path)" + CLI->>GM: "check_clean_working_directory()" + GM-->>CLI: "(is_clean, status_message)" + alt "Working directory dirty and not forced" + GM-->>CLI: "raise RepositoryError" + else "Clean or forced" + CLI->>GM: "create_documentation_branch(force)" + GM->>Repo: "create_head(branch_name)" + GM->>Repo: "checkout()" + GM-->>CLI: "branch_name" + CLI->>DocGen: "generate()" + DocGen-->>CLI: "DocumentationJob" + CLI->>GM: "commit_documentation(docs_path, message)" + GM->>Repo: "index.add([docs_path])" + GM->>Repo: "index.commit(message)" + GM-->>CLI: "commit_hash" + CLI->>GM: "get_github_pr_url(branch_name)" + GM-->>CLI: "pr_url or None" + end diff --git a/docs/diagrams/architecture/git_integration-3.mmd b/docs/diagrams/architecture/git_integration-3.mmd new file mode 100644 index 00000000..da45ed9e --- /dev/null +++ b/docs/diagrams/architecture/git_integration-3.mmd @@ -0,0 +1,13 @@ +flowchart TD + Start["create_documentation_branch(force)"] --> CheckForce{{force?}} + CheckForce -->|"false"| CheckClean["check_clean_working_directory()"] + CheckForce -->|"true"| GenName["Generate timestamped branch name"] + CheckClean --> IsClean{{"is_clean?"}} + IsClean -->|"false"| RaiseErr["raise RepositoryError with remediation steps"] + IsClean -->|"true"| GenName + GenName --> CheckExists{{"branch name exists?"}} + CheckExists -->|"true"| AppendCounter["Append incrementing counter suffix"] + AppendCounter --> CheckExists + CheckExists -->|"false"| CreateBranch["repo.create_head(branch_name)"] + CreateBranch --> Checkout["new_branch.checkout()"] + Checkout --> ReturnName["return branch_name"] diff --git a/docs/diagrams/architecture/git_integration.mmd b/docs/diagrams/architecture/git_integration.mmd new file mode 100644 index 00000000..b8d0ca9d --- /dev/null +++ b/docs/diagrams/architecture/git_integration.mmd @@ -0,0 +1,18 @@ +classDiagram + class GitManager { + +Path repo_path + +Repo repo + +__init__(repo_path) + +check_clean_working_directory() Tuple + +create_documentation_branch(force) str + +commit_documentation(docs_path, message) str + +get_remote_url(remote_name) str + +get_current_branch() str + +get_commit_hash() str + +branch_exists(branch_name) bool + +get_github_pr_url(branch_name) str + } + class RepositoryError { + +__init__(message) + } + GitManager ..> RepositoryError : raises diff --git a/docs/diagrams/architecture/graph_construction-2.mmd b/docs/diagrams/architecture/graph_construction-2.mmd new file mode 100644 index 00000000..a6ed0636 --- /dev/null +++ b/docs/diagrams/architecture/graph_construction-2.mmd @@ -0,0 +1,19 @@ +flowchart TD + Start["parse_repository()"] --> Check{{"len(repo_paths) == 1?"}} + Check -->|"Yes"| Single["_parse_single_repository()"] + Check -->|"No"| Multi["_parse_multiple_repositories()"] + + Single --> S1["analysis_service._analyze_structure()"] + S1 --> S2["analysis_service._analyze_call_graph()"] + S2 --> S3["_build_components_from_analysis()"] + S3 --> SResult["self.components"] + + Multi --> M1["For each repo path: compute namespace"] + M1 --> M2["analysis_service._analyze_structure()"] + M2 --> M3["analysis_service._analyze_call_graph()"] + M3 --> M4["_build_namespaced_components()"] + M4 --> M5["Merge into all_components"] + M5 --> M6{{"More paths?"}} + M6 -->|"Yes"| M1 + M6 -->|"No"| M7["_resolve_cross_namespace_dependencies()"] + M7 --> MResult["self.components"] diff --git a/docs/diagrams/architecture/graph_construction-3.mmd b/docs/diagrams/architecture/graph_construction-3.mmd new file mode 100644 index 00000000..1fc2338c --- /dev/null +++ b/docs/diagrams/architecture/graph_construction-3.mmd @@ -0,0 +1,15 @@ +flowchart TD + A["build_dependency_graph()"] --> B["Ensure dependency_graph_dir exists"] + B --> C["Compute sanitized dependency_graph_path"] + C --> D["Resolve include/exclude patterns from Config"] + D --> E["Build repo_paths from config.all_source_paths"] + E --> F["Instantiate DependencyParser"] + F --> G["parser.parse_repository()"] + G --> H["Log component type breakdown"] + H --> I["parser.save_dependency_graph(path)"] + I --> J["build_graph_from_components(components)"] + J --> K["validate_graph_completeness(components, graph)"] + K --> L["get_leaf_nodes(graph, components)"] + L --> M["Determine valid leaf types from available component types"] + M --> N["Filter leaf_nodes: skip invalid / wrong-type / not-found"] + N --> O["Return (components, keep_leaf_nodes)"] diff --git a/docs/diagrams/architecture/graph_construction-4.mmd b/docs/diagrams/architecture/graph_construction-4.mmd new file mode 100644 index 00000000..5e5d67a9 --- /dev/null +++ b/docs/diagrams/architecture/graph_construction-4.mmd @@ -0,0 +1,8 @@ +flowchart TD + Start["For each leaf_node in leaf_nodes"] --> V1{{"Is leaf_node a valid non-empty string without error keywords?"}} + V1 -->|"No"| SkipInvalid["skipped_invalid += 1"] + V1 -->|"Yes"| V2{{"leaf_node in components?"}} + V2 -->|"No"| SkipNotFound["skipped_not_found += 1"] + V2 -->|"Yes"| V3{{"component_type in valid_types?"}} + V3 -->|"No"| SkipType["skipped_type += 1"] + V3 -->|"Yes"| Keep["keep_leaf_nodes.append(leaf_node)"] diff --git a/docs/diagrams/architecture/graph_construction-5.mmd b/docs/diagrams/architecture/graph_construction-5.mmd new file mode 100644 index 00000000..68967328 --- /dev/null +++ b/docs/diagrams/architecture/graph_construction-5.mmd @@ -0,0 +1,22 @@ +sequenceDiagram + participant Caller as "Pipeline Caller" + participant Builder as "DependencyGraphBuilder" + participant Parser as "DependencyParser" + participant Svc as "AnalysisService" + participant FS as "File System" + + Caller->>Builder: build_dependency_graph() + Builder->>Parser: new DependencyParser(repo_paths, patterns) + Builder->>Parser: parse_repository() + Parser->>Svc: _analyze_structure(repo_path) + Svc-->>Parser: structure_result + Parser->>Svc: _analyze_call_graph(file_tree, repo_path) + Svc-->>Parser: call_graph_result + Parser->>Parser: _build_components_from_analysis() / _build_namespaced_components() + Parser-->>Builder: components (Dict[str, Node]) + Builder->>Parser: save_dependency_graph(path) + Parser->>FS: write JSON + Builder->>Builder: build_graph_from_components(components) + Builder->>Builder: validate_graph_completeness(components, graph) + Builder->>Builder: get_leaf_nodes(graph, components) + Builder-->>Caller: (components, leaf_nodes) diff --git a/docs/diagrams/architecture/graph_construction.mmd b/docs/diagrams/architecture/graph_construction.mmd new file mode 100644 index 00000000..32fb052a --- /dev/null +++ b/docs/diagrams/architecture/graph_construction.mmd @@ -0,0 +1,11 @@ +flowchart TD + Config["Config"] --> Builder["DependencyGraphBuilder"] + Builder -->|"instantiates"| Parser["DependencyParser"] + Parser -->|"uses"| AnalysisSvc["AnalysisService"] + AnalysisSvc -->|"structure + call graph"| Parser + Parser -->|"produces"| Nodes["Node objects (Dict[str, Node])"] + Builder -->|"build_graph_from_components()"| TraversalGraph["In-memory traversal graph"] + Builder -->|"validate_graph_completeness()"| Validation["Graph validation"] + Builder -->|"get_leaf_nodes()"| LeafFilter["Leaf node filtering"] + Parser -->|"save_dependency_graph()"| JSONFile["dependency_graph.json"] + LeafFilter --> Output["(components, leaf_nodes)"] diff --git a/docs/diagrams/architecture/html_generation-2.mmd b/docs/diagrams/architecture/html_generation-2.mmd new file mode 100644 index 00000000..12aad20b --- /dev/null +++ b/docs/diagrams/architecture/html_generation-2.mmd @@ -0,0 +1,19 @@ +sequenceDiagram + participant CLIGen as "CLIDocumentationGenerator" + participant HTMLGen as "HTMLGenerator" + participant FS as "safe_read / safe_write" + participant Git as "GitPython" + + CLIGen->>HTMLGen: HTMLGenerator() + CLIGen->>HTMLGen: detect_repository_info(repo_path) + HTMLGen->>Git: Repo(repo_path) + Git-->>HTMLGen: remote URL, name + HTMLGen-->>CLIGen: name, url, github_pages_url + CLIGen->>HTMLGen: generate(output_path, title, repository_url, docs_dir) + HTMLGen->>FS: safe_read(module_tree.json) + HTMLGen->>FS: safe_read(metadata.json) + HTMLGen->>FS: safe_read(viewer_template.html) + HTMLGen->>HTMLGen: _build_info_content(metadata) + HTMLGen->>HTMLGen: substitute placeholders + HTMLGen->>FS: safe_write(index.html) + HTMLGen-->>CLIGen: index.html written diff --git a/docs/diagrams/architecture/html_generation-3.mmd b/docs/diagrams/architecture/html_generation-3.mmd new file mode 100644 index 00000000..d4be53b1 --- /dev/null +++ b/docs/diagrams/architecture/html_generation-3.mmd @@ -0,0 +1,17 @@ +flowchart TD + Start["generate() called"] --> CheckDocsDir{{"docs_dir provided?"}} + CheckDocsDir -->|"yes"| AutoLoadTree["load_module_tree(docs_dir)"] + CheckDocsDir -->|"yes"| AutoLoadMeta["load_metadata(docs_dir)"] + CheckDocsDir -->|"no"| UseProvided["use provided module_tree/metadata"] + AutoLoadTree --> Defaults + AutoLoadMeta --> Defaults + UseProvided --> Defaults + Defaults["Apply defaults for module_tree/config"] --> LoadTemplate["Load viewer_template.html"] + LoadTemplate -->|"missing"| RaiseErr["raise FileSystemError"] + LoadTemplate -->|"found"| BuildInfo["_build_info_content(metadata)"] + BuildInfo --> BuildRepoLink["Build repository link HTML"] + BuildRepoLink --> ComputeBasePath["Compute docs_base_path"] + ComputeBasePath --> SerializeJson["Serialize config/module_tree/metadata to JSON"] + SerializeJson --> Replace["Replace template placeholders"] + Replace --> WriteOut["safe_write(output_path)"] + WriteOut --> Done["index.html generated"] diff --git a/docs/diagrams/architecture/html_generation.mmd b/docs/diagrams/architecture/html_generation.mmd new file mode 100644 index 00000000..fe3dd6cf --- /dev/null +++ b/docs/diagrams/architecture/html_generation.mmd @@ -0,0 +1,10 @@ +flowchart TD + Docs["Documentation Output Directory"] --> ModuleTreeJson["module_tree.json"] + Docs --> MetadataJson["metadata.json"] + Template["viewer_template.html"] --> Generator["HTMLGenerator"] + ModuleTreeJson --> Generator + MetadataJson --> Generator + RepoPath["Repository Path"] --> DetectInfo["detect_repository_info()"] + DetectInfo --> Generator + Generator -->|"safe_write()"| IndexHtml["index.html"] + Generator -->|"on failure"| FSError["FileSystemError"] diff --git a/docs/diagrams/architecture/java_analyzer-2.mmd b/docs/diagrams/architecture/java_analyzer-2.mmd new file mode 100644 index 00000000..aca32f56 --- /dev/null +++ b/docs/diagrams/architecture/java_analyzer-2.mmd @@ -0,0 +1,16 @@ +sequenceDiagram + participant Caller as "analyze_java_file()" + participant Analyzer as "TreeSitterJavaAnalyzer" + participant Parser as "tree_sitter.Parser" + participant NodesPass as "_extract_nodes()" + participant RelsPass as "_extract_relationships()" + + Caller->>Analyzer: "__init__(file_path, content, repo_path)" + Analyzer->>Analyzer: "_analyze()" + Analyzer->>Parser: "parse(content)" + Parser-->>Analyzer: "AST root node" + Analyzer->>NodesPass: "_extract_nodes(root, top_level_nodes, lines)" + NodesPass-->>Analyzer: "self.nodes populated" + Analyzer->>RelsPass: "_extract_relationships(root, top_level_nodes)" + RelsPass-->>Analyzer: "self.call_relationships populated" + Analyzer-->>Caller: "nodes, call_relationships" diff --git a/docs/diagrams/architecture/java_analyzer-3.mmd b/docs/diagrams/architecture/java_analyzer-3.mmd new file mode 100644 index 00000000..ccd70e87 --- /dev/null +++ b/docs/diagrams/architecture/java_analyzer-3.mmd @@ -0,0 +1,6 @@ +flowchart LR + A["class_declaration
with superclass"] -->|"1: Inheritance"| R1["CallRelationship
class extends BaseClass"] + B["class/enum/record
with super_interfaces"] -->|"2: Interface Implementation"| R2["CallRelationship
class implements Interface"] + C["field_declaration"] -->|"3: Field Type Use"| R3["CallRelationship
class has field of Type"] + D["method_invocation"] -->|"4: Method Call"| R4["CallRelationship
caller calls object.method()"] + E["object_creation_expression"] -->|"5: Object Creation"| R5["CallRelationship
class creates new Type()"] diff --git a/docs/diagrams/architecture/java_analyzer.mmd b/docs/diagrams/architecture/java_analyzer.mmd new file mode 100644 index 00000000..0f8c5faf --- /dev/null +++ b/docs/diagrams/architecture/java_analyzer.mmd @@ -0,0 +1,19 @@ +flowchart TD + subgraph Pipeline["Dependency Analysis Pipeline"] + CGA["CallGraphAnalyzer"] -->|"dispatches .java files"| JA["TreeSitterJavaAnalyzer"] + end + + subgraph JavaAnalyzerModule["Java Analyzer"] + JA -->|"1: parse source"| TS["tree-sitter-java grammar"] + TS -->|"AST"| EN["_extract_nodes()"] + TS -->|"AST"| ER["_extract_relationships()"] + EN -->|"appends"| Nodes["self.nodes: List[Node]"] + ER -->|"appends"| Rels["self.call_relationships: List[CallRelationship]"] + end + + Nodes -->|"consumed by"| DP["DependencyParser"] + Rels -->|"consumed by"| DP + DP -->|"builds"| Graph["Dependency Graph"] + + Node["Node model"] -.->|"defines schema for"| Nodes + CallRel["CallRelationship model"] -.->|"defines schema for"| Rels diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers-2.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers-2.mmd new file mode 100644 index 00000000..62ab2cab --- /dev/null +++ b/docs/diagrams/architecture/javascript_typescript_analyzers-2.mmd @@ -0,0 +1,45 @@ +classDiagram + class TreeSitterJSAnalyzer { + +file_path: Path + +content: str + +repo_path: str + +nodes: List~Node~ + +call_relationships: List~CallRelationship~ + +top_level_nodes: dict + +analyze() None + -_extract_functions(node) None + -_extract_call_relationships(node) None + -_extract_jsdoc_type_dependencies(node, caller) None + } + class TreeSitterTSAnalyzer { + +file_path: Path + +content: str + +repo_path: str + +nodes: List~Node~ + +call_relationships: List~CallRelationship~ + +top_level_nodes: dict + +analyze() None + -_extract_all_entities(node, all_entities, depth) None + -_filter_top_level_declarations(all_entities) None + -_extract_all_relationships(node, all_entities) None + } + class Node { + +id: str + +name: str + +component_type: str + +source_code: str + +start_line: int + +end_line: int + +node_type: str + +base_classes: List~str~ + } + class CallRelationship { + +caller: str + +callee: str + +call_line: int + +is_resolved: bool + } + TreeSitterJSAnalyzer --> Node : produces + TreeSitterJSAnalyzer --> CallRelationship : produces + TreeSitterTSAnalyzer --> Node : produces + TreeSitterTSAnalyzer --> CallRelationship : produces diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers-3.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers-3.mmd new file mode 100644 index 00000000..82f1d2c5 --- /dev/null +++ b/docs/diagrams/architecture/javascript_typescript_analyzers-3.mmd @@ -0,0 +1,19 @@ +flowchart TD + Start["analyze()"] --> Parse["Parser.parse(content)"] + Parse --> ExtractFn["_extract_functions(root_node)"] + ExtractFn --> TraverseFn["_traverse_for_functions(node)"] + TraverseFn -->|"class/abstract/interface"| ExtractClass["_extract_class_declaration()"] + ExtractClass --> ExtractMethods["_extract_methods_from_class()"] + TraverseFn -->|"function_declaration"| ExtractFunc["_extract_function_declaration()"] + TraverseFn -->|"export_statement"| ExtractExport["_extract_exported_function()"] + TraverseFn -->|"lexical_declaration"| ExtractArrow["_extract_arrow_function_from_declaration()"] + ExtractFn --> ExtractCalls["_extract_call_relationships(root_node)"] + ExtractCalls --> TraverseCalls["_traverse_for_calls(node, current_top_level)"] + TraverseCalls -->|"call_expression"| CallRel["_extract_call_from_node()"] + TraverseCalls -->|"new_expression"| NewRel["Constructor relationship"] + TraverseCalls -->|"class_heritage"| InheritRel["Inheritance relationship"] + TraverseCalls -->|"JSDoc comment"| JsdocRel["_parse_jsdoc_types()"] + CallRel --> AddRel["_add_relationship() (dedup)"] + NewRel --> AddRel + InheritRel --> AddRel + JsdocRel --> AddRel diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers-4.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers-4.mmd new file mode 100644 index 00000000..26d20004 --- /dev/null +++ b/docs/diagrams/architecture/javascript_typescript_analyzers-4.mmd @@ -0,0 +1,17 @@ +flowchart TD + A["analyze()"] --> B["Pass 1: _extract_all_entities()
builds all_entities dict"] + B --> C["Pass 2: _filter_top_level_declarations()
_is_actually_top_level() check"] + C --> D["Node objects appended to self.nodes
and self.top_level_nodes"] + D --> E["_extract_constructor_dependencies()
(classes only)"] + B --> F["Pass 3: _extract_all_relationships()"] + F --> G["_traverse_for_relationships()
tracks current_top_level"] + G -->|"call_expression"| H["_extract_call_relationship()"] + G -->|"new_expression"| I["_extract_new_relationship()"] + G -->|"member_expression"| J["_extract_member_relationship()"] + G -->|"type_annotation / type_arguments"| K["_extract_type_relationship()"] + G -->|"extends_clause / implements_clause"| L["_extract_inheritance_relationship()"] + H --> M["_add_relationship()"] + I --> M + J --> M + K --> M + L --> M diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers-5.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers-5.mmd new file mode 100644 index 00000000..eb338103 --- /dev/null +++ b/docs/diagrams/architecture/javascript_typescript_analyzers-5.mmd @@ -0,0 +1,19 @@ +sequenceDiagram + participant DP as "DependencyParser" + participant JS as "TreeSitterJSAnalyzer" + participant TS as "TreeSitterTSAnalyzer" + participant DGB as "DependencyGraphBuilder" + + DP->>JS: analyze_javascript_file_treesitter(path, content, repo_path) + JS->>JS: Parser.parse(content) + JS->>JS: _extract_functions() / _extract_call_relationships() + JS-->>DP: (nodes, call_relationships) + + DP->>TS: analyze_typescript_file_treesitter(path, content, repo_path) + TS->>TS: Parser.parse(content) + TS->>TS: _extract_all_entities() / _filter_top_level_declarations() / _extract_all_relationships() + TS-->>DP: (nodes, call_relationships) + + DP->>DGB: aggregated nodes + relationships + DGB->>DGB: resolve callee ids against repository Node index + DGB-->>DP: Repository dependency graph diff --git a/docs/diagrams/architecture/javascript_typescript_analyzers.mmd b/docs/diagrams/architecture/javascript_typescript_analyzers.mmd new file mode 100644 index 00000000..47fbe960 --- /dev/null +++ b/docs/diagrams/architecture/javascript_typescript_analyzers.mmd @@ -0,0 +1,10 @@ +flowchart LR + DP["DependencyParser"] -->|"dispatches .js/.jsx/.mjs/.cjs files"| JSA["TreeSitterJSAnalyzer"] + DP -->|"dispatches .ts/.tsx files"| TSA["TreeSitterTSAnalyzer"] + JSA -->|"produces"| Nodes["Node objects"] + JSA -->|"produces"| Rels["CallRelationship objects"] + TSA -->|"produces"| Nodes + TSA -->|"produces"| Rels + Nodes --> DGB["DependencyGraphBuilder"] + Rels --> DGB + DGB --> Graph["Repository Dependency Graph"] diff --git a/docs/diagrams/architecture/job_models-2.mmd b/docs/diagrams/architecture/job_models-2.mmd new file mode 100644 index 00000000..fc6bbd1a --- /dev/null +++ b/docs/diagrams/architecture/job_models-2.mmd @@ -0,0 +1,7 @@ +stateDiagram-v2 + [*] --> PENDING : "DocumentationJob() created" + PENDING --> RUNNING : "start()" + RUNNING --> COMPLETED : "complete()" + RUNNING --> FAILED : "fail(error_message)" + COMPLETED --> [*] + FAILED --> [*] diff --git a/docs/diagrams/architecture/job_models-3.mmd b/docs/diagrams/architecture/job_models-3.mmd new file mode 100644 index 00000000..5cc40d39 --- /dev/null +++ b/docs/diagrams/architecture/job_models-3.mmd @@ -0,0 +1,20 @@ +sequenceDiagram + participant Caller as "CLI Component" + participant Job as "DocumentationJob" + participant Coerce as "Coercion Helpers" + + Caller->>Job: "DocumentationJob(...)" + Job-->>Caller: "job instance (status=PENDING)" + Caller->>Job: "job.start()" + Caller->>Job: "job.to_dict() / job.to_json()" + Job-->>Caller: "dict / JSON string" + Note over Caller: "Persisted to disk or transmitted" + + Caller->>Job: "DocumentationJob.from_dict(raw_data)" + Job->>Coerce: "_coerce_job_status(raw_data.status)" + Job->>Coerce: "_coerce_int(raw_data.module_count)" + Job->>Coerce: "_coerce_generation_options(raw_data.generation_options)" + Job->>Coerce: "_coerce_llm_config(raw_data.llm_config)" + Job->>Coerce: "_coerce_statistics(raw_data.statistics)" + Coerce-->>Job: "validated nested objects" + Job-->>Caller: "reconstructed DocumentationJob" diff --git a/docs/diagrams/architecture/job_models-4.mmd b/docs/diagrams/architecture/job_models-4.mmd new file mode 100644 index 00000000..0fc4c4ce --- /dev/null +++ b/docs/diagrams/architecture/job_models-4.mmd @@ -0,0 +1,15 @@ +flowchart TD + JobModels["Job Models"] + CliCore["Cli Core (parent)"] + Generation["Generation"] + Configuration["Configuration"] + GitIntegration["Git Integration"] + HtmlGeneration["Html Generation"] + Utils["Utils"] + + CliCore --> JobModels + Generation -->|"creates & drives"| JobModels + Configuration -->|"populates LLMConfig"| JobModels + GitIntegration -->|"populates commit_hash / branch_name"| JobModels + HtmlGeneration -->|"reads files_generated"| JobModels + Utils -->|"reports on status/statistics"| JobModels diff --git a/docs/diagrams/architecture/job_models.mmd b/docs/diagrams/architecture/job_models.mmd new file mode 100644 index 00000000..80bc5699 --- /dev/null +++ b/docs/diagrams/architecture/job_models.mmd @@ -0,0 +1,52 @@ +classDiagram + class JobStatus { + <> + PENDING + RUNNING + COMPLETED + FAILED + } + class GenerationOptions { + +bool create_branch + +bool github_pages + +bool no_cache + +str custom_output + } + class LLMConfig { + +str main_model + +str cluster_model + +str base_url + } + class JobStatistics { + +int total_files_analyzed + +int leaf_nodes + +int max_depth + +int total_tokens_used + } + class DocumentationJob { + +str job_id + +str repository_path + +str repository_name + +str output_directory + +str commit_hash + +str branch_name + +str timestamp_start + +str timestamp_end + +JobStatus status + +str error_message + +List files_generated + +int module_count + +GenerationOptions generation_options + +LLMConfig llm_config + +JobStatistics statistics + +start() + +complete() + +fail(error_message) + +to_dict() dict + +to_json() str + +from_dict(data) DocumentationJob + } + DocumentationJob --> JobStatus : "status" + DocumentationJob --> GenerationOptions : "generation_options" + DocumentationJob --> LLMConfig : "llm_config (optional)" + DocumentationJob --> JobStatistics : "statistics" diff --git a/docs/diagrams/architecture/llm-services-2.mmd b/docs/diagrams/architecture/llm-services-2.mmd new file mode 100644 index 00000000..9606e040 --- /dev/null +++ b/docs/diagrams/architecture/llm-services-2.mmd @@ -0,0 +1,20 @@ +flowchart LR + subgraph LlmServices["Llm Services"] + Factory["Model factories
(main / fallback / cluster)"] + Counting["CountingFallbackModel"] + DirectCall["call_llm()"] + end + + subgraph DocGen["Documentation Generator"] + DG["DocumentationGenerator"] + Orchestrator["AgentOrchestrator"] + end + + ConfigCore["Config"] --> Factory + Factory --> Counting + Orchestrator -->|"create_fallback_models(config)"| Factory + Orchestrator -->|"used by pydantic-ai Agent"| Counting + DG -->|"call_llm(prompt, config)"| DirectCall + + DirectCall --> OpenAISDK["openai.OpenAI client"] + Counting --> PydanticAI["pydantic-ai FallbackModel / OpenAIProvider"] diff --git a/docs/diagrams/architecture/llm-services-3.mmd b/docs/diagrams/architecture/llm-services-3.mmd new file mode 100644 index 00000000..7d804584 --- /dev/null +++ b/docs/diagrams/architecture/llm-services-3.mmd @@ -0,0 +1,22 @@ +sequenceDiagram + participant Caller + participant Factory as "create_main_model() / create_fallback_model() / create_cluster_model()" + participant Config as "Config" + participant Provider as "OpenAIProvider" + participant Model as "OpenAIModel" + + Caller->>Factory: create_X_model(config) + Factory->>Config: getattr(config, "X_max_tokens", None) + Factory->>Config: getattr(config, "X_temperature", 0.0) + Factory->>Config: getattr(config, "X_temperature_supported", True) + Factory->>Config: getattr(config, "X_base_url", None) + alt base_url missing + Factory-->>Caller: raise ValueError("X_base_url is required...") + end + Factory->>Config: getattr(config, "X_api_key", None) + alt api_key missing + Factory-->>Caller: raise ValueError("X_api_key is required...") + end + Factory->>Provider: OpenAIProvider(base_url, api_key) + Factory->>Model: OpenAIModel(model_name, provider, settings) + Model-->>Caller: configured OpenAIModel diff --git a/docs/diagrams/architecture/llm-services-4.mmd b/docs/diagrams/architecture/llm-services-4.mmd new file mode 100644 index 00000000..2a21e453 --- /dev/null +++ b/docs/diagrams/architecture/llm-services-4.mmd @@ -0,0 +1,16 @@ +sequenceDiagram + participant Orchestrator as "AgentOrchestrator" + participant LlmServices as "llm_services" + participant MainModel as "OpenAIModel (main)" + participant FallbackModelObj as "OpenAIModel (fallback)" + participant CFM as "CountingFallbackModel" + + Orchestrator->>LlmServices: create_fallback_models(config) + LlmServices->>LlmServices: create_main_model(config) + LlmServices-->>MainModel: main + LlmServices->>LlmServices: create_fallback_model(config) + LlmServices-->>FallbackModelObj: fallback + LlmServices->>CFM: CountingFallbackModel(main, fallback) + CFM-->>Orchestrator: fallback_models + + Note over Orchestrator,CFM: Passed as the model backend
for pydantic-ai Agent diff --git a/docs/diagrams/architecture/llm-services-5.mmd b/docs/diagrams/architecture/llm-services-5.mmd new file mode 100644 index 00000000..e86222d9 --- /dev/null +++ b/docs/diagrams/architecture/llm-services-5.mmd @@ -0,0 +1,24 @@ +sequenceDiagram + participant Caller + participant CallLLM as "call_llm()" + participant ClientFactory as "create_openai_client()" + participant Config as "Config" + participant OpenAISDK as "openai.OpenAI" + + Caller->>CallLLM: call_llm(prompt, config, model, temperature) + CallLLM->>CallLLM: resolve stage (main / cluster / fallback) + CallLLM->>ClientFactory: create_openai_client(config, model) + ClientFactory->>Config: resolve base_url / api_key / api_version for stage + alt base_url or api_key missing + ClientFactory-->>CallLLM: raise ValueError(...) + end + ClientFactory->>OpenAISDK: OpenAI(base_url, api_key, default_headers) + OpenAISDK-->>CallLLM: client + CallLLM->>CallLLM: get_model_max_token_field(stage) + CallLLM->>OpenAISDK: client.chat.completions.create(**kwargs) + alt success + OpenAISDK-->>CallLLM: response + CallLLM-->>Caller: response_content + else OpenAIError + CallLLM-->>Caller: raise RuntimeError(context-wrapped) + end diff --git a/docs/diagrams/architecture/llm-services.mmd b/docs/diagrams/architecture/llm-services.mmd new file mode 100644 index 00000000..aa28dfd3 --- /dev/null +++ b/docs/diagrams/architecture/llm-services.mmd @@ -0,0 +1,24 @@ +flowchart TD + Config["Config"] --> CreateMain["create_main_model()"] + Config --> CreateFallback["create_fallback_model()"] + Config --> CreateCluster["create_cluster_model()"] + Config --> CreateClient["create_openai_client()"] + + CreateMain --> MainModel["OpenAIModel (main)"] + CreateFallback --> FallbackModel_["OpenAIModel (fallback)"] + + MainModel --> CreateChain["create_fallback_models()"] + FallbackModel_ --> CreateChain + + CreateChain --> CountingModel["CountingFallbackModel"] + CountingModel --> IncCounter["increment_request_counter()"] + IncCounter --> Counter[("_request_counter")] + + CountingModel --> Agent["pydantic-ai Agent"] + + CreateCluster --> ClusterModel["OpenAIModel (cluster)"] + ClusterModel --> ClusterAgent["Clustering LLM calls"] + + CreateClient --> OpenAIClient["openai.OpenAI client"] + OpenAIClient --> CallLLM["call_llm()"] + CallLLM --> Response["LLM response text"] diff --git a/docs/diagrams/architecture/logging-config-2.mmd b/docs/diagrams/architecture/logging-config-2.mmd new file mode 100644 index 00000000..9387e9d5 --- /dev/null +++ b/docs/diagrams/architecture/logging-config-2.mmd @@ -0,0 +1,21 @@ +sequenceDiagram + participant Caller as "Backend Component" + participant Setup as "setup_module_logging()" + participant Logger as "logging.Logger" + participant Handler as "StreamHandler(stdout)" + participant Fmt as "ColoredFormatter" + + Caller->>Setup: setup_module_logging("my_module", level) + Setup->>Logger: logging.getLogger("my_module") + Setup->>Handler: create StreamHandler(sys.stdout) + Setup->>Fmt: create ColoredFormatter() + Setup->>Handler: setFormatter(Fmt) + Setup->>Logger: handlers.clear() + Setup->>Logger: addHandler(Handler) + Setup->>Logger: propagate = False + Setup-->>Caller: return configured logger + Caller->>Logger: logger.info("message") + Logger->>Handler: emit(record) + Handler->>Fmt: format(record) + Fmt-->>Handler: colored log line + Handler-->>Caller: printed to stdout diff --git a/docs/diagrams/architecture/logging-config-3.mmd b/docs/diagrams/architecture/logging-config-3.mmd new file mode 100644 index 00000000..f5e7ee0e --- /dev/null +++ b/docs/diagrams/architecture/logging-config-3.mmd @@ -0,0 +1,13 @@ +classDiagram + class Formatter { + <> + +format(record) + +formatTime(record, datefmt) + +formatException(exc_info) + } + class ColoredFormatter { + +COLORS : dict + +COMPONENT_COLORS : dict + +format(record) str + } + Formatter <|-- ColoredFormatter diff --git a/docs/diagrams/architecture/logging-config-4.mmd b/docs/diagrams/architecture/logging-config-4.mmd new file mode 100644 index 00000000..9d71e38e --- /dev/null +++ b/docs/diagrams/architecture/logging-config-4.mmd @@ -0,0 +1,12 @@ +graph TD + Backend["Backend Core"] --> LoggingConfig["Logging Config"] + Backend --> DepAnalyzer["Dependency Analyzer Core"] + Backend --> TreeSitter["Tree-Sitter Analyzers"] + Backend --> DepModels["Dependency Analyzer Models"] + Backend --> DocGen["Documentation Generator"] + Backend --> LLMServices["LLM Services"] + Backend --> AgentTools["Agent Tools Core"] + + DepAnalyzer -.->|"colored console output"| LoggingConfig + TreeSitter -.->|"colored console output"| LoggingConfig + DocGen -.->|"colored console output"| LoggingConfig diff --git a/docs/diagrams/architecture/logging-config.mmd b/docs/diagrams/architecture/logging-config.mmd new file mode 100644 index 00000000..1a14fc08 --- /dev/null +++ b/docs/diagrams/architecture/logging-config.mmd @@ -0,0 +1,14 @@ +flowchart TD + A["logging.LogRecord"] --> B["ColoredFormatter.format(record)"] + B --> C["Look up level color in COLORS"] + B --> D["Format timestamp HH:MM:SS"] + D --> E["Wrap timestamp in blue"] + C --> F["Wrap levelname in level color"] + C --> G["Wrap message in level color"] + E --> H["Concatenate: timestamp + level + message"] + F --> H + G --> H + H --> I{"record.exc_info present?"} + I -->|"Yes"| J["Append formatException() output"] + I -->|"No"| K["Return colored log line"] + J --> K diff --git a/docs/diagrams/architecture/php_analyzer-2.mmd b/docs/diagrams/architecture/php_analyzer-2.mmd new file mode 100644 index 00000000..d4697603 --- /dev/null +++ b/docs/diagrams/architecture/php_analyzer-2.mmd @@ -0,0 +1,41 @@ +classDiagram + class NamespaceResolver { + +string current_namespace + +dict use_map + +register_namespace(ns) + +register_use(fqn, alias) + +resolve(name) string + } + class TreeSitterPHPAnalyzer { + +Path file_path + +string content + +string repo_path + +List nodes + +List call_relationships + +NamespaceResolver namespace_resolver + +_analyze() + +_extract_namespace_info(node, depth) + +_extract_nodes(node, lines, depth, parent_class) + +_extract_relationships(node, depth) + +_is_template_file() bool + +_get_component_id(name, parent_class) string + } + class Node { + +string id + +string name + +string component_type + +string file_path + +Set depends_on + +string source_code + +int start_line + +int end_line + } + class CallRelationship { + +string caller + +string callee + +int call_line + +bool is_resolved + } + TreeSitterPHPAnalyzer --> NamespaceResolver : "delegates name resolution" + TreeSitterPHPAnalyzer --> Node : "creates" + TreeSitterPHPAnalyzer --> CallRelationship : "creates" diff --git a/docs/diagrams/architecture/php_analyzer-3.mmd b/docs/diagrams/architecture/php_analyzer-3.mmd new file mode 100644 index 00000000..3ee94d5f --- /dev/null +++ b/docs/diagrams/architecture/php_analyzer-3.mmd @@ -0,0 +1,20 @@ +sequenceDiagram + participant Caller as "Caller" + participant Analyzer as "TreeSitterPHPAnalyzer" + participant Parser as "tree_sitter Parser" + participant Resolver as "NamespaceResolver" + + Caller->>Analyzer: "__init__(file_path, content, repo_path)" + Analyzer->>Analyzer: "_is_template_file()" + alt is template file + Analyzer-->>Caller: "return (skip analysis)" + else not a template + Analyzer->>Parser: "parse(content)" + Parser-->>Analyzer: "AST root node" + Analyzer->>Analyzer: "_extract_namespace_info(root)" + Analyzer->>Resolver: "register_namespace() / register_use()" + Analyzer->>Analyzer: "_extract_nodes(root, lines)" + Analyzer->>Analyzer: "_extract_relationships(root)" + Analyzer->>Resolver: "resolve(type_name)" + Analyzer-->>Caller: "nodes, call_relationships populated" + end diff --git a/docs/diagrams/architecture/php_analyzer-4.mmd b/docs/diagrams/architecture/php_analyzer-4.mmd new file mode 100644 index 00000000..e3ac6b68 --- /dev/null +++ b/docs/diagrams/architecture/php_analyzer-4.mmd @@ -0,0 +1,7 @@ +flowchart LR + File["PHP source file"] --> Analyzer["TreeSitterPHPAnalyzer"] + Analyzer -->|"nodes"| NodeList["List of Node objects
(classes, methods, functions, etc.)"] + Analyzer -->|"relationships"| RelList["List of CallRelationship objects
(extends, implements, new, static call)"] + NodeList --> Parser["DependencyParser
(Graph Construction)"] + RelList --> Parser + Parser --> Graph["Dependency Graph
(components + depends_on edges)"] diff --git a/docs/diagrams/architecture/php_analyzer.mmd b/docs/diagrams/architecture/php_analyzer.mmd new file mode 100644 index 00000000..b02e83ed --- /dev/null +++ b/docs/diagrams/architecture/php_analyzer.mmd @@ -0,0 +1,9 @@ +flowchart TD + Caller["Dependency Analysis Pipeline"] -->|"analyze_php_file(path, content, repo_path)"| Analyze["analyze_php_file()"] + Analyze --> Analyzer["TreeSitterPHPAnalyzer"] + Analyzer -->|"uses"| Resolver["NamespaceResolver"] + Analyzer -->|"parses with"| TreeSitter["tree_sitter_php / Parser"] + Analyzer -->|"produces"| Nodes["List of Node"] + Analyzer -->|"produces"| Rels["List of CallRelationship"] + Nodes --> Models["Dependency Analyzer Models"] + Rels --> Models diff --git a/docs/diagrams/architecture/python_analyzer-2.mmd b/docs/diagrams/architecture/python_analyzer-2.mmd new file mode 100644 index 00000000..364a74b2 --- /dev/null +++ b/docs/diagrams/architecture/python_analyzer-2.mmd @@ -0,0 +1,4 @@ +flowchart LR + FilePath["file_path"] -->|"os.path.relpath"| RelPath["relative_path"] + RelPath -->|"strip .py/.pyx, replace separators with dots"| ModulePath["module.path"] + ModulePath -->|"+ '::' + ComponentName"| ComponentID["component_id"] diff --git a/docs/diagrams/architecture/python_analyzer-3.mmd b/docs/diagrams/architecture/python_analyzer-3.mmd new file mode 100644 index 00000000..b6853911 --- /dev/null +++ b/docs/diagrams/architecture/python_analyzer-3.mmd @@ -0,0 +1,17 @@ +sequenceDiagram + participant Visitor as PythonASTAnalyzer + participant AST as ast.NodeVisitor + participant Nodes as top_level_nodes + participant Rels as call_relationships + + AST->>Visitor: visit_ClassDef(node) + Visitor->>Visitor: extract base_classes + Visitor->>Nodes: register class Node + Visitor->>Rels: append inheritance CallRelationship (if resolved) + Visitor->>Visitor: set current_class_name + Visitor->>AST: generic_visit(node) + AST->>Visitor: visit_Call(node) [inside class/function body] + Visitor->>Visitor: _get_call_name(node.func) + Visitor->>Nodes: lookup call_name + Visitor->>Rels: append CallRelationship (resolved or unresolved) + Visitor->>Visitor: clear current_class_name diff --git a/docs/diagrams/architecture/python_analyzer-4.mmd b/docs/diagrams/architecture/python_analyzer-4.mmd new file mode 100644 index 00000000..8988784c --- /dev/null +++ b/docs/diagrams/architecture/python_analyzer-4.mmd @@ -0,0 +1,6 @@ +flowchart TD + Start["analyze() called"] --> Parse["ast.parse(content)"] + Parse -->|"success"| Visit["self.visit(tree)"] + Parse -->|"SyntaxError"| WarnLog["log warning, skip file"] + Visit -->|"success"| Done["nodes + call_relationships populated"] + Visit -->|"unexpected Exception"| ErrLog["log error with traceback"] diff --git a/docs/diagrams/architecture/python_analyzer.mmd b/docs/diagrams/architecture/python_analyzer.mmd new file mode 100644 index 00000000..c5db2c86 --- /dev/null +++ b/docs/diagrams/architecture/python_analyzer.mmd @@ -0,0 +1,11 @@ +flowchart TD + RepoAnalyzer["RepoAnalyzer"] -->|"reads .py file"| PyFile["Python Source File"] + RepoAnalyzer -->|"instantiate"| PythonASTAnalyzer["PythonASTAnalyzer"] + PyFile -->|"content"| PythonASTAnalyzer + PythonASTAnalyzer -->|"ast.parse()"| AST["Python AST"] + AST -->|"NodeVisitor traversal"| PythonASTAnalyzer + PythonASTAnalyzer -->|"produces"| Nodes["List of Node"] + PythonASTAnalyzer -->|"produces"| Relationships["List of CallRelationship"] + Nodes --> DependencyGraphBuilder["DependencyGraphBuilder"] + Relationships --> DependencyGraphBuilder + DependencyGraphBuilder --> Graph["Dependency Graph"] diff --git a/docs/diagrams/architecture/sample_fixtures-2.mmd b/docs/diagrams/architecture/sample_fixtures-2.mmd new file mode 100644 index 00000000..116beb42 --- /dev/null +++ b/docs/diagrams/architecture/sample_fixtures-2.mmd @@ -0,0 +1,28 @@ +classDiagram + class APIController { + +service: MainService + +request_count: int + +__init__() + +handle_request(endpoint, data) dict + } + class MainService { + +name: str + +active: bool + +__init__(name) + +start(config) bool + +process_request(data) dict + +stop() + } + class PluginInterface { + +initialize() bool + +execute(context) dict + } + class DataPlugin { + +name: str + +initialized: bool + +__init__(name) + +initialize() bool + +execute(context) dict + } + APIController --> MainService : "uses" + DataPlugin --|> PluginInterface : "implements" diff --git a/docs/diagrams/architecture/sample_fixtures-3.mmd b/docs/diagrams/architecture/sample_fixtures-3.mmd new file mode 100644 index 00000000..d2ebf7b6 --- /dev/null +++ b/docs/diagrams/architecture/sample_fixtures-3.mmd @@ -0,0 +1,11 @@ +sequenceDiagram + participant Client + participant Controller as APIController + participant Service as MainService + + Client->>Controller: handle_request("/process", data) + Controller->>Controller: request_count += 1 + Controller->>Service: process_request(data) + Service->>Service: check active flag + Service-->>Controller: {"status": "success", "result": processed} + Controller-->>Client: response dict diff --git a/docs/diagrams/architecture/sample_fixtures-4.mmd b/docs/diagrams/architecture/sample_fixtures-4.mmd new file mode 100644 index 00000000..968fb7b7 --- /dev/null +++ b/docs/diagrams/architecture/sample_fixtures-4.mmd @@ -0,0 +1,5 @@ +stateDiagram-v2 + [*] --> Uninitialized: "DataPlugin(name)" + Uninitialized --> Initialized: "initialize()" + Initialized --> Initialized: "execute(context)" + Uninitialized --> Error: "execute() before initialize()" diff --git a/docs/diagrams/architecture/sample_fixtures-5.mmd b/docs/diagrams/architecture/sample_fixtures-5.mmd new file mode 100644 index 00000000..d1775cc3 --- /dev/null +++ b/docs/diagrams/architecture/sample_fixtures-5.mmd @@ -0,0 +1,4 @@ +graph LR + Controller["main.controller.APIController"] -->|"imports"| Service["main.service.MainService"] + Service -->|"imports"| Deps["deps.helper (external)"] + Plugin["external.plugin.DataPlugin"] -->|"implements"| Iface["external.plugin.PluginInterface"] diff --git a/docs/diagrams/architecture/sample_fixtures.mmd b/docs/diagrams/architecture/sample_fixtures.mmd new file mode 100644 index 00000000..4a9397b3 --- /dev/null +++ b/docs/diagrams/architecture/sample_fixtures.mmd @@ -0,0 +1,4 @@ +graph TD + Parent["Test Multi Path"] --> SampleFixtures["Sample Fixtures"] + Parent --> TestSuites["Test Suites"] + SampleFixtures -.->|"exercised by"| TestSuites diff --git a/docs/diagrams/architecture/test-clustering-2.mmd b/docs/diagrams/architecture/test-clustering-2.mmd new file mode 100644 index 00000000..667f93aa --- /dev/null +++ b/docs/diagrams/architecture/test-clustering-2.mmd @@ -0,0 +1,23 @@ +classDiagram + class TestResultsBase { + +add_test(name, passed, details) + +print_summary() bool + } + class DebugResults + class ForcedResults + class LocalResults + class IntegrationResults { + +passed int + +failed int + } + class ValidationResults { + +success bool + } + class IdBasedResults + + TestResultsBase <|-- DebugResults + TestResultsBase <|-- ForcedResults + TestResultsBase <|-- LocalResults + TestResultsBase <|-- IntegrationResults + TestResultsBase <|-- ValidationResults + TestResultsBase <|-- IdBasedResults diff --git a/docs/diagrams/architecture/test-clustering-3.mmd b/docs/diagrams/architecture/test-clustering-3.mmd new file mode 100644 index 00000000..bc44214d --- /dev/null +++ b/docs/diagrams/architecture/test-clustering-3.mmd @@ -0,0 +1,15 @@ +sequenceDiagram + participant Dev as Developer + participant Debug as "test_clustering_debug.py" + participant Patch as "Monkey-patched LLM Client" + participant Cluster as "cluster_modules()" + + Dev->>Debug: python3 test_clustering_debug.py + Debug->>Patch: patch create_llm_client() + Debug->>Cluster: cluster_modules(leaf_nodes, components, config) + Cluster->>Patch: client.call(prompt) + Patch-->>Cluster: raw LLM response text + Patch-->>Debug: capture response (global variable) + Cluster-->>Debug: module_tree dict + Debug->>Dev: print captured response + module_tree summary + Debug->>Dev: TestResults.print_summary() diff --git a/docs/diagrams/architecture/test-clustering.mmd b/docs/diagrams/architecture/test-clustering.mmd new file mode 100644 index 00000000..baeef457 --- /dev/null +++ b/docs/diagrams/architecture/test-clustering.mmd @@ -0,0 +1,37 @@ +flowchart TD + subgraph scripts["Test Clustering Scripts"] + Debug["test_clustering_debug.py"] + Forced["test_clustering_forced.py"] + Local["test_clustering_local.py"] + Integration["test_clustering_integration.py"] + Validation["test_clustering_validation.py"] + IdBased["test_id_based_clustering.py"] + end + + subgraph backend["Backend Clustering Pipeline"] + ClusterFn["cluster_modules()"] + IdMap["create_component_id_map()"] + Normalize["normalize_component_ids_by_lookup()"] + LLMClient["LLM Client"] + end + + ConfigMod["Config"] + NodeMod["Node"] + + Debug -->|"invokes"| ClusterFn + Forced -->|"invokes"| ClusterFn + Local -->|"invokes"| ClusterFn + ClusterFn -->|"calls"| LLMClient + + Integration -->|"invokes directly"| IdMap + Integration -->|"invokes directly"| Normalize + + Validation -->|"simulates validation logic of"| ClusterFn + IdBased -->|"simulates parsing/validation logic of"| ClusterFn + + Debug -->|"constructs"| NodeMod + Forced -->|"constructs"| NodeMod + Local -->|"constructs"| NodeMod + Debug -->|"constructs"| ConfigMod + Forced -->|"constructs"| ConfigMod + Local -->|"constructs"| ConfigMod diff --git a/docs/diagrams/architecture/test-multi-path-2.mmd b/docs/diagrams/architecture/test-multi-path-2.mmd new file mode 100644 index 00000000..5b0d9b95 --- /dev/null +++ b/docs/diagrams/architecture/test-multi-path-2.mmd @@ -0,0 +1,12 @@ +sequenceDiagram + participant Suite as "Test Script" + participant Cfg as "Config" + participant Builder as "DependencyGraphBuilder" + + Suite->>Cfg: Create Config with repo_path and additional_source_paths + Suite->>Cfg: validate_source_paths() (optional, for invalid-path test) + Suite->>Builder: DependencyGraphBuilder(config) + Suite->>Builder: build_dependency_graph() + Builder-->>Suite: Returns (components, leaf_nodes) + Suite->>Suite: Assert namespaces, counts, and cross-path edges + Suite->>Suite: Print summary and exit code diff --git a/docs/diagrams/architecture/test-multi-path.mmd b/docs/diagrams/architecture/test-multi-path.mmd new file mode 100644 index 00000000..993d8fc5 --- /dev/null +++ b/docs/diagrams/architecture/test-multi-path.mmd @@ -0,0 +1,25 @@ +flowchart TD + subgraph fixtures["Sample Application Fixtures"] + Service["MainService"] + Controller["APIController"] + Plugin["DataPlugin / PluginInterface"] + end + + subgraph suites["Test Suites and Runners"] + MultiPathSuite["test_multi_path.py suite"] + IntegrationRunner["IntegrationTestRunner"] + end + + subgraph external_deps["External Dependencies"] + ConfigCls["Config"] + Builder["DependencyGraphBuilder"] + end + + MultiPathSuite -->|"constructs"| ConfigCls + IntegrationRunner -->|"constructs"| ConfigCls + ConfigCls -->|"declares additional_source_paths"| Builder + Builder -->|"parses"| Service + Builder -->|"parses"| Controller + Builder -->|"parses"| Plugin + MultiPathSuite -->|"asserts on"| Builder + IntegrationRunner -->|"asserts on"| Builder diff --git a/docs/diagrams/architecture/test_suites-2.mmd b/docs/diagrams/architecture/test_suites-2.mmd new file mode 100644 index 00000000..da3a803e --- /dev/null +++ b/docs/diagrams/architecture/test_suites-2.mmd @@ -0,0 +1,19 @@ +flowchart TD + Start["run_all_tests()"] --> T1["test_single_path()"] + Start --> T2["test_multiple_paths()"] + Start --> T3["test_component_namespacing()"] + Start --> T4["test_cross_path_dependencies()"] + Start --> T5["test_invalid_path_handling()"] + Start --> T6["test_empty_additional_paths()"] + Start --> T7["test_relative_vs_absolute_paths()"] + T1 --> Builder["DependencyGraphBuilder.build_dependency_graph()"] + T2 --> Builder + T3 --> Builder + T4 --> Builder + T6 --> Builder + T7 --> Builder + T5 --> Validate["Config.validate_source_paths()"] + Builder --> Results["TestResults.add_test()"] + Validate --> Results + Results --> Summary["TestResults.print_summary()"] + Summary --> ExitCode["Process exit code (0 or 1)"] diff --git a/docs/diagrams/architecture/test_suites-3.mmd b/docs/diagrams/architecture/test_suites-3.mmd new file mode 100644 index 00000000..f559e4ad --- /dev/null +++ b/docs/diagrams/architecture/test_suites-3.mmd @@ -0,0 +1,25 @@ +sequenceDiagram + participant Main as "main()" + participant Runner as "IntegrationTestRunner" + participant Cfg as "Config" + participant Builder as "DependencyGraphBuilder" + participant Results as "IntegrationTestResults" + + Main->>Runner: run() + Runner->>Runner: setup_test_environment() + Note over Runner: Creates main/, deps/, vendor/ with cross-importing fixture files + Runner->>Cfg: create_config() + Cfg-->>Runner: Config(repo_path, additional_source_paths) + Runner->>Runner: validate_paths() + Runner->>Builder: execute_dependency_parser() + Builder->>Builder: build_dependency_graph() + Builder-->>Runner: components, leaf_nodes + Runner->>Results: verify_namespaces() + Runner->>Results: verify_cross_path_dependencies() + Runner->>Results: verify_no_warnings() + Runner->>Results: verify_file_counts() + Runner->>Runner: print_detailed_output() + Runner->>Results: print_summary() + Results-->>Runner: all_passed() + Runner->>Runner: cleanup() + Runner-->>Main: exit code 0, 1, or 2 diff --git a/docs/diagrams/architecture/test_suites-4.mmd b/docs/diagrams/architecture/test_suites-4.mmd new file mode 100644 index 00000000..df7aea08 --- /dev/null +++ b/docs/diagrams/architecture/test_suites-4.mmd @@ -0,0 +1,15 @@ +flowchart LR + subgraph TS["Test Suites"] + SM["Scenario Test Suite (test_multi_path.py)"] + IT["Integration Test Runner (integration_test.py)"] + end + subgraph CC["Config Core"] + Config["Config"] + end + subgraph DA["Dependency Analyzer Core"] + DGB["DependencyGraphBuilder"] + end + SM --> Config + IT --> Config + SM --> DGB + IT --> DGB diff --git a/docs/diagrams/architecture/test_suites.mmd b/docs/diagrams/architecture/test_suites.mmd new file mode 100644 index 00000000..3b7dc665 --- /dev/null +++ b/docs/diagrams/architecture/test_suites.mmd @@ -0,0 +1,46 @@ +classDiagram + class Colors { + +GREEN + +RED + +YELLOW + +BLUE + +BOLD + +END + } + class ScenarioTestResults { + -results dict + +add_test(name, passed) + +print_summary() int + } + class IntegrationTestRunner { + -test_dir + -main_path + -deps_path + -vendor_path + -config + -builder + -components + -leaf_nodes + -results + +setup_test_environment() + +create_config() + +validate_paths() + +execute_dependency_parser() + +verify_namespaces() + +verify_cross_path_dependencies() + +verify_no_warnings() + +verify_file_counts() + +print_detailed_output() + +cleanup() + +run() int + } + class IntegrationTestResults { + -tests list + +add_test(name, passed, details) + +print_summary() + +all_passed() bool + } + IntegrationTestRunner --> IntegrationTestResults : records into + IntegrationTestRunner --> Config : creates + IntegrationTestRunner --> DependencyGraphBuilder : invokes + ScenarioTestResults ..> Colors : uses for output diff --git a/docs/diagrams/architecture/tree-sitter-analyzers.mmd b/docs/diagrams/architecture/tree-sitter-analyzers.mmd new file mode 100644 index 00000000..4fd82f40 --- /dev/null +++ b/docs/diagrams/architecture/tree-sitter-analyzers.mmd @@ -0,0 +1,22 @@ +flowchart TD + Caller["Dependency Parser"] -->|"dispatches by file extension"| Router["Language Router"] + Router -->|".c / .h"| CAnalyzer["TreeSitterCAnalyzer"] + Router -->|".cpp / .hpp / .cc"| CppAnalyzer["TreeSitterCppAnalyzer"] + Router -->|".cs"| CSharpAnalyzer["TreeSitterCSharpAnalyzer"] + Router -->|".java"| JavaAnalyzer["TreeSitterJavaAnalyzer"] + Router -->|".js / .jsx / .mjs"| JSAnalyzer["TreeSitterJSAnalyzer"] + Router -->|".ts / .tsx"| TSAnalyzer["TreeSitterTSAnalyzer"] + Router -->|".php"| PHPAnalyzer["TreeSitterPHPAnalyzer"] + Router -->|".py"| PyAnalyzer["PythonASTAnalyzer"] + + CAnalyzer --> Output["Node and CallRelationship lists"] + CppAnalyzer --> Output + CSharpAnalyzer --> Output + JavaAnalyzer --> Output + JSAnalyzer --> Output + TSAnalyzer --> Output + PHPAnalyzer --> Output + PyAnalyzer --> Output + + Output --> Builder["Dependency Graph Builder"] + Output --> CallGraph["Call Graph Analyzer"] diff --git a/docs/diagrams/architecture/utils-2.mmd b/docs/diagrams/architecture/utils-2.mmd new file mode 100644 index 00000000..064dda9f --- /dev/null +++ b/docs/diagrams/architecture/utils-2.mmd @@ -0,0 +1,22 @@ +sequenceDiagram + participant Pipeline as "CLI Pipeline" + participant Logger as "CLILogger" + participant Tracker as "ProgressTracker" + participant ModBar as "ModuleProgressBar" + + Pipeline->>Logger: create_logger(verbose) + Pipeline->>Tracker: start_stage(1, "Dependency Analysis") + Pipeline->>Logger: info("Analyzing repository...") + Tracker-->>Pipeline: update_stage(progress) + Pipeline->>Tracker: complete_stage() + + Pipeline->>Tracker: start_stage(3, "Documentation Generation") + Pipeline->>ModBar: new ModuleProgressBar(total_modules) + loop "for each module" + Pipeline->>ModBar: update(module_name, cached) + Pipeline->>Logger: debug("module details") + end + Pipeline->>ModBar: finish() + Pipeline->>Tracker: complete_stage("Generation finished") + + Pipeline->>Logger: success("Documentation generated") diff --git a/docs/diagrams/architecture/utils-3.mmd b/docs/diagrams/architecture/utils-3.mmd new file mode 100644 index 00000000..9ef7a30b --- /dev/null +++ b/docs/diagrams/architecture/utils-3.mmd @@ -0,0 +1,11 @@ +flowchart TD + A["create_logger(verbose)"] --> B["ProgressTracker(total_stages=5)"] + B --> C["tracker.start_stage(1, 'Dependency Analysis')"] + C --> D["perform analysis; call tracker.update_stage()"] + D --> E["tracker.complete_stage()"] + E --> F["tracker.start_stage(3, 'Documentation Generation')"] + F --> G["ModuleProgressBar(total_modules)"] + G --> H["for each module: bar.update(name, cached)"] + H --> I["bar.finish()"] + I --> J["tracker.complete_stage()"] + J --> K["logger.success('Done')"] diff --git a/docs/diagrams/architecture/utils.mmd b/docs/diagrams/architecture/utils.mmd new file mode 100644 index 00000000..69a2014a --- /dev/null +++ b/docs/diagrams/architecture/utils.mmd @@ -0,0 +1,36 @@ +classDiagram + class CLILogger { + +bool verbose + +datetime start_time + +debug(message) void + +info(message) void + +success(message) void + +warning(message) void + +error(message) void + +step(message, step, total) void + +elapsed_time() str + } + + class ProgressTracker { + +int total_stages + +int current_stage + +float stage_progress + +float start_time + +bool verbose + +STAGE_WEIGHTS dict + +STAGE_NAMES dict + +start_stage(stage, description) void + +update_stage(progress, message) void + +complete_stage(message) void + +get_overall_progress() float + +get_eta() str + } + + class ModuleProgressBar { + +int total_modules + +int current_module + +bool verbose + +bar + +update(module_name, cached) void + +finish() void + } diff --git a/docs/getting-started/.gitignore b/docs/getting-started/.gitignore new file mode 100644 index 00000000..a5d01be5 --- /dev/null +++ b/docs/getting-started/.gitignore @@ -0,0 +1,7 @@ +# VoltAgent temp files +temp/ + +# JSON intermediate files (except schema/config) +*.json +!*-schema.json +!*-config.json diff --git a/docs/getting-started/first-steps.md b/docs/getting-started/first-steps.md new file mode 100644 index 00000000..192c1604 --- /dev/null +++ b/docs/getting-started/first-steps.md @@ -0,0 +1,95 @@ +# First Steps + +You've installed CodeWiki, configured your LLM credentials, and generated your first documentation set. Here's what to explore next. + +## 1. Inspect and Confirm Your Configuration + +Use `codewiki config show` to review everything CodeWiki currently knows about your setup — models, base URLs, token limits, temperature settings, and agent instructions. API keys are always masked (only the first/last few characters shown). + +```bash +codewiki config show +codewiki config show --json +``` + +Use `codewiki config validate` any time you change providers or suspect something is misconfigured. It checks the config file, verifies all three API keys are present, validates base URL formats, confirms models are set, and (unless `--quick` is passed) performs a live connectivity test against each configured provider. + +```bash +codewiki config validate +codewiki config validate --quick # Skip live API connectivity test +codewiki config validate --verbose # Step-by-step diagnostic output +``` + +## 2. Customize What Gets Documented + +The `generate` command accepts several options that narrow or reshape the analysis without touching your saved configuration: + +```bash +# Only analyze C# files, skip test projects +codewiki generate --include "*.cs" --exclude "*Tests*,*Specs*,test_*" + +# Focus documentation on specific modules/paths +codewiki generate --focus "src/core,src/api" --doc-type architecture + +# Add free-form custom instructions for the documentation agent +codewiki generate --instructions "Focus on public APIs and include usage examples" + +# Include additional source directories (e.g., vendored dependencies) +codewiki generate --additional-paths "vendor/packages,external/deps" +``` + +If you want these choices to become your **default** behavior for every future run (rather than one-off flags), persist them with: + +```bash +codewiki config agent --include "*.cs" --exclude "*Tests*,*Specs*" +codewiki config agent --doc-type architecture +codewiki config agent --instructions "Focus on public APIs and include usage examples" + +# Clear all saved agent instructions +codewiki config agent --clear +``` + +`--doc-type` accepts one of: `api`, `architecture`, `user-guide`, or `developer`. + +## 3. Tune Token Budgets and Depth for Large Repositories + +If your repository is very large or the LLM response is being truncated, adjust token and depth limits either per-run or persistently: + +```bash +# Per-run override +codewiki generate --max-tokens 32768 --max-token-per-module 40000 --max-token-per-leaf-module 20000 --max-depth 3 + +# Persist as defaults +codewiki config set --cluster-max-tokens 128000 --main-max-tokens 128000 \ + --max-token-per-module 40000 --max-token-per-leaf-module 20000 --max-depth 3 +``` + +`--max-depth` controls how many levels of hierarchical module decomposition are produced (default: 2). + +## 4. Explore the Git and GitHub Pages Workflow + +If you're documenting a Git-tracked project, CodeWiki can create a dedicated branch for the generated docs and prepare a GitHub Pages-ready static site: + +```bash +codewiki generate --create-branch --github-pages +``` + +`--create-branch` requires a clean working tree and creates a timestamped branch. `--github-pages` renders a self-contained `index.html` from the generated `module_tree.json` and `metadata.json`, suitable for publishing directly. + +For CI/CD pipelines where you don't want interactive prompts, add `--force` to overwrite existing documentation without prompting, and `--no-cache` to force a full regeneration. + +## 5. Explore the Generated Output Structure + +After a run, look inside your output directory (default `./docs`) for: + +- Individual Markdown files per analyzed module (leaves generated first, then parent overview pages) +- `module_tree.json` — the hierarchical module structure used for navigation +- `metadata.json` — job status and generation statistics +- `index.html` (only if `--github-pages` was used) + +If you used `--diagrams-output`, Mermaid diagrams extracted from the generated Markdown are also saved separately as `.mmd` files. + +## Where to Get Help + +- Run `codewiki --help`, `codewiki generate --help`, or `codewiki config --help` / `codewiki config set --help` for full flag references directly in your terminal — these are the most up-to-date source of truth for available options. +- For questions, feedback, or community discussion, join the OpenMSP Slack community: [https://www.openmsp.ai/](https://www.openmsp.ai/) ([join link](https://join.slack.com/t/openmsp/shared_invite/zt-36bl7mx0h-3~U2nFH6nqHqoTPXMaHEHA)). +- If you plan to contribute code or documentation improvements back to CodeWiki itself, continue on to the development section of this documentation. diff --git a/docs/getting-started/introduction.md b/docs/getting-started/introduction.md new file mode 100644 index 00000000..81f5512a --- /dev/null +++ b/docs/getting-started/introduction.md @@ -0,0 +1,85 @@ +# Introduction to CodeWiki + +## What is CodeWiki? + +**CodeWiki** is an AI-powered documentation generator for source code repositories. Point it at a codebase — locally via a command-line interface, or remotely via a web application — and it will: + +1. Analyze the repository's file structure and cross-file call relationships +2. Build a dependency graph of functions, classes, and modules +3. Cluster related components into meaningful, hierarchical modules +4. Use LLM-backed agents to write structured Markdown documentation for each module +5. Produce navigable output, including an optional static HTML viewer suitable for GitHub Pages + +CodeWiki is built as a Python package (`codewiki`, requires Python `>=3.12`) and supports analysis of **Python, Java, JavaScript, TypeScript, C, C++, C#, and PHP** source files through dedicated tree-sitter based language analyzers. + +## Key Features + +- **Two entry points, one engine** — a `codewiki` CLI for local repositories and a FastAPI web application for submitting GitHub repository URLs. Both drive the same backend documentation pipeline. +- **Multi-language dependency analysis** — tree-sitter powered analyzers extract call graphs and structural relationships across eight languages. +- **LLM-driven module clustering** — instead of a flat file listing, components are grouped into meaningful modules (e.g., "Auth Module", "API Module") by an LLM, then documented leaf-first before parent overviews are generated. +- **Per-provider LLM configuration** — separate model, API key, base URL, token limit, and temperature settings for the *cluster*, *main*, and *fallback* LLM roles, so you can mix providers (e.g., a cheaper model for clustering, a stronger model for generation). +- **Secure credential storage** — API keys are stored in the OS keyring (macOS Keychain, Windows Credential Manager, Linux Secret Service), never in plaintext configuration files. +- **Git-aware workflow** — the CLI can validate a clean working tree, create a timestamped documentation branch, and prepare it for a pull request. +- **Static HTML output** — an optional, self-contained `index.html` viewer can be generated for GitHub Pages deployment. +- **Caching for the web app** — the FastAPI frontend caches generated documentation by repository URL (with configurable expiry) to avoid redundant regeneration. +- **Multi-path analysis** — repositories whose source is split across multiple root directories (e.g., a monorepo with `main/`, `deps/`, `vendor/`) can be analyzed as a single unified documentation set via `additional_source_paths`. + +## Target Audience + +CodeWiki is for: + +- **Individual developers** who want up-to-date architecture documentation for a repository without manually writing it. +- **Teams and maintainers** who want a repeatable, CI-friendly way to regenerate documentation as code evolves (`--force` flag for non-interactive use). +- **Open-source project maintainers** who want a GitHub Pages-ready documentation site generated directly from their codebase. +- **Platform operators** who want to offer documentation-as-a-service through a hosted web application backed by GitHub repository submissions. + +## High-Level Architecture + +```mermaid +flowchart TD + User["Developer or Web User"] --> Entry{{"Choose entry point?"}} + Entry -->|CLI| CLI["CLI Core"] + Entry -->|Web| Frontend["Frontend Core"] + + CLI --> RuntimeConfig["Config Core"] + Frontend --> RuntimeConfig + + CLI --> GitOps["Git Integration"] + Frontend --> RepoProcessor["GitHub Repository Processor"] + + GitOps --> Source["Source Repository"] + RepoProcessor --> Source + + RuntimeConfig --> Generator["DocumentationGenerator"] + Source --> Generator + + Generator --> Analysis["Dependency Analysis"] + Analysis --> Parsers["Language Parsers"] + Parsers --> Graph["Dependency Graph"] + + Graph --> Clustering["Module Clustering"] + Clustering --> Agents["LLM Agent Orchestration"] + Agents --> Docs["Markdown Documentation"] + + Docs --> Metadata["Module Tree and Metadata"] + Metadata --> HTML["Optional HTML Viewer"] + Docs --> Output["Generated Documentation Output"] + HTML --> Output +``` + +## Core Modules at a Glance + +| Module | Purpose | +|---|---| +| **CLI Core** | Command-line orchestration: local config, Git integration, terminal progress, static HTML generation. | +| **Backend Core** | Repository analysis, dependency-graph construction, module clustering, LLM agent orchestration, documentation generation. | +| **Frontend Core** | FastAPI web application: repository submission, background job processing, caching, doc serving. | +| **Config Core** | Shared runtime `Config` model for source paths, output locations, provider settings, token limits, and agent instructions. | + +> **Note:** CodeWiki itself is a documentation-generation *tool* — the "modules" above describe CodeWiki's own internal architecture, not application features you configure for an end-user product. When you run CodeWiki against your own repository, it analyzes *your* code and produces documentation about *your* project. + +## Where to Go Next + +- [Prerequisites](prerequisites.md) — required software, versions, and environment setup before installing CodeWiki. +- [Quick Start](quick-start.md) — a five-minute walkthrough of installing CodeWiki and generating your first documentation set. +- [First Steps](first-steps.md) — what to explore after your first successful run. diff --git a/docs/getting-started/prerequisites.md b/docs/getting-started/prerequisites.md new file mode 100644 index 00000000..4869d69e --- /dev/null +++ b/docs/getting-started/prerequisites.md @@ -0,0 +1,71 @@ +# Prerequisites + +Before installing CodeWiki, make sure your environment meets the following requirements. + +## Required Software + +| Software | Minimum Version | Why It's Needed | +|---|---|---| +| Python | `>=3.12` | CodeWiki is a Python package (`pyproject.toml` requires `requires-python = ">=3.12"`). | +| pip | Bundled with Python 3.12 | Used to install the `codewiki` package and its dependencies from `requirements.txt` / `pyproject.toml`. | +| Git | Any recent version | Required for repository cloning (web app), git-branch workflows (`--create-branch`), and commit-hash/branch detection. | +| Node.js | `>=14.0.0` | Declared as a build requirement in `pyproject.toml` (`[external] build-requires`) — used by `mermaid-py` to validate Mermaid diagrams embedded in generated documentation. | +| Docker & Docker Compose | Recent version | Optional — only needed if you want to run the web application via `docker/docker-compose.yml` instead of running it directly with Python. | + +## System Requirements + +- **OS**: Linux, macOS, or Windows. Keyring-backed credential storage uses the OS-native secret store: macOS Keychain, Windows Credential Manager, or a Linux Secret Service implementation (e.g., GNOME Keyring). If no keyring backend is available, CodeWiki degrades gracefully (API keys must then be supplied another way each run). +- **Disk space**: Sufficient space to clone the target repository plus generated output (`output/cache`, `output/temp`, `output/docs`, `output/dependency_graphs`). +- **Network access**: Outbound HTTPS access to your configured LLM provider endpoint(s) (e.g., OpenAI, Anthropic, or an OpenAI-compatible base URL) is required at generation time. + +## Account / Access Requirements + +CodeWiki does not include a bundled LLM — you must bring your own provider credentials for **each** of the three configurable model roles: + +| Role | Purpose | +|---|---| +| **Cluster model** | Groups discovered code components into a hierarchical module tree. Recommended: a strong/top-tier model, since clustering quality drives overall documentation structure. | +| **Main model** | Writes the actual Markdown documentation for each module. | +| **Fallback model** | Used automatically if the main model call fails or is rate-limited. | + +Each role has its own API key, base URL, API version (for Anthropic-style APIs), max-token setting, and temperature setting — they do not need to be the same provider. + +## Environment Variables + +CodeWiki's CLI does **not** rely on ad-hoc environment variables for normal operation — persistent configuration is stored in `~/.codewiki/config.json` (non-secret settings) and the OS keyring (API keys), managed via `codewiki config set`. + +For the **test/diagnostic scripts** included in the repository (e.g., `test_clustering_*.py`), the following environment variables (or a local `.env.local` file) may be read directly: + +| Variable | Purpose | +|---|---| +| `OPENAI_API_KEY` | API key for OpenAI-compatible providers used by ad-hoc clustering test scripts. | +| `ANTHROPIC_API_KEY` | API key for Anthropic providers used by ad-hoc clustering test scripts. | +| `MAIN_API_KEY` / `CLUSTER_API_KEY` / `FALLBACK_API_KEY` | Per-role overrides used by the same test scripts. | + +For the **Docker Compose** deployment of the web application, environment values are loaded from an `.env` file referenced in `docker/docker-compose.yml` (`env_file: ../.env`), and `APP_PORT` controls the host port mapping (defaults to `8000`). + +## Verification Commands + +Run these commands to confirm your environment is ready before installing CodeWiki: + +```bash +# Check Python version (must be 3.12 or higher) +python3 --version + +# Check pip is available +pip3 --version + +# Check Git is installed +git --version + +# Check Node.js is installed (required by mermaid-py for diagram validation) +node --version + +# Optional: check Docker and Docker Compose (only needed for the web app container) +docker --version +docker compose version +``` + +> **Note:** If `python3 --version` reports an older version than 3.12, install a compatible Python before proceeding — CodeWiki's `pyproject.toml` will refuse to install otherwise. + +Once these checks pass, continue to the [Quick Start](quick-start.md) guide to install and run CodeWiki for the first time. diff --git a/docs/getting-started/quick-start.md b/docs/getting-started/quick-start.md new file mode 100644 index 00000000..330b15fc --- /dev/null +++ b/docs/getting-started/quick-start.md @@ -0,0 +1,114 @@ +# Quick Start + +This guide gets you from zero to a generated documentation set in about five minutes, using the `codewiki` CLI against a local repository. + +> If you'd rather run CodeWiki as a hosted web service (submit a GitHub URL, poll for job status, view cached results), see the Docker-based setup mentioned at the end of this guide instead. + +## Step 1: Install CodeWiki + +Clone the repository and install the package (editable install is convenient for exploring the source): + +```bash +git clone https://github.com/flamingo-stack/CodeWiki.git +cd CodeWiki +pip install -e . +``` + +This registers the `codewiki` console command, defined in `pyproject.toml` as: + +```text +[project.scripts] +codewiki = "codewiki.cli.main:cli" +``` + +Verify the install: + +```bash +codewiki --version +codewiki version +``` + +## Step 2: Configure Your LLM Credentials + +CodeWiki needs API credentials for at least a **main model** and a **cluster model** (a fallback model is optional but recommended). Credentials are stored securely in your OS keyring; non-secret settings go to `~/.codewiki/config.json`. + +```bash +codewiki config set \ + --cluster-api-key "sk-your-cluster-provider-key" \ + --main-api-key "sk-your-main-provider-key" \ + --cluster-model "your-cluster-model-name" \ + --main-model "your-main-model-name" \ + --cluster-base-url "https://api.your-provider.com/v1" \ + --main-base-url "https://api.your-provider.com/v1" +``` + +> **Note:** Replace the model names, base URLs, and API keys with values for your actual LLM provider. CodeWiki does not ship with default credentials — you must supply your own. + +Confirm the configuration was saved and is complete: + +```bash +codewiki config validate +``` + +## Step 3: Generate Documentation for a Repository + +Navigate to any Git repository you want to document, then run: + +```bash +cd /path/to/your/project +codewiki generate +``` + +By default, output is written to `./docs`. The CLI runs through four staged checks and then the documentation pipeline itself: + +```text +Validating configuration... +Validating repository... +Analyzing dependencies... +Generating documentation... +``` + +## Expected Output + +After a successful run, you should see a `docs/` directory in your project containing: + +- Markdown files for each analyzed module (leaf modules first, then parent overview pages) +- A `module_tree.json` describing the hierarchical module structure +- A `metadata.json` describing job statistics and status + +## Example: Verbose Run with Custom Output + +```bash +codewiki generate --output ./generated-docs --verbose +``` + +Add `--verbose` any time you want detailed stage-by-stage progress and debug information printed to your terminal. + +## Example: Generate a GitHub Pages Site + +```bash +codewiki generate --github-pages --create-branch +``` + +This additionally renders a self-contained `index.html` viewer (from the generated `module_tree.json` and `metadata.json`) and creates a timestamped Git branch for the documentation changes, ready to push and open a pull request. + +## Running the Web Application Instead + +If you prefer the hosted web workflow (submit a GitHub repo URL through a browser, track job status, and view cached results), you can run the FastAPI app directly: + +```bash +python codewiki/run_web_app.py +``` + +Or via Docker Compose: + +```bash +cd docker +docker compose up --build +``` + +The web app listens on port `8000` by default (configurable via the `APP_PORT` environment variable read by `docker/docker-compose.yml`). + +## Next Steps + +Once you've generated your first documentation set, continue to [First Steps](first-steps.md) to learn about customizing what gets documented, exploring the CLI's other options, and where to find help. diff --git a/docs/reference/architecture/.gitignore b/docs/reference/architecture/.gitignore new file mode 100644 index 00000000..b4f4c7a6 --- /dev/null +++ b/docs/reference/architecture/.gitignore @@ -0,0 +1,8 @@ +# CodeWiki temp files (dependency graphs can be 7GB+) +temp/ +dependency_graphs/ + +# JSON intermediate files (except schema/config) +*.json +!*-schema.json +!*-config.json diff --git a/docs/reference/architecture/README.md b/docs/reference/architecture/README.md new file mode 100644 index 00000000..6d9394e0 --- /dev/null +++ b/docs/reference/architecture/README.md @@ -0,0 +1,123 @@ +# CodeWiki Overview + +CodeWiki is an AI-powered documentation generator for source repositories. It analyzes a codebase, builds dependency relationships, groups components into meaningful modules, and uses LLM-backed agents to produce structured Markdown documentation with navigation metadata and optional static HTML output. + +The repository supports two primary entry points: + +- **CLI workflow** for local repository documentation, configuration management, Git workflows, progress reporting, and static-site generation. +- **Web workflow** for submitting GitHub repositories, queueing asynchronous generation jobs, caching results, and serving generated documentation. + +## End-to-End Architecture + +```mermaid +flowchart TD + User["Developer or Web User"] --> Entry{{"Choose entry point?"}} + Entry -->|CLI| CLI["CLI Core"] + Entry -->|Web| Frontend["Frontend Core"] + + CLI --> RuntimeConfig["Config Core"] + Frontend --> RuntimeConfig + + CLI --> GitOps["Git Integration"] + Frontend --> RepoProcessor["GitHub Repository Processor"] + + GitOps --> Source["Source Repository"] + RepoProcessor --> Source + + RuntimeConfig --> Generator["DocumentationGenerator"] + Source --> Generator + + Generator --> Analysis["Dependency Analysis"] + Analysis --> Parsers["Language Parsers"] + Parsers --> Graph["Dependency Graph"] + + Graph --> Clustering["Module Clustering"] + Clustering --> Agents["LLM Agent Orchestration"] + Agents --> Docs["Markdown Documentation"] + + Docs --> Metadata["Module Tree and Metadata"] + Metadata --> HTML["Optional HTML Viewer"] + Docs --> Output["Generated Documentation Output"] + HTML --> Output +``` + +## Documentation Generation Flow + +```mermaid +sequenceDiagram + participant Caller + participant Config as "Config Core" + participant Generator as "DocumentationGenerator" + participant Analyzer as "Dependency Analyzer" + participant Agents as "AgentOrchestrator" + participant Output as "Documentation Output" + + Caller->>Config: Build validated runtime configuration + Caller->>Generator: Start documentation run + Generator->>Analyzer: Analyze repository files and calls + Analyzer-->>Generator: Dependency graph and leaf components + Generator->>Generator: Cluster components into modules + Generator->>Agents: Generate leaf module documentation + Agents-->>Generator: Generated module content + Generator->>Output: Write Markdown, module tree, and metadata + Generator-->>Caller: Generation complete +``` + +## Core Modules + +| Module | Purpose | Source | +|---|---|---| +| [CLI Core](cli-core.md) | Command-line orchestration, local configuration, Git integration, terminal progress, and static HTML generation. | [`codewiki/cli`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/cli) | +| [Backend Core](backend-core.md) | Repository analysis, dependency-graph construction, module clustering, LLM agent orchestration, and documentation generation. | [`codewiki/src/be`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/src/be) | +| [Frontend Core](frontend-core.md) | FastAPI web application, repository submission, background job processing, caching, and documentation serving. | [`codewiki/src/fe`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/src/fe) | +| [Config Core](config-core.md) | Shared runtime `Config` model for source paths, output locations, provider settings, token limits, and agent instructions. | [`codewiki/src`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/src) | +| [Test Multi Path](test-multi-path.md) | Fixtures and executable checks for analysis across multiple source roots. | [`test-multi-path`](https://github.com/flamingo-stack/CodeWiki/tree/main/test-multi-path) | +| [Test Clustering](test-clustering.md) | Diagnostic and validation scripts for LLM-based module clustering and component-ID normalization. | [`test_clustering`](https://github.com/flamingo-stack/CodeWiki/tree/main/test_clustering) | + +## Module Relationships + +```mermaid +flowchart LR + Config["Config Core"] --> CLI["CLI Core"] + Config --> Frontend["Frontend Core"] + Config --> Backend["Backend Core"] + + CLI --> Backend + Frontend --> Backend + + Backend --> AgentTools["Agent Tools Core"] + Backend --> Analyzer["Dependency Analyzer Core"] + Backend --> Language["Tree-sitter Analyzers"] + Backend --> LLM["LLM Services"] + + TestPaths["Test Multi Path"] --> Config + TestPaths --> Analyzer + + TestCluster["Test Clustering"] --> Config + TestCluster --> Backend +``` + +## Backend Documentation References + +The Backend Core module is decomposed into focused subsystems: + +- [Agent Tools Core](backend-core/agent-tools-core/agent-tools-core.md) — controlled repository inspection and documentation editing tools. +- [Dependency Analyzer Core](backend-core/dependency-analyzer-core/dependency-analyzer-core.md) — file discovery, AST parsing, call analysis, and graph construction. +- [Tree Sitter Analyzers](backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md) — language support for C, C++, C#, Java, JavaScript, TypeScript, PHP, and Python. +- [Dependency Analyzer Models](backend-core/dependency-analyzer-models/dependency-analyzer-models.md) — repository, node, relationship, and analysis-result contracts. +- [Documentation Generator](backend-core/documentation-generator/documentation-generator.md) — top-level pipeline coordination and documentation output. +- [LLM Services](backend-core/llm-services/llm-services.md) — LLM model selection, request counting, and fallback handling. +- [Logging Config](backend-core/logging-config/logging-config.md) — shared colorized logging support. + +## CLI Documentation References + +- [Generation](cli-core/generation/generation.md) — CLI adapter for staged documentation generation. +- [Configuration](cli-core/configuration/configuration.md) — persisted settings, keyring-backed credentials, and agent instructions. +- [Job Models](cli-core/job_models/job_models.md) — generation job status, statistics, and LLM configuration models. +- [Git Integration](cli-core/git_integration/git_integration.md) — clean-tree checks, branch creation, commits, and remote URL handling. +- [HTML Generation](cli-core/html_generation/html_generation.md) — static documentation viewer generation. +- [Utils](cli-core/utils/utils.md) — terminal logging and progress tracking. + +## Summary + +CodeWiki separates user-facing workflows from its reusable generation engine. CLI and web layers construct a shared runtime configuration and delegate to Backend Core, which transforms source code into dependency-aware, module-oriented documentation. Test modules provide targeted coverage for multi-root analysis and the LLM-driven clustering stage. \ No newline at end of file diff --git a/docs/reference/architecture/backend-core/agent-tools-core/agent-tools-core.md b/docs/reference/architecture/backend-core/agent-tools-core/agent-tools-core.md new file mode 100644 index 00000000..0301cf95 --- /dev/null +++ b/docs/reference/architecture/backend-core/agent-tools-core/agent-tools-core.md @@ -0,0 +1,123 @@ +# Agent Tools Core + +## Purpose + +The Agent Tools Core module provides the foundational toolset used by the CodeWiki documentation-generation agent to interact with the filesystem during automated documentation authoring. It defines: + +- **`CodeWikiDeps`** — the shared dependency container passed to every agent tool invocation, carrying paths, the component registry, module-tree position, and LLM configuration. +- **`EditTool`** (and its helpers `Filemap`, `WindowExpander`) — a filesystem editor that lets the documentation agent view source/repository files and view, create, and edit documentation files, mirroring the SWE-agent `str_replace_editor` tool contract used by Anthropic-compatible agent frameworks. + +This module is a direct child of [Backend Core](../backend-core.md) and is consumed by the [Agent Orchestrator](../backend-core.md) during the documentation generation pipeline orchestrated for each module in the [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md) output tree. + +## Architecture Overview + +```mermaid +flowchart TD + Orchestrator["AgentOrchestrator"] -->|"builds"| Deps["CodeWikiDeps"] + Orchestrator -->|"invokes tool with"| ToolCall["str_replace_editor Tool Call"] + ToolCall -->|"uses ctx.deps"| Deps + ToolCall -->|"instantiates"| EditTool["EditTool"] + EditTool -->|"large .py view"| Filemap["Filemap"] + EditTool -->|"view/edit windows"| WindowExpander["WindowExpander"] + EditTool -->|"reads"| RepoFS[("Repository Files")] + EditTool -->|"reads/writes"| DocsFS[("Generated Docs Files")] + ToolCall -->|"post-edit .md check"| MermaidValidator["validate_mermaid_diagrams"] + Deps -->|"references"| Config["Config"] + Deps -->|"references"| Node["Node (component registry)"] +``` + +The module bridges the AI agent's tool-calling layer (built on `pydantic_ai`) and the local filesystem, enforcing safety constraints (absolute paths only, read-only access to the source repository, write access limited to the docs output directory) so that the agent can safely author markdown documentation while inspecting source code. + +## Core Components + +### `CodeWikiDeps` — Agent Dependency Container + +`CodeWikiDeps` (in `codewiki/src/be/agent_tools/deps.py`) is a `dataclass` that bundles everything a tool call needs to operate correctly within the current documentation generation run: + +| Field | Purpose | +|---|---| +| `absolute_docs_path` | Root directory where generated documentation is written | +| `absolute_repo_path` | Root directory of the analyzed source repository (read-only) | +| `registry` | Shared mutable dict used for cross-call state (e.g., file edit history, mermaid validation retry counters) | +| `components` | Mapping of component identifiers to `Node` objects from the dependency graph | +| `path_to_current_module` / `current_module_name` | Position of the module currently being documented within the module tree | +| `module_tree` | The full hierarchical module tree being documented | +| `max_depth` / `current_depth` | Recursion bounds for nested sub-module documentation generation | +| `config` | LLM configuration (`Config`) used to drive model behavior | +| `custom_instructions` | Optional user-supplied instructions injected into the agent's prompt | + +This container is instantiated once per documentation job/module by the orchestrator and passed as `ctx.deps` into every `pydantic_ai.RunContext` during a tool call, giving each tool a consistent view of the run's state without global variables. + +### `EditTool`, `Filemap`, and `WindowExpander` — Filesystem Editor + +The `str_replace_editor` tool (in `codewiki/src/be/agent_tools/str_replace_editor.py`) is adapted from the SWE-agent reference implementation and exposes five commands to the agent: + +- **`view`** — Displays a file (with `cat -n`-style line numbers) or lists a directory up to two levels deep. +- **`create`** — Creates a new file, auto-creating any missing parent directories (used for hierarchical docs output). +- **`str_replace`** — Replaces a unique occurrence of `old_str` with `new_str` in a file; rejects ambiguous or missing matches. +- **`insert`** — Inserts text at a specific line number. +- **`undo_edit`** — Reverts the most recent edit, using a per-file history stack persisted in `CodeWikiDeps.registry`. + +**Safety model:** The tool enforces `working_dir` semantics — when `working_dir="repo"`, only `view` is permitted (the source repository is never mutated); when `working_dir="docs"`, all commands are available against the documentation output tree. All paths must resolve to absolute paths under the appropriate root, and a leading-slash stripping fix prevents `Path` composition bugs where an absolute-looking relative path could escape the intended docs root. + +```mermaid +sequenceDiagram + participant Agent + participant Tool as "str_replace_editor" + participant Edit as "EditTool" + participant FS as "Filesystem" + participant Validator as "Mermaid Validator" + + Agent->>Tool: command="create", working_dir="docs", path="module.md" + Tool->>Tool: resolve absolute_path under docs root + Tool->>Edit: EditTool(registry, docs_path) + Edit->>Edit: validate_path(command, path) + Edit->>FS: create parent dirs + write_file + FS-->>Edit: written + Edit-->>Tool: success log + Tool->>Validator: validate_mermaid_diagrams(path) + Validator-->>Tool: validation result + Tool-->>Agent: combined result string +``` + +#### Supporting Helper Classes + +- **`Filemap`**: Uses `tree-sitter` to parse Python source and elide long function bodies (`>= 5` lines) when a `.py` file exceeds the response length limit, producing a condensed "filemap" view so the agent can navigate large files without exhausting context. This is currently gated behind the `USE_FILEMAP` flag. +- **`WindowExpander`**: Expands a requested `view_range` or edit snippet window outward to natural code boundaries (blank lines, `def`/`class`/decorator lines for Python) so that partial views don't cut a function or class definition in half. Expansion size is controlled by `MAX_WINDOW_EXPANSION_VIEW` / `MAX_WINDOW_EXPANSION_EDIT_CONFIRM` (both set to `0` by default, effectively disabling automatic expansion in the current configuration). + +### Mermaid Validation Hook + +After any non-`view` command that touches a `.md` file under `working_dir="docs"`, the tool asynchronously calls `validate_mermaid_diagrams` to catch malformed Mermaid diagrams as soon as the agent writes them. A per-file retry counter is stored in `ctx.deps.registry` (`mermaid_attempts:{path}`) to cap validation retries at `MAX_MERMAID_ATTEMPTS = 3`, preventing infinite fix-and-retry loops when the agent cannot resolve a diagram syntax error. + +```mermaid +flowchart LR + Edit["Edit .md file"] --> Check{{"path ends with .md?"}} + Check -->|"no"| Done["Return result"] + Check -->|"yes"| Attempts["Read mermaid_attempts counter"] + Attempts --> Limit{{"attempts >= MAX?"}} + Limit -->|"yes"| Skip["Skip validation, warn agent"] + Limit -->|"no"| Validate["validate_mermaid_diagrams"] + Validate --> HasError{{"errors found?"}} + HasError -->|"yes"| Increment["Increment counter in registry"] + HasError -->|"no"| Reset["Reset counter to 0"] + Increment --> Done + Reset --> Done + Skip --> Done +``` + +### Optional Linting Integration + +The module includes `flake8`-based linting utilities (`Flake8Error`, `format_flake8_output`, `flake8`) that can compare pre- and post-edit lint output for Python files and surface newly-introduced errors to the agent after a `str_replace` edit. This is gated by the `USE_LINTER` flag and is primarily relevant when the agent edits Python source rather than markdown documentation. + +## Integration with the Wider System + +- **[Backend Core](../backend-core.md)**: The parent module's `AgentOrchestrator` constructs `CodeWikiDeps` for each documentation job and registers the `str_replace_editor_tool` (a `pydantic_ai.Tool`) with the agent, along with the component registry produced by the [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md) and its supporting [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) (`Node`, `Repository`, etc.). +- **Configuration**: `CodeWikiDeps.config` is populated from the top-level `Config` object (see [Config Core](../../config-core.md)), which supplies LLM provider settings used throughout the generation pipeline. +- **Documentation Output**: Every markdown file produced by the [Documentation Generator](../documentation-generator/documentation-generator.md) subsystem is ultimately written to disk via `EditTool.create_file` / `EditTool.str_replace`, making this module the final write path for all agent-authored documentation. + +## Key Design Considerations + +1. **Read-only source, writable docs**: The `working_dir` parameter is the sole gate that separates safe read-only inspection of the analyzed repository from the mutable documentation workspace, preventing the agent from accidentally modifying source code. +2. **Absolute path enforcement**: All paths passed to `EditTool` must be absolute, with an explicit fix to prevent `pathlib.Path`'s "absolute path discards base" behavior from silently escaping the intended root directory. +3. **Bounded retries for generated content**: Both the mermaid-validation retry counter and the lint-comparison logic are designed to give the agent actionable feedback without entering unbounded correction loops. +4. **Stateful history via registry**: Rather than relying on in-process instance state (which would not survive across tool calls within the same agent run), file edit history and validation counters are persisted in the shared `CodeWikiDeps.registry` dict, keyed by file path. diff --git a/docs/reference/architecture/backend-core/backend-core.md b/docs/reference/architecture/backend-core/backend-core.md new file mode 100644 index 00000000..3d844e7f --- /dev/null +++ b/docs/reference/architecture/backend-core/backend-core.md @@ -0,0 +1,98 @@ +# Backend Core + +Backend Core is CodeWiki’s server-side documentation pipeline. Located in `codewiki/src/be`, it analyzes source repositories, constructs dependency graphs, organizes components into modules, invokes LLM-backed agents to write documentation, and provides the filesystem tools and logging needed to run those workflows safely. + +Its primary entry point is `DocumentationGenerator`, which coordinates repository analysis and hierarchical documentation output. `AgentOrchestrator` manages agent-driven generation for individual modules, while `CountingFallbackModel` provides resilient LLM execution. + +## Architecture + +```mermaid +flowchart TD + Input["Source Repository"] --> Generator["DocumentationGenerator"] + Generator --> GraphBuilder["DependencyGraphBuilder"] + GraphBuilder --> Parser["DependencyParser"] + Parser --> Analysis["AnalysisService"] + Analysis --> RepoAnalyzer["RepoAnalyzer"] + Analysis --> CallAnalyzer["CallGraphAnalyzer"] + CallAnalyzer --> LanguageAnalyzers["Tree-sitter and Python Analyzers"] + LanguageAnalyzers --> Models["Dependency Analyzer Models"] + + GraphBuilder --> Components["Component Dependency Graph"] + Components --> Generator + + Generator --> Orchestrator["AgentOrchestrator"] + Orchestrator --> AgentTools["Agent Tools Core"] + Orchestrator --> LLM["CountingFallbackModel"] + LLM --> Provider["Configured LLM Provider"] + + AgentTools --> Docs["Generated Markdown Documentation"] + Generator --> Docs +``` + +## Documentation Generation Flow + +```mermaid +sequenceDiagram + participant Caller + participant Generator as "DocumentationGenerator" + participant Builder as "DependencyGraphBuilder" + participant Agent as "AgentOrchestrator" + participant Tools as "Agent Tools" + participant Output as "Documentation Files" + + Caller->>Generator: run() + Generator->>Builder: build_dependency_graph() + Builder-->>Generator: components and leaf nodes + Generator->>Generator: cluster modules and order leaves first + Generator->>Agent: process leaf module + Agent->>Tools: inspect repository and write docs + Tools-->>Agent: tool results + Agent-->>Generator: module documentation complete + Generator->>Output: write parent overviews and metadata + Generator-->>Caller: documentation complete +``` + +## Core Components + +| Component | Responsibility | +|---|---| +| `AgentOrchestrator` | Creates and runs documentation-writing agents for module-level generation. | +| `DocumentationGenerator` | Coordinates graph building, module clustering, leaf-first generation, parent overviews, and metadata output. | +| `AnalysisService` | Orchestrates repository structure analysis and call-graph extraction. | +| `RepoAnalyzer` | Discovers repository files and builds filtered file-tree representations. | +| `CallGraphAnalyzer` | Routes source files to language analyzers and aggregates call relationships. | +| `DependencyParser` | Converts analysis results into namespaced dependency-graph components. | +| `DependencyGraphBuilder` | Builds, validates, filters, and persists the repository dependency graph. | +| `CountingFallbackModel` | Wraps LLM models with request counting and automatic fallback behavior. | +| `CodeWikiDeps` | Carries shared run context and configuration into agent tools. | +| `EditTool` | Provides controlled repository viewing and documentation-file editing for agents. | +| `ColoredFormatter` | Produces readable, colorized backend console logs. | + +## Backend Subsystems + +- [Agent Tools Core](agent-tools-core/agent-tools-core.md) — Safe filesystem interaction through `CodeWikiDeps`, `EditTool`, `Filemap`, and `WindowExpander`. +- [Dependency Analyzer Core](dependency-analyzer-core/dependency-analyzer-core.md) — End-to-end repository analysis and dependency-graph construction. +- [Tree Sitter Analyzers](tree-sitter-analyzers/tree-sitter-analyzers.md) — Language-specific parsing for C, C++, C#, Java, JavaScript, TypeScript, PHP, and Python. +- [Dependency Analyzer Models](dependency-analyzer-models/dependency-analyzer-models.md) — Shared `Node`, `CallRelationship`, `Repository`, and analysis-result contracts. +- [Documentation Generator](documentation-generator/documentation-generator.md) — The top-level documentation generation workflow. +- [LLM Services](llm-services/llm-services.md) — Model factories, fallback handling, token configuration, and direct LLM calls. +- [Logging Config](logging-config/logging-config.md) — Shared colorized logging utilities. + +## Component Relationships + +```mermaid +flowchart LR + Models["Dependency Analyzer Models"] --> Analyzer["Dependency Analyzer Core"] + Language["Tree-sitter Analyzers"] --> Analyzer + Logging["Logging Config"] -.-> Analyzer + + Analyzer --> Generator["Documentation Generator"] + LLM["LLM Services"] --> Generator + Generator --> Orchestrator["AgentOrchestrator"] + Orchestrator --> Tools["Agent Tools Core"] + Tools --> Output["Markdown Documentation"] +``` + +## Source Location + +Backend Core source is maintained under [`codewiki/src/be`](https://github.com/flamingo-stack/CodeWiki/tree/main/codewiki/src/be). The module is consumed by the CLI and frontend layers to transform a repository into structured, navigable documentation. \ No newline at end of file diff --git a/docs/reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md b/docs/reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md new file mode 100644 index 00000000..ff7af08d --- /dev/null +++ b/docs/reference/architecture/backend-core/dependency-analyzer-core/analysis_pipeline.md @@ -0,0 +1,274 @@ +# Analysis Pipeline + +The Analysis Pipeline module is the orchestration layer responsible for turning a raw repository (either a cloned GitHub repository or a local folder) into a structured, language-aware call graph. It coordinates repository structure discovery, multi-language source parsing, and call-relationship resolution, producing the `AnalysisResult` data that downstream documentation and visualization components consume. + +This module sits inside the dependency analyzer subsystem of the backend and provides the primary entry point for "what does this codebase look like and how do its functions call each other?" + +## Purpose and Scope + +The Analysis Pipeline answers three questions for any supported repository: + +1. **What files exist?** — via structure analysis and pattern-based filtering. +2. **What functions/methods/classes exist in each file?** — via per-language AST parsing. +3. **How do those functions call each other?** — via cross-file/cross-language call relationship resolution. + +The pipeline does **not** implement language-specific parsing logic itself; it delegates that to the [Tree-Sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) module. It also does not define the data model shapes it operates on — those come from [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md). Instead, this module focuses purely on **orchestration**: cloning, filtering, dispatching, aggregating, and cleaning up. + +## Core Components + +| Component | Responsibility | +|---|---| +| `AnalysisService` | Top-level façade that orchestrates the full analysis workflow: clone → structure analysis → call graph analysis → README extraction → result assembly → cleanup | +| `RepoAnalyzer` | Builds a filtered file tree for a repository (or multiple repositories), applying include/exclude glob patterns and computing summary statistics | +| `CallGraphAnalyzer` | Central multi-language orchestrator that extracts code files from a file tree, dispatches each file to the correct language analyzer, aggregates functions/relationships, resolves call targets, deduplicates edges, and generates visualization data | + +## Architecture + +`AnalysisService` is the public API of this module. It composes a `RepoAnalyzer` (created per-call, since include/exclude patterns can vary) and a long-lived `CallGraphAnalyzer` instance. `CallGraphAnalyzer` never parses source code itself — it lazily imports the language-specific analyzer function (e.g. `analyze_python_file`, `analyze_javascript_file_treesitter`) from the [Tree-Sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) module at call time, keeping this module decoupled from heavy per-language parser dependencies. + +```mermaid +flowchart TD + Client["Caller (e.g. Documentation Generator)"] --> AS["AnalysisService"] + AS -->|"clone_repository()"| Clone["Repository Cloning (analysis.cloning)"] + AS -->|"analyze_repository_structure()"| RA["RepoAnalyzer"] + AS -->|"analyze_code_files()"| CGA["CallGraphAnalyzer"] + RA -->|"file_tree"| CGA + CGA -->|"extract_code_files()"| Extract["Code File Extraction"] + CGA -->|"dispatch by language"| Analyzers["Language Analyzers"] + Analyzers --> TS["Tree-Sitter Analyzers module"] + CGA -->|"Node / CallRelationship"| Models["Dependency Analyzer Models module"] + AS -->|"AnalysisResult"| Models + AS -->|"Repository"| Models +``` + +## Component Relationships + +```mermaid +classDiagram + class AnalysisService { + +call_graph_analyzer: CallGraphAnalyzer + +analyze_local_repository(repo_path, max_files, languages) dict + +analyze_repository_full(github_url, include_patterns, exclude_patterns) AnalysisResult + +analyze_repository_structure_only(github_url, include_patterns, exclude_patterns) dict + +cleanup_all() + -_clone_repository(github_url) str + -_analyze_structure(repo_dir, include, exclude) dict + -_analyze_call_graph(file_tree, repo_dir) dict + -_read_readme_file(repo_dir) str + -_cleanup_repository(temp_dir) + } + class RepoAnalyzer { + +include_patterns: list + +exclude_patterns: list + +analyze_repository_structure(repo_dir) dict + -_build_file_tree(repo_dir) dict + -_should_exclude_path(path, filename) bool + -_should_include_file(path, filename) bool + -_count_files(tree) int + -_calculate_size(tree) float + } + class CallGraphAnalyzer { + +functions: dict~str, Node~ + +call_relationships: list~CallRelationship~ + +analyze_code_files(code_files, base_dir) dict + +extract_code_files(file_tree) list + -_analyze_code_file(repo_dir, file_info) + -_resolve_call_relationships() + -_deduplicate_relationships() + -_generate_visualization_data() dict + +generate_llm_format() dict + -_select_most_connected_nodes(target_count) + } + class AnalysisResult { + +repository: Repository + +functions: list + +relationships: list + +file_tree: dict + +summary: dict + +visualization: dict + +readme_content: str + } + class Node { + +id: str + +name: str + +node_type: str + +file_path: str + +component_id: str + +docstring: str + +parameters: list + } + class CallRelationship { + +caller: str + +callee: str + +call_line: int + +is_resolved: bool + } + + AnalysisService --> RepoAnalyzer : uses + AnalysisService --> CallGraphAnalyzer : uses + AnalysisService --> AnalysisResult : produces + CallGraphAnalyzer --> Node : produces + CallGraphAnalyzer --> CallRelationship : produces + AnalysisResult --> Node : contains + AnalysisResult --> CallRelationship : contains +``` + +`AnalysisResult`, `Node`, `CallRelationship`, and `Repository` are defined in the [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) module and are reused here as the canonical output contract. + +## Workflow: Full Repository Analysis + +`analyze_repository_full` is the primary entry point used when a complete call graph (functions + relationships + visualization) is required for a GitHub repository. It is a sequential pipeline with cleanup guaranteed on both success and failure paths. + +```mermaid +sequenceDiagram + participant Caller + participant AS as AnalysisService + participant Clone as "Cloning Utility" + participant RA as RepoAnalyzer + participant CGA as CallGraphAnalyzer + participant FS as Filesystem + + Caller->>AS: analyze_repository_full(github_url) + AS->>Clone: clone_repository(github_url) + Clone-->>AS: temp_dir + AS->>AS: parse_github_url(github_url) + AS->>RA: analyze_repository_structure(temp_dir) + RA->>FS: walk directory tree + RA-->>AS: file_tree + summary + AS->>CGA: extract_code_files(file_tree) + CGA-->>AS: code_files + AS->>CGA: analyze_code_files(code_files, temp_dir) + CGA->>CGA: dispatch per language, parse each file + CGA->>CGA: resolve_call_relationships() + CGA->>CGA: deduplicate_relationships() + CGA->>CGA: generate_visualization_data() + CGA-->>AS: functions, relationships, call_graph, visualization + AS->>FS: read README file + AS->>AS: build AnalysisResult + AS->>Clone: cleanup_repository(temp_dir) + AS-->>Caller: AnalysisResult +``` + +If any step raises an exception, `AnalysisService` still cleans up the temporary clone directory before re-raising a `RuntimeError`, preventing orphaned checkouts from accumulating on disk. + +## Workflow: Structure-Only and Local Analysis + +Two lighter-weight workflows exist alongside the full analysis path: + +- **`analyze_repository_structure_only`** clones a GitHub repository and runs only `RepoAnalyzer`, skipping call graph generation entirely. This is used when only the file tree and summary statistics are needed (e.g. for quick previews). +- **`analyze_local_repository`** skips cloning altogether and operates directly on a local folder path, optionally filtering by `languages` and capping the number of files via `max_files`. It returns a simplified dict (`nodes`, `relationships`, `summary`) rather than a full `AnalysisResult`. + +```mermaid +flowchart LR + subgraph Full["analyze_repository_full"] + F1["Clone"] --> F2["Structure Analysis"] --> F3["Call Graph Analysis"] --> F4["README Read"] --> F5["AnalysisResult"] + end + subgraph StructureOnly["analyze_repository_structure_only"] + S1["Clone"] --> S2["Structure Analysis"] --> S3["Structure Dict"] + end + subgraph Local["analyze_local_repository"] + L1["Local Path"] --> L2["Structure Analysis"] --> L3["Filter by language / max_files"] --> L4["Call Graph Analysis"] --> L5["Simplified Dict"] + end +``` + +## Repository Structure Analysis (`RepoAnalyzer`) + +`RepoAnalyzer` walks a repository directory recursively and builds a nested `file_tree` dictionary, applying two categories of patterns: + +- **Include patterns**: if explicitly provided, they *replace* the defaults and restrict results to matching files only. +- **Exclude patterns**: if provided, they are *merged* with a built-in default ignore list (e.g. `.git`, `node_modules`). + +It defends against unsafe filesystem traversal by rejecting symlinks and any resolved path that escapes the repository root. + +`RepoAnalyzer` also supports **multi-repository analysis**: when given a list of paths instead of a single path, it builds a namespaced, merged tree — each repository's subtree is wrapped with a namespace derived from its folder name, and summary statistics (`total_files`, `total_size_kb`, `repositories`, `namespaces`) are aggregated across all inputs. + +```mermaid +flowchart TD + Input["repo_dir: str or list[str]"] --> Check{{"Single path?"}} + Check -->|"Yes"| Single["_build_file_tree(repo_dir)"] + Check -->|"No"| Multi["_analyze_multiple_repositories(repo_dirs)"] + Multi --> Loop["For each repo_dir: compute namespace, build tree"] + Loop --> Wrap["Wrap tree with namespace prefix"] + Wrap --> Merge["Merge into single root tree"] + Single --> Summary1["Compute total_files, total_size_kb"] + Merge --> Summary2["Aggregate totals across repositories"] + Summary1 --> Out["file_tree + summary"] + Summary2 --> Out +``` + +Path filtering combines glob matching (`fnmatch`) on both the full relative path and the bare filename, along with segment-level and prefix checks, so patterns like `node_modules`, `*.test.js`, or `build/` are all honored. + +## Call Graph Analysis (`CallGraphAnalyzer`) + +`CallGraphAnalyzer` is the multi-language orchestrator. Its responsibilities, in order: + +1. **Extraction** — `extract_code_files` walks the file tree and filters files by known code extensions (mapped to language names via a shared extension table). +2. **Dispatch** — `_analyze_code_file` routes each file to a private per-language method (`_analyze_python_file`, `_analyze_javascript_file`, `_analyze_typescript_file`, `_analyze_java_file`, `_analyze_csharp_file`, `_analyze_c_file`, `_analyze_cpp_file`, `_analyze_php_file`). Each of these lazily imports the corresponding analyzer function from the [Tree-Sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) module and normalizes the returned functions into the shared `Node` map keyed by function ID. +3. **Resolution** — `_resolve_call_relationships` builds a lookup table from function ID, bare name, `component_id`, and trailing method name, then attempts to match every call relationship's `callee` string against it, marking matches as `is_resolved`. +4. **Deduplication** — `_deduplicate_relationships` removes duplicate `(caller, callee)` edges, keeping only the first occurrence. +5. **Visualization** — `_generate_visualization_data` emits Cytoscape.js-compatible `elements` (nodes classified by type/language, edges limited to resolved relationships) plus a summary of node/edge counts. + +```mermaid +flowchart TD + FT["file_tree"] --> Extract["extract_code_files()"] + Extract --> CodeFiles["code_files list"] + CodeFiles --> Dispatch{{"Dispatch by language"}} + Dispatch -->|"python"| PyA["Python AST Analyzer"] + Dispatch -->|"javascript"| JsA["JS Tree-Sitter Analyzer"] + Dispatch -->|"typescript"| TsA["TS Tree-Sitter Analyzer"] + Dispatch -->|"java"| JavaA["Java Tree-Sitter Analyzer"] + Dispatch -->|"csharp"| CsA["C# Tree-Sitter Analyzer"] + Dispatch -->|"c"| CA["C Tree-Sitter Analyzer"] + Dispatch -->|"cpp"| CppA["C++ Tree-Sitter Analyzer"] + Dispatch -->|"php"| PhpA["PHP Tree-Sitter Analyzer"] + PyA --> Agg["functions dict, call_relationships list"] + JsA --> Agg + TsA --> Agg + JavaA --> Agg + CsA --> Agg + CA --> Agg + CppA --> Agg + PhpA --> Agg + Agg --> Resolve["_resolve_call_relationships()"] + Resolve --> Dedup["_deduplicate_relationships()"] + Dedup --> Viz["_generate_visualization_data()"] + Viz --> Result["functions, relationships, call_graph, visualization"] +``` + +Each per-language method (e.g. `_analyze_python_file`) follows the same normalization pattern: on success, returned `functions`/`relationships` are merged into `self.functions` (keyed by `func.id` or a synthesized `path:name` fallback) and appended to `self.call_relationships`; on failure, the exception is logged and analysis continues for remaining files, ensuring a single malformed file cannot abort the whole run. + +`_filter_supported_languages` (in `AnalysisService`) additionally narrows the extracted code files to a fixed set of supported languages (`python`, `javascript`, `typescript`, `java`, `csharp`, `c`, `cpp`, `php`, `go`, `rust`) before invoking the call graph analyzer, and reports how many files were skipped as unsupported. + +## Call Relationship Resolution Logic + +Resolving a raw `callee` string (as extracted from source, e.g. a bare function name, a dotted method reference, or a fully-qualified component ID) into an actual function node is the most delicate part of the pipeline, since different language analyzers may extract call targets in different formats. + +```mermaid +flowchart TD + Start["callee: str"] --> Direct{{"callee in func_lookup?"}} + Direct -->|"Yes"| Resolved["Mark resolved, rewrite callee to func_id"] + Direct -->|"No"| HasDot{{"Contains '.'?"}} + HasDot -->|"No"| Unresolved["Leave unresolved"] + HasDot -->|"Yes"| MethodName["Extract trailing segment after last '.'"] + MethodName --> MethodLookup{{"method_name in func_lookup?"}} + MethodLookup -->|"Yes"| Resolved + MethodLookup -->|"No"| Unresolved +``` + +The lookup table (`func_lookup`) is populated with multiple keys per function — its ID, bare name, `component_id`, and the last dotted segment of `component_id` — to maximize the chance of matching call sites extracted with varying levels of qualification across languages. + +## Cleanup and Resource Management + +`AnalysisService` tracks every temporary clone directory it creates in `_temp_directories`. Both `analyze_repository_full` and `analyze_repository_structure_only` clean up their temp directory in the `except` branch as well as on the success path, and `cleanup_all()` (also invoked from `__del__`) provides a final safety net to remove any directories that were never explicitly cleaned up — guarding against disk space leaks from crashed or aborted analysis runs. + +## Backward-Compatible Function API + +Two module-level functions, `analyze_repository` and `analyze_repository_structure_only`, wrap `AnalysisService` for callers that expect the older tuple-returning function signature (`(result, temp_dir)`), returning `None` in place of `temp_dir` since cleanup is now handled internally by the service. + +## Relationship to Other Modules + +- **[Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md)** — supplies the `AnalysisResult`, `Node`, `CallRelationship`, and `Repository` data structures produced by this pipeline. +- **[Tree-Sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md)** — supplies the actual per-language parsing functions (`analyze_python_file`, `analyze_javascript_file_treesitter`, etc.) that `CallGraphAnalyzer` dispatches to. +- **[Graph Construction](../graph_construction/graph_construction.md)** — a sibling module handling AST parsing and dependency graph assembly (`DependencyParser`, `DependencyGraphBuilder`); the Analysis Pipeline focuses on repository-level orchestration and call graph resolution, while graph construction focuses on assembling the broader dependency graph from parsed nodes. +- **[Dependency Analyzer Core](../dependency-analyzer-core.md)** — the parent module that groups the Analysis Pipeline together with Graph Construction as the two core analysis capabilities of the dependency analyzer subsystem. +- **[Documentation Generator](../../documentation-generator/documentation-generator.md)** — a consumer of `AnalysisResult` data produced by this pipeline, using it as input for generating documentation content. diff --git a/docs/reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md b/docs/reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md new file mode 100644 index 00000000..3e4335de --- /dev/null +++ b/docs/reference/architecture/backend-core/dependency-analyzer-core/dependency-analyzer-core.md @@ -0,0 +1,124 @@ +# Dependency Analyzer Core + +## Overview + +The Dependency Analyzer Core module is the central orchestration engine of the CodeWiki backend's static analysis pipeline. It coordinates the full lifecycle of turning a raw source repository into a structured, queryable **dependency graph** of code components (functions, classes, methods, interfaces, etc.) connected by call/usage relationships. + +This module answers three fundamental questions for any supported repository: + +1. **What files exist and which are relevant?** — repository structure discovery and filtering. +2. **What code components exist, and who calls whom?** — multi-language AST/call-graph extraction. +3. **How do these components form a navigable, deduplicated dependency graph?** — component construction, namespacing, and graph persistence. + +The resulting dependency graph is the primary input consumed by downstream stages of the CodeWiki backend, including documentation generation and LLM-driven summarization (see the [Documentation Generator](../documentation-generator/documentation-generator.md) module in `backend-core`). + +## Position in the System + +Dependency Analyzer Core lives inside the `backend-core` module, alongside several closely related sibling modules: + +- [Tree-sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) — language-specific parsers (Python, JavaScript, TypeScript, Java, C#, C, C++, PHP) invoked by this module's call graph orchestrator. +- [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) — the shared `Node`, `CallRelationship`, `Repository`, `AnalysisResult`, and `NodeSelection` data models produced and consumed by this module. +- [Logging Config](../logging-config/logging-config.md) — shared logging formatting utilities used across the analysis pipeline. +- [Documentation Generator](../documentation-generator/documentation-generator.md) — consumes the dependency graph produced here to generate documentation. +- [Agent Tools Core](../agent-tools-core/agent-tools-core.md) — provides file-editing and repository interaction tools used by the broader backend agent workflow. +- [LLM Services](../llm-services/llm-services.md) — provides model routing/fallback used elsewhere in `backend-core`. + +For the full picture of how these modules interrelate, see the parent [Backend Core](../backend-core.md) documentation. + +## Architecture + +The module is organized into two cooperating sub-systems: + +1. **Analysis Pipeline** — discovers repository structure and extracts functions/classes and their call relationships for a single repository path. +2. **Graph Construction** — wraps the analysis pipeline to build fully-qualified, namespaced `Node` components (supporting both single-repository and multi-repository/dependency scenarios), resolves cross-namespace edges, validates graph completeness, and persists the final dependency graph to disk. + +```mermaid +flowchart TD + subgraph AnalysisPipeline["Analysis Pipeline"] + direction TB + RepoAnalyzer["RepoAnalyzer"] -->|"file_tree"| AnalysisService["AnalysisService"] + CallGraphAnalyzer["CallGraphAnalyzer"] -->|"functions + relationships"| AnalysisService + end + + subgraph GraphConstruction["Graph Construction"] + direction TB + DependencyParser["DependencyParser"] -->|"components (Node)"| DependencyGraphBuilder["DependencyGraphBuilder"] + end + + Config["Config"] -.->|"repo paths, patterns"| DependencyGraphBuilder + DependencyGraphBuilder -->|"drives"| DependencyParser + DependencyParser -->|"uses"| AnalysisService + CallGraphAnalyzer -->|"delegates per-language"| TreeSitterAnalyzers["Tree-sitter Analyzers"] + AnalysisService -->|"produces"| Models["Node / CallRelationship / AnalysisResult"] + DependencyGraphBuilder -->|"writes"| GraphFile[("dependency_graph.json")] + + click TreeSitterAnalyzers "../tree-sitter-analyzers/tree-sitter-analyzers.md" + click Models "../dependency-analyzer-models/dependency-analyzer-models.md" +``` + +### High-Level Data Flow + +```mermaid +sequenceDiagram + participant Builder as "DependencyGraphBuilder" + participant Parser as "DependencyParser" + participant Service as "AnalysisService" + participant RepoA as "RepoAnalyzer" + participant CallA as "CallGraphAnalyzer" + participant TS as "Tree-sitter Analyzers" + + Builder->>Parser: parse_repository() + Parser->>Service: _analyze_structure(repo_path) + Service->>RepoA: analyze_repository_structure() + RepoA-->>Service: file_tree + summary + Parser->>Service: _analyze_call_graph(file_tree, repo_path) + Service->>CallA: extract_code_files() / analyze_code_files() + CallA->>TS: analyze_python_file / analyze_javascript_file_treesitter / ... + TS-->>CallA: functions, relationships + CallA-->>Service: functions, relationships, visualization + Service-->>Parser: call_graph_result + Parser->>Parser: build namespaced Node components + Parser-->>Builder: components (Dict[str, Node]) + Builder->>Builder: build_graph_from_components / validate / filter leaves + Builder-->>Builder: (components, leaf_nodes) +``` + +## Sub-modules + +### Analysis Pipeline + +Handles repository cloning support, file-tree discovery with include/exclude filtering, and multi-language call graph extraction by delegating to per-language analyzers. Composed of `AnalysisService`, `RepoAnalyzer`, and `CallGraphAnalyzer`. + +See [Analysis Pipeline](analysis_pipeline.md) for full details. + +### Graph Construction + +Turns raw analysis results into fully-qualified, namespaced `Node` components, resolves single- and multi-repository (dependency-aware) call relationships, and drives graph validation, leaf-node filtering, and persistence to JSON. Composed of `DependencyParser` and `DependencyGraphBuilder`. + +See [Graph Construction](graph_construction.md) for full details. + +## Key Concepts + +### Fully-Qualified Domain Names (FQDNs) + +Every extracted code component is assigned a unique identifier of the form `namespace.module.path::ComponentName`. The `namespace` prefix is derived from the source repository's directory name and prevents ID collisions when analyzing multiple repositories together (e.g., a primary repo plus its dependencies). See [Graph Construction](graph_construction.md) for details on how `DependencyParser` builds and resolves these identifiers. + +### Multi-Language Support + +The pipeline supports Python, JavaScript, TypeScript, Java, C#, C, C++, PHP, Go, and Rust. Python uses a native AST analyzer; all other languages are handled via tree-sitter based analyzers described in [Tree-sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md). + +### Single-Path vs. Multi-Path Analysis + +`DependencyGraphBuilder` and `DependencyParser` both support analyzing either a single repository or multiple repository paths simultaneously (for example, a main repository plus one or more external dependency sources). In multi-path mode, each source directory is assigned its own namespace, and cross-namespace dependency edges are resolved after all components have been extracted. See [Graph Construction](graph_construction.md). + +## Summary + +| Component | Responsibility | +|---|---| +| `AnalysisService` | Top-level orchestration: repository cloning, structure + call-graph analysis, README extraction, cleanup | +| `RepoAnalyzer` | Builds filtered file trees for one or more repository paths | +| `CallGraphAnalyzer` | Routes files to per-language analyzers, resolves and deduplicates call relationships, builds visualization data | +| `DependencyParser` | Converts analysis results into namespaced `Node` components; resolves cross-namespace dependencies | +| `DependencyGraphBuilder` | Drives parsing end-to-end, validates graph completeness, filters leaf nodes, persists the dependency graph | + +For details on each area, see [Analysis Pipeline](analysis_pipeline.md) and [Graph Construction](graph_construction.md). diff --git a/docs/reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md b/docs/reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md new file mode 100644 index 00000000..12ef0668 --- /dev/null +++ b/docs/reference/architecture/backend-core/dependency-analyzer-core/graph_construction.md @@ -0,0 +1,218 @@ +# Graph Construction + +The Graph Construction module is the stage-two engine of CodeWiki's dependency analysis pipeline. It transforms raw structural and call-graph data produced by the [Analysis Pipeline](../analysis_pipeline/analysis_pipeline.md) into a fully-resolved, persistable dependency graph made of typed `Node` objects. It is responsible for assigning fully-qualified domain names (FQDNs) to every code component, namespacing components when multiple repositories are analyzed together, resolving both intra- and cross-namespace dependency edges, and filtering the resulting graph down to a set of "leaf" components suitable for downstream clustering and documentation generation. + +This module contains two core components: + +- **`DependencyParser`** — converts raw analysis output (functions, classes, relationships) into namespaced `Node` objects and resolves dependency edges between them, supporting both single-repository and multi-repository ("multi-path") modes. +- **`DependencyGraphBuilder`** — orchestrates the end-to-end graph-building workflow: invoking the parser, persisting the graph to disk, building an in-memory traversal graph, validating its completeness, and filtering leaf nodes for downstream consumption. + +## Purpose and Scope + +The Graph Construction module sits between raw source analysis and the higher-level clustering/documentation stages of CodeWiki. Its responsibilities are strictly bounded to: + +1. **Component identity assignment** — converting analyzer-produced identifiers into canonical FQDNs of the form `{namespace}.{module.path}::{ComponentName}`. +2. **Namespace management** — when analyzing multiple source directories (multi-path mode), prefixing components by their originating repository/directory to avoid ID collisions. +3. **Dependency edge resolution** — mapping raw caller/callee identifiers extracted by language analyzers into resolved FQDN-to-FQDN edges, including matching across namespace boundaries by component name when a direct ID match is unavailable. +4. **Graph persistence** — serializing the final component graph to a deterministic (sorted) JSON representation on disk. +5. **Leaf node filtering** — identifying and validating the "leaf" components (typically classes, interfaces, structs, or functions for C-style codebases) that anchor the hierarchical documentation generation process. + +It does **not** perform language-specific parsing itself (that responsibility belongs to the [Tree-Sitter Analyzers](../../tree-sitter-analyzers/tree-sitter-analyzers.md)) nor does it run the structural/call-graph analysis (handled by `AnalysisService` in the [Analysis Pipeline](../analysis_pipeline/analysis_pipeline.md)). It also delegates the `Node` data model itself to [Dependency Analyzer Models](../../dependency-analyzer-models/dependency-analyzer-models.md). + +## Architecture + +```mermaid +flowchart TD + Config["Config"] --> Builder["DependencyGraphBuilder"] + Builder -->|"instantiates"| Parser["DependencyParser"] + Parser -->|"uses"| AnalysisSvc["AnalysisService"] + AnalysisSvc -->|"structure + call graph"| Parser + Parser -->|"produces"| Nodes["Node objects (Dict[str, Node])"] + Builder -->|"build_graph_from_components()"| TraversalGraph["In-memory traversal graph"] + Builder -->|"validate_graph_completeness()"| Validation["Graph validation"] + Builder -->|"get_leaf_nodes()"| LeafFilter["Leaf node filtering"] + Parser -->|"save_dependency_graph()"| JSONFile["dependency_graph.json"] + LeafFilter --> Output["(components, leaf_nodes)"] +``` + +`DependencyGraphBuilder` is constructed with a [Config](../../../config-core.md) instance that supplies the repository path(s), output directories, and include/exclude filter patterns. It then instantiates a `DependencyParser` scoped to either a single repository path (`str`) or a list of paths (multi-path mode), delegating all structural/call-graph analysis to `AnalysisService` from the [Analysis Pipeline](../analysis_pipeline/analysis_pipeline.md). + +## Core Components + +### DependencyParser + +`DependencyParser` is the component-extraction engine. It wraps an `AnalysisService` instance and converts its raw output into the canonical `Node` representation used throughout the rest of CodeWiki. + +**Initialization** accepts either a single repository path or a list of paths, along with optional `include_patterns` / `exclude_patterns` file filters: + +```python +parser = DependencyParser( + repo_path=["/path/to/main-repo", "/path/to/dependency-repo"], + include_patterns=["*.py", "*.ts"], + exclude_patterns=["*Tests*"] +) +components = parser.parse_repository() +``` + +**Key responsibilities:** + +| Method | Purpose | +|---|---| +| `parse_repository()` | Entry point; dispatches to single- or multi-path parsing based on how many repo paths were configured. | +| `_parse_single_repository()` | Backward-compatible path: analyzes structure and call graph for one repository, then builds components via `_build_components_from_analysis()`. | +| `_parse_multiple_repositories()` | Analyzes each configured path independently, namespaces its components, merges all namespaces together, then resolves cross-namespace dependencies. | +| `_build_namespaced_components()` | First pass creates `Node` objects with FQDN keys (`{namespace}.{original_id}`); second pass wires up `depends_on` edges within the same namespace. | +| `_resolve_cross_namespace_dependencies()` | For any dependency edge that didn't resolve within its own namespace, attempts a name-based match against components in *other* namespaces. | +| `_build_components_from_analysis()` | Single-path equivalent of namespaced component construction; also tracks legacy `file_path:name` IDs for backward-compatible dependency resolution. | +| `save_dependency_graph()` | Serializes all components (sorted by ID, with `depends_on` sets converted to sorted lists) to a JSON file for deterministic, diffable output. | + +#### FQDN Construction + +Every component is identified by a Fully Qualified Domain Name in the form: + +```text +{namespace}.{module.path}::{ComponentName} +``` + +- `namespace` is derived from the last path segment of the source directory being analyzed (e.g., `openframe-frontend`, `ui-kit`), computed by `_get_namespace_from_path()`. +- `module.path::ComponentName` is the raw identifier produced by the language-specific analyzer (see [Tree-Sitter Analyzers](../../tree-sitter-analyzers/tree-sitter-analyzers.md)). +- In single-path mode, `is_from_deps` is always `False`. In multi-path mode, the first configured path (`repo_index == 0`) is treated as the primary repository, while subsequent paths are marked `is_from_deps=True`. + +Each `Node` also retains `short_id` (the original, un-namespaced identifier) purely for display purposes, alongside the `namespace` string itself — see the [Dependency Analyzer Models](../../dependency-analyzer-models/dependency-analyzer-models.md) documentation for the full `Node` schema. + +#### Single-Path vs. Multi-Path Parsing + +```mermaid +flowchart TD + Start["parse_repository()"] --> Check{{"len(repo_paths) == 1?"}} + Check -->|"Yes"| Single["_parse_single_repository()"] + Check -->|"No"| Multi["_parse_multiple_repositories()"] + + Single --> S1["analysis_service._analyze_structure()"] + S1 --> S2["analysis_service._analyze_call_graph()"] + S2 --> S3["_build_components_from_analysis()"] + S3 --> SResult["self.components"] + + Multi --> M1["For each repo path: compute namespace"] + M1 --> M2["analysis_service._analyze_structure()"] + M2 --> M3["analysis_service._analyze_call_graph()"] + M3 --> M4["_build_namespaced_components()"] + M4 --> M5["Merge into all_components"] + M5 --> M6{{"More paths?"}} + M6 -->|"Yes"| M1 + M6 -->|"No"| M7["_resolve_cross_namespace_dependencies()"] + M7 --> MResult["self.components"] +``` + +In multi-path mode, each configured source directory is analyzed independently by `AnalysisService`, then merged into a single component dictionary keyed by namespaced FQDN. A `namespace_mapping` dict (original ID → FQDN) is built incrementally as each repository is processed, enabling within-namespace dependency resolution during `_build_namespaced_components()`. + +#### Cross-Namespace Dependency Resolution + +After all repositories are parsed and merged, `_resolve_cross_namespace_dependencies()` performs a second reconciliation pass. For every component's `depends_on` set: + +1. If the dependency ID already matches an entry in `all_components`, it is kept as-is. +2. Otherwise, the parser extracts the trailing component name (`dep_id.split(".")[-1]`) and searches all components for a name match belonging to a *different* namespace, logging the resolution as a cross-namespace dependency. +3. If no match is found at all, the original (unresolved) dependency ID is preserved rather than dropped, ensuring no data loss even when resolution fails. + +This name-based fallback allows dependency edges to survive even when different analyzers or namespaces produce slightly different identifier formats for the same logical target. + +### DependencyGraphBuilder + +`DependencyGraphBuilder` is the orchestration layer that wraps `DependencyParser` with graph persistence, traversal-graph construction, validation, and leaf-node filtering. It is the primary entry point used by the rest of the backend pipeline to obtain a ready-to-cluster dependency graph. + +```python +builder = DependencyGraphBuilder(config) +components, leaf_nodes = builder.build_dependency_graph() +``` + +**Workflow performed by `build_dependency_graph()`:** + +```mermaid +flowchart TD + A["build_dependency_graph()"] --> B["Ensure dependency_graph_dir exists"] + B --> C["Compute sanitized dependency_graph_path"] + C --> D["Resolve include/exclude patterns from Config"] + D --> E["Build repo_paths from config.all_source_paths"] + E --> F["Instantiate DependencyParser"] + F --> G["parser.parse_repository()"] + G --> H["Log component type breakdown"] + H --> I["parser.save_dependency_graph(path)"] + I --> J["build_graph_from_components(components)"] + J --> K["validate_graph_completeness(components, graph)"] + K --> L["get_leaf_nodes(graph, components)"] + L --> M["Determine valid leaf types from available component types"] + M --> N["Filter leaf_nodes: skip invalid / wrong-type / not-found"] + N --> O["Return (components, keep_leaf_nodes)"] +``` + +**Key behaviors:** + +- **Source path resolution**: Uses `config.all_source_paths` (primary `repo_path` plus any `additional_source_paths`) to decide whether to invoke single-path or multi-path parsing on `DependencyParser`. See [Config](../../../config-core.md) for how these paths are validated and exposed. +- **Deterministic output naming**: The dependency graph JSON file is named `{sanitized_repo_name}_dependency_graph.json`, where the repository's base directory name is sanitized to alphanumeric characters and underscores. +- **Graph traversal preparation**: After parsing, `build_graph_from_components()` and `get_leaf_nodes()` (internal graph/topology utilities) convert the flat `Node` dictionary into a traversable graph structure and extract nodes with no outgoing dependencies as candidate "leaves." +- **Post-build validation**: `validate_graph_completeness()` is invoked immediately after graph construction to catch structural inconsistencies before leaf filtering proceeds. +- **Leaf node type filtering**: Leaf nodes are only retained if their `component_type` is one of `class`, `interface`, or `struct` — unless *none* of the parsed components have those types (e.g., a purely procedural/C-style codebase), in which case `function` is also accepted. Nodes that are empty, contain error-like keywords (`error`, `exception`, `failed`, `invalid`), or are absent from the parsed `components` dictionary are skipped and logged with a specific reason. + +#### Leaf Node Filtering Logic + +```mermaid +flowchart TD + Start["For each leaf_node in leaf_nodes"] --> V1{{"Is leaf_node a valid non-empty string without error keywords?"}} + V1 -->|"No"| SkipInvalid["skipped_invalid += 1"] + V1 -->|"Yes"| V2{{"leaf_node in components?"}} + V2 -->|"No"| SkipNotFound["skipped_not_found += 1"] + V2 -->|"Yes"| V3{{"component_type in valid_types?"}} + V3 -->|"No"| SkipType["skipped_type += 1"] + V3 -->|"Yes"| Keep["keep_leaf_nodes.append(leaf_node)"] +``` + +The resulting `keep_leaf_nodes` list, together with the full `components` dictionary, is returned to the caller and forms the input for the clustering and documentation-generation stages that follow in the broader backend pipeline. + +## Data Model + +Both `DependencyParser` and `DependencyGraphBuilder` operate on the `Node` model defined in [Dependency Analyzer Models](../../dependency-analyzer-models/dependency-analyzer-models.md). The fields most relevant to graph construction are: + +| Field | Description | +|---|---| +| `id` | Primary key; the FQDN (`{namespace}.{original_id}`). | +| `short_id` | Original, un-namespaced identifier from the language analyzer. | +| `namespace` | Source directory namespace (e.g., repository directory name). | +| `is_from_deps` | `True` if the component came from a non-primary (dependency) source path in multi-path mode. | +| `component_type` | Classifies the node (`class`, `interface`, `struct`, `function`, `method`, etc.); drives leaf-node filtering. | +| `depends_on` | Set of FQDNs this component depends on; populated and reconciled by `DependencyParser`. | + +## Component Interactions + +```mermaid +sequenceDiagram + participant Caller as "Pipeline Caller" + participant Builder as "DependencyGraphBuilder" + participant Parser as "DependencyParser" + participant Svc as "AnalysisService" + participant FS as "File System" + + Caller->>Builder: build_dependency_graph() + Builder->>Parser: new DependencyParser(repo_paths, patterns) + Builder->>Parser: parse_repository() + Parser->>Svc: _analyze_structure(repo_path) + Svc-->>Parser: structure_result + Parser->>Svc: _analyze_call_graph(file_tree, repo_path) + Svc-->>Parser: call_graph_result + Parser->>Parser: _build_components_from_analysis() / _build_namespaced_components() + Parser-->>Builder: components (Dict[str, Node]) + Builder->>Parser: save_dependency_graph(path) + Parser->>FS: write JSON + Builder->>Builder: build_graph_from_components(components) + Builder->>Builder: validate_graph_completeness(components, graph) + Builder->>Builder: get_leaf_nodes(graph, components) + Builder-->>Caller: (components, leaf_nodes) +``` + +## Relationship to the Broader Pipeline + +- **Upstream**: Relies on `AnalysisService` from the [Analysis Pipeline](../analysis_pipeline/analysis_pipeline.md) for structural file-tree scanning and call-graph extraction, which in turn delegates language-specific parsing to the [Tree-Sitter Analyzers](../../tree-sitter-analyzers/tree-sitter-analyzers.md). +- **Configuration**: Reads repository paths, filter patterns, and output directories from the [Config](../../../config-core.md) object. +- **Data Model**: Produces and manipulates `Node` instances defined in [Dependency Analyzer Models](../../dependency-analyzer-models/dependency-analyzer-models.md). +- **Downstream**: The `(components, leaf_nodes)` tuple returned by `DependencyGraphBuilder.build_dependency_graph()` feeds into the clustering and documentation-generation stages that consume the persisted dependency graph and the filtered leaf set for hierarchical documentation planning. + +For the broader dependency-analysis subsystem this module belongs to, see the [Dependency Analyzer Core](../dependency-analyzer-core.md) overview. diff --git a/docs/reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md b/docs/reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md new file mode 100644 index 00000000..a65c76c6 --- /dev/null +++ b/docs/reference/architecture/backend-core/dependency-analyzer-models/dependency-analyzer-models.md @@ -0,0 +1,194 @@ +# Dependency Analyzer Models + +## Introduction + +The Dependency Analyzer Models module defines the core Pydantic data contracts used throughout the dependency analysis pipeline of the backend. It contains no business logic of its own — instead, it establishes the shared vocabulary (data shapes) that every other backend component relies on to represent source code entities, their relationships, the repositories being analyzed, and the final analysis output. + +Because these models are pure data definitions with no external side effects, this module acts as the **foundation layer** of the backend dependency-analysis subsystem. Every producer (parsers, analyzers, graph builders) and every consumer (documentation generator, agent orchestrator, CLI) of analysis data depends on these types either directly or transitively. + +The module is composed of two files: + +- `models/core.py` — the fundamental building blocks: `Node`, `CallRelationship`, `Repository` +- `models/analysis.py` — the higher-level, aggregate result types: `AnalysisResult`, `NodeSelection` + +Given its small size and purely declarative nature, this module is documented as a single cohesive unit rather than split into further sub-modules. + +## Purpose in the Overall System + +The dependency-analysis pipeline works roughly as follows: + +1. Language-specific tree-sitter/AST analyzers parse source files and extract code entities and call edges. +2. The `DependencyParser` and `DependencyGraphBuilder` (part of the Dependency Analyzer Core module) assemble these extracted entities into `Node` and `CallRelationship` instances. +3. The `AnalysisService`, `RepoAnalyzer`, and `CallGraphAnalyzer` (also part of Dependency Analyzer Core) orchestrate the end-to-end analysis of a repository, producing an `AnalysisResult`. +4. Downstream consumers — the Documentation Generator, the Agent Orchestrator, and the CLI — read the `AnalysisResult` (and its nested `Node`/`CallRelationship`/`Repository` data) to generate documentation, drive AI agents, or render reports. + +This module's models are therefore the "wire format" passed between all of those stages. + +```mermaid +flowchart TD + Analyzers["Tree-sitter Analyzers"] -->|"produce"| Node["Node"] + Analyzers -->|"produce"| CallRel["CallRelationship"] + Parser["Dependency Parser / Graph Builder"] -->|"assembles"| Node + Parser -->|"assembles"| CallRel + RepoModel["Repository"] -->|"describes source of"| Node + AnalysisSvc["Analysis Service / Repo Analyzer / Call Graph Analyzer"] -->|"aggregates"| Node + AnalysisSvc -->|"aggregates"| CallRel + AnalysisSvc -->|"aggregates"| RepoModel + AnalysisSvc -->|"produces"| Result["AnalysisResult"] + Selection["NodeSelection"] -->|"filters"| Result + Result -->|"consumed by"| DocGen["Documentation Generator"] + Result -->|"consumed by"| Agent["Agent Orchestrator"] + Result -->|"consumed by"| CLIGen["CLI Documentation Generator"] +``` + +## Core Components + +### `Node` + +`Node` (defined in `models/core.py`) is the canonical representation of a single source-code entity discovered during analysis — a function, class, method, or other addressable code unit. + +Key fields: + +- `id`: Fully-qualified domain name (FQDN) in the format `{namespace}.{original_id}`, used as the unique identifier across the entire system. +- `name`, `display_name`, `class_name`: Human-readable naming information; `get_display_name()` falls back to `name` when `display_name` is not set. +- `component_type`, `node_type`: Classify the kind of entity (e.g., function, class, method). +- `file_path`, `relative_path`: Location of the entity within the analyzed repository. +- `depends_on`: A `Set[str]` of other node IDs this node depends on — the backbone of the dependency graph. +- `source_code`, `start_line`, `end_line`: Raw source snippet and its location for documentation/reference purposes. +- `has_docstring`, `docstring`: Extracted documentation comments. +- `parameters`, `base_classes`: Signature and inheritance metadata (when applicable). +- `component_id`: Optional external identifier correlating a node with a documentation component. +- FQDN metadata: `short_id` (original ID without namespace), `namespace` (e.g., `main`, `deps`, `ui-kit`), and `is_from_deps` (whether the node originates from a dependency rather than the main repository). + +`Node` is intentionally permissive — most fields beyond `id`, `name`, `component_type`, `file_path`, and `relative_path` are optional or default-valued, allowing different language analyzers (Python, Java, C/C++/C#, JavaScript/TypeScript, PHP) to populate only the metadata that is available for their language. + +### `CallRelationship` + +`CallRelationship` models a directed edge between two `Node` instances, representing a call or reference from one code entity to another. + +Fields: + +- `caller`: ID of the node initiating the call. +- `callee`: ID of the node being called. +- `call_line`: Optional line number where the call occurs, useful for navigation and documentation cross-referencing. +- `is_resolved`: Whether the callee could be definitively resolved to a known `Node` (as opposed to remaining an unresolved/external reference). + +Collections of `CallRelationship` objects, combined with `Node.depends_on` sets, form the call graph that the Call Graph Analyzer and Dependency Graph Builder operate on. + +### `Repository` + +`Repository` captures metadata about the source repository under analysis: + +- `url`: Origin location of the repository (e.g., a Git remote URL). +- `name`: Human-readable repository name. +- `clone_path`: Local filesystem path where the repository was cloned/checked out for analysis. +- `analysis_id`: Identifier correlating this repository record with a specific analysis run. + +This model is embedded directly inside `AnalysisResult` to associate analysis output with its source. + +### `AnalysisResult` + +`AnalysisResult` (defined in `models/analysis.py`) is the top-level aggregate produced at the end of a full repository analysis. It bundles together everything downstream consumers need: + +- `repository`: The `Repository` that was analyzed. +- `functions`: A `List[Node]` of all extracted code entities. +- `relationships`: A `List[CallRelationship]` describing the call graph. +- `file_tree`: A `Dict[str, Any]` representing the hierarchical file/directory structure of the repository. +- `summary`: A `Dict[str, Any]` with aggregate statistics or high-level findings. +- `visualization`: An optional `Dict[str, Any]` holding precomputed visualization data (defaults to an empty dict). +- `readme_content`: An optional string containing the repository's README content, when available. + +This model is the primary data structure passed from the analysis pipeline (Analysis Service, Repo Analyzer, Call Graph Analyzer) to consumers such as the Documentation Generator and the Agent Orchestrator. + +### `NodeSelection` + +`NodeSelection` supports partial/filtered exports of an analysis result — for example, when a user wants documentation generated for only a subset of discovered nodes. + +Fields: + +- `selected_nodes`: A `List[str]` of node IDs to include (defaults to an empty list). +- `include_relationships`: Whether relationships between selected nodes should also be included (defaults to `True`). +- `custom_names`: A `Dict[str, str]` mapping node IDs to user-supplied display names, allowing renaming without mutating the underlying `Node` data. + +## Data Model Relationships + +```mermaid +classDiagram + class Node { + +str id + +str name + +str component_type + +str file_path + +str relative_path + +Set~str~ depends_on + +Optional~str~ source_code + +int start_line + +int end_line + +bool has_docstring + +str docstring + +Optional~List~str~~ parameters + +Optional~str~ node_type + +Optional~List~str~~ base_classes + +Optional~str~ class_name + +Optional~str~ display_name + +Optional~str~ component_id + +str short_id + +str namespace + +bool is_from_deps + +get_display_name() str + } + + class CallRelationship { + +str caller + +str callee + +Optional~int~ call_line + +bool is_resolved + } + + class Repository { + +str url + +str name + +str clone_path + +str analysis_id + } + + class AnalysisResult { + +Repository repository + +List~Node~ functions + +List~CallRelationship~ relationships + +Dict file_tree + +Dict summary + +Dict visualization + +Optional~str~ readme_content + } + + class NodeSelection { + +List~str~ selected_nodes + +bool include_relationships + +Dict~str,str~ custom_names + } + + AnalysisResult "1" --> "1" Repository : describes + AnalysisResult "1" --> "*" Node : contains + AnalysisResult "1" --> "*" CallRelationship : contains + CallRelationship "*" --> "1" Node : caller/callee reference (by id) + NodeSelection "1" --> "*" Node : references (by id) +``` + +## Usage Across the System + +These models form the shared contract consumed by several other backend modules: + +- The tree-sitter language analyzers and the AST/graph construction components (`DependencyParser`, `DependencyGraphBuilder`) construct `Node` and `CallRelationship` instances as they walk source files. +- The analysis pipeline components (`AnalysisService`, `RepoAnalyzer`, `CallGraphAnalyzer`) assemble these into a single `Repository`-scoped `AnalysisResult`. +- The Documentation Generator consumes `AnalysisResult` to produce human-readable documentation, optionally filtering via `NodeSelection` for partial exports. +- The Agent Orchestrator and its tooling read `Node` and `CallRelationship` data to answer questions about code structure and to drive AI-assisted documentation generation. +- The CLI's documentation generation adapter ultimately surfaces `AnalysisResult` data to end users. + +Because these are plain Pydantic `BaseModel` classes, they also provide built-in serialization/validation, making it straightforward to persist analysis results as JSON, pass them between processes, or validate data received from external sources. + +## Design Notes + +- **FQDN-based identity**: The `id` field on `Node` always follows the `{namespace}.{original_id}` convention, ensuring uniqueness even when analyzing a main repository alongside its dependencies. The `namespace` and `is_from_deps` fields make it possible to distinguish first-party code from vendored/dependency code without losing the original identifier (`short_id`). +- **Permissive defaults**: Nearly every field beyond the minimal identity fields on `Node` has a sensible default (empty set, `None`, empty string), which allows the model to accommodate the varying levels of metadata different language analyzers can extract. +- **Separation of raw graph data from aggregate results**: `Node`/`CallRelationship`/`Repository` represent the atomic units of the dependency graph, while `AnalysisResult` represents the fully-assembled output of a completed analysis run. `NodeSelection` sits alongside `AnalysisResult` as a lightweight filter/view specification rather than a graph primitive. diff --git a/docs/reference/architecture/backend-core/documentation-generator/documentation-generator.md b/docs/reference/architecture/backend-core/documentation-generator/documentation-generator.md new file mode 100644 index 00000000..488c05a1 --- /dev/null +++ b/docs/reference/architecture/backend-core/documentation-generator/documentation-generator.md @@ -0,0 +1,228 @@ +# Documentation Generator + +The Documentation Generator module contains the `DocumentationGenerator` class — the top-level orchestrator that drives the entire end-to-end documentation generation pipeline for a target repository. It coordinates dependency graph construction, module clustering, hierarchical (dynamic-programming style) documentation generation via AI agents, and final metadata assembly. + +This module is the primary entry point invoked by both the [CLI Core](../../cli-core.md) (`CLIDocumentationGenerator`) and the [Frontend Core](../../frontend-core.md) (`BackgroundWorker`) layers whenever a documentation job needs to be executed against a cloned or provided repository. + +## Purpose and Responsibilities + +`DocumentationGenerator` ties together several backend subsystems to transform raw source code into a hierarchical set of Markdown documentation files: + +- **Dependency graph construction**: Delegates to `DependencyGraphBuilder` (see [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md)) to parse the repository and produce a component graph plus a filtered list of "leaf nodes" (classes, interfaces, structs, or functions worth documenting). +- **Module clustering**: Groups leaf-level components into a hierarchical module tree using `cluster_modules`, with a **synthetic module fallback** that prevents context-overflow failures when clustering yields no groups. +- **Documentation generation ordering**: Computes a topological "leaf-modules-first" processing order so that child module documentation is always available before its parent module summary is generated. +- **Agent-driven content generation**: Delegates the actual LLM-backed documentation writing to `AgentOrchestrator` (see [Backend Core](../../backend-core.md)) for leaf modules, and to its own `generate_parent_module_docs` routine (using `REPO_OVERVIEW_PROMPT` / `MODULE_OVERVIEW_PROMPT`) for parent/overview modules. +- **Hierarchical file layout**: Computes nested working directories (`module_path` → `docs_dir/A/B/C/C.md`) so that every module's documentation lives in its own directory alongside its children. +- **Metadata generation**: After generation completes, writes a `metadata.json` describing generation statistics and the list of all generated Markdown files. + +## Position in the System + +```mermaid +flowchart TD + CLI["CLI Core: CLIDocumentationGenerator"] --> DG["Documentation Generator: DocumentationGenerator"] + Web["Frontend Core: BackgroundWorker"] --> DG + DG --> GraphBuilder["Dependency Analyzer Core: DependencyGraphBuilder"] + DG --> Cluster["cluster_modules"] + DG --> Orchestrator["Backend Core: AgentOrchestrator"] + DG --> LLM["LLM Services: call_llm / CountingFallbackModel"] + GraphBuilder --> Analyzers["Tree-sitter Analyzers"] + GraphBuilder --> Models["Dependency Analyzer Models"] + Orchestrator --> AgentTools["Agent Tools Core"] + DG --> ConfigCore["Config Core: Config"] +``` + +- **Upstream callers**: [CLI Core](../../cli-core.md) and [Frontend Core](../../frontend-core.md) both construct a `Config` (see [Config Core](../../../config-core.md)) and instantiate `DocumentationGenerator(config, commit_id)` before calling `run()`. +- **Downstream collaborators**: + - `DependencyGraphBuilder` from [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md) parses the repository via `DependencyParser` and the [Tree-sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md). + - `AgentOrchestrator` (in [Backend Core](../../backend-core.md)) creates and runs `pydantic-ai` agents backed by `CountingFallbackModel` from [LLM Services](../llm-services/llm-services.md), using tools defined in [Agent Tools Core](../agent-tools-core/agent-tools-core.md). + - `Node`, `AnalysisResult` and related types from [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) represent the parsed components passed throughout the pipeline. + +## Core Component + +### `DocumentationGenerator` + +```python +class DocumentationGenerator: + def __init__(self, config: Config, commit_id: str = None): + self.config = config + self.commit_id = commit_id + self.graph_builder = DependencyGraphBuilder(config) + self.agent_orchestrator = AgentOrchestrator(config) +``` + +On construction, `DocumentationGenerator` immediately wires up its two main collaborators: + +| Attribute | Type | Role | +|---|---|---| +| `config` | `Config` | Repository path, model configuration, output directories, agent instructions | +| `commit_id` | `str \| None` | Optional git commit SHA recorded in generation metadata | +| `graph_builder` | `DependencyGraphBuilder` | Parses the repository and produces components + leaf nodes | +| `agent_orchestrator` | `AgentOrchestrator` | Runs LLM agents to write documentation for leaf modules | + +#### Public / Key Methods + +| Method | Purpose | +|---|---| +| `run()` | Entry point — runs the complete pipeline: graph build → clustering (with synthetic fallback) → module documentation generation → metadata creation | +| `generate_module_documentation(components, leaf_nodes)` | Iterates the module tree in leaf-first order, invoking either the agent orchestrator (leaf modules) or `generate_parent_module_docs` (parent modules), and finally the repository overview | +| `generate_parent_module_docs(module_path, working_dir)` | Builds a summarized child-doc structure and calls the LLM (`REPO_OVERVIEW_PROMPT` / `MODULE_OVERVIEW_PROMPT`) to synthesize a parent/overview document | +| `get_processing_order(module_tree, parent_path)` | Computes a topological (post-order/leaf-first) traversal of the module tree | +| `is_leaf_module(module_info)` | Determines whether a module has no children (i.e., should be documented directly from components) | +| `build_overview_structure(module_tree, module_path, working_dir)` | Produces a JSON-serializable subtree with one level of children's already-generated docs embedded, marking the current target module | +| `_get_nested_working_dir(base_dir, module_path)` | Computes the hierarchical output directory for a given module path (e.g. `docs_dir/Backend/Auth/JWT/`) | +| `create_documentation_metadata(working_dir, components, num_leaf_nodes)` | Walks the output directory for generated `.md` files and writes `metadata.json` | + +## End-to-End Generation Flow + +```mermaid +sequenceDiagram + participant Caller as "CLI/Web Caller" + participant DG as "DocumentationGenerator" + participant GB as "DependencyGraphBuilder" + participant Cluster as "cluster_modules" + participant AO as "AgentOrchestrator" + participant LLM as "call_llm" + participant FS as "file_manager" + + Caller->>DG: run() + DG->>GB: build_dependency_graph() + GB-->>DG: components, leaf_nodes + alt "module tree cache exists" + DG->>FS: load_json(first_module_tree.json) + else "no cache" + DG->>Cluster: cluster_modules(leaf_nodes, components, config) + Cluster-->>DG: module_tree + DG->>FS: save_json(first_module_tree.json) + end + alt "module_tree empty but leaf_nodes exist" + DG->>DG: build synthetic modules by top-level directory + DG->>FS: save_json(synthetic module_tree) + end + DG->>FS: save_json(module_tree.json) + DG->>DG: generate_module_documentation(components, leaf_nodes) + loop "for each module in leaf-first order" + alt "is_leaf_module" + DG->>AO: process_module(name, components, ids, path, dir) + AO->>LLM: agent.run(user_prompt) + LLM-->>AO: markdown docs + AO->>FS: save module_name.md + else "parent module" + DG->>DG: build_overview_structure(...) + DG->>LLM: call_llm(MODULE_OVERVIEW_PROMPT) + LLM-->>DG: "..." + DG->>FS: save parent module_name.md + end + end + DG->>DG: generate_parent_module_docs([], working_dir) + DG->>LLM: call_llm(REPO_OVERVIEW_PROMPT) + DG->>FS: save overview.md + DG->>DG: create_documentation_metadata(working_dir, components, len(leaf_nodes)) + DG-->>Caller: "documentation complete" +``` + +### 1. Dependency Graph Construction + +`run()` first delegates to `self.graph_builder.build_dependency_graph()`, returning: +- `components`: a dict mapping component IDs to parsed `Node` objects (see [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md)) +- `leaf_nodes`: a filtered list of component IDs that are candidates for direct documentation (classes/interfaces/structs, or functions for C-style codebases) + +### 2. Module Clustering (with Synthetic Fallback) + +If a cached `first_module_tree.json` exists, it is reused. Otherwise `cluster_modules(leaf_nodes, components, config)` groups leaf nodes into a semantic module hierarchy. + +A **synthetic module patch** guards against an important failure mode: if clustering returns an empty tree despite having leaf nodes (which would otherwise trigger an expensive "whole repository in one pass" fallback that can exceed LLM context limits), `DocumentationGenerator` instead groups leaf nodes by their top-level source directory and constructs flat synthetic modules directly. This synthetic tree is persisted back to the cache file so subsequent runs skip re-clustering. + +### 3. Leaf-First Processing Order + +`get_processing_order` performs a recursive collection over the module tree: for any module with children, it recurses into the children **first** and appends the parent path afterward. This guarantees a true dynamic-programming order — every child module's Markdown file already exists on disk by the time its parent module is processed, since `build_overview_structure` reads children's generated docs directly from `working_dir/child_name/child_name.md`. + +```mermaid +flowchart LR + subgraph Tree["Module Tree Example"] + Root["Backend"] --> Auth["Authentication"] + Root --> API["API Layer"] + Auth --> JWT["JWT"] + Auth --> OAuth["OAuth"] + end + subgraph Order["Processing Order"] + O1["1: JWT (leaf)"] --> O2["2: OAuth (leaf)"] + O2 --> O3["3: Authentication (parent)"] + O3 --> O4["4: API Layer (leaf/parent)"] + O4 --> O5["5: Backend (parent)"] + end +``` + +### 4. Per-Module Documentation Generation + +For each module in the processing order, `generate_module_documentation`: +1. Resolves `module_info` by walking the module tree along `module_path`. +2. Skips modules already processed (idempotency across retries). +3. Computes the nested working directory via `_get_nested_working_dir` (e.g. `docs_dir/Backend/Authentication/JWT/`) and ensures it exists. +4. **Leaf modules** — delegates to `self.agent_orchestrator.process_module(...)`, which builds a `pydantic-ai` agent (simple or "complex" with sub-module tools) and runs it against the module's component list. See [Backend Core](../../backend-core.md) for orchestrator internals. +5. **Parent modules** — delegates to `self.generate_parent_module_docs(...)`, which synthesizes an overview from already-generated child documentation. +6. Errors for any single module are logged with a full traceback but do **not** halt the overall run — processing continues with the next module (graceful degradation). + +After all modules are processed, a final call to `generate_parent_module_docs([], working_dir)` produces the top-level repository `overview.md`. + +### 5. Parent / Overview Document Synthesis + +`generate_parent_module_docs`: +1. Loads the canonical `module_tree.json` from the **base** docs directory (not the nested `working_dir`, since `module_tree.json` is always written at the top level). +2. Returns early if `overview.md` or the module's own `.md` already exists (idempotent caching). +3. Calls `build_overview_structure` to construct a subtree containing the target module marked with `is_target_for_overview_generation` and each direct child's already-generated Markdown embedded under a `docs` key. +4. Serializes this structure to JSON and formats it into `MODULE_OVERVIEW_PROMPT` (for non-root modules) or `REPO_OVERVIEW_PROMPT` (for the repository root, when `module_path` is empty). +5. Invokes `call_llm(prompt, self.config)` (see [LLM Services](../llm-services/llm-services.md)), extracts the content between `` / `` tags (falling back to the raw response if tags are missing), and writes it to disk via `file_manager.save_text`. + +### 6. Small-Repository Fast Path + +If clustering (including the synthetic fallback) still yields **zero** modules, `run()` falls back to treating the entire repository as a single module: it calls `agent_orchestrator.process_module` directly with all `leaf_nodes`, then renames the resulting `.md` to `overview.md`. + +### 7. Metadata Generation + +Once all documentation is written, `create_documentation_metadata` walks the output directory tree, records every generated `.md` file (relative path), and writes a `metadata.json` capturing: +- Generation timestamp, model used, generator version, repo path, and commit ID +- Component/leaf-node/max-depth statistics +- The full list of generated files + +## Directory Layout Convention + +`DocumentationGenerator` enforces a strict hierarchical output convention via `_get_nested_working_dir`: every module — leaf or parent, at any depth — gets its own subdirectory named after itself, containing a same-named Markdown file: + +```text +docs_dir/ +├── overview.md # repository-level overview +├── module_tree.json # final clustered/processed module tree +├── first_module_tree.json # initial clustering result (cache) +├── metadata.json # generation metadata +├── Backend/ +│ ├── Backend.md # parent module overview +│ ├── Authentication/ +│ │ ├── Authentication.md +│ │ ├── JWT/ +│ │ │ └── JWT.md # leaf module docs +│ │ └── OAuth/ +│ │ └── OAuth.md +│ └── API/ +│ └── API.md +``` + +This matches the same hierarchical linking convention used across all generated module documentation in this system. + +## Error Handling and Resilience + +- **Per-module isolation**: Exceptions raised while processing an individual module in `generate_module_documentation` are caught, logged with a traceback, and the loop continues — a single failing module does not abort the entire documentation run. +- **Idempotent caching**: Both `process_module` (in `AgentOrchestrator`) and `generate_parent_module_docs` check for existing output files before invoking the LLM, allowing interrupted runs to be resumed cheaply. +- **Context-overflow prevention**: The synthetic module fallback in `run()` avoids the case where an empty clustering result would force the entire repository through a single oversized LLM call. +- **Tag-tolerant parsing**: `generate_parent_module_docs` gracefully handles LLM responses that omit the expected `` wrapper tags by falling back to the full raw response instead of failing. + +## Related Modules + +- [Backend Core](../../backend-core.md) — parent module; hosts `AgentOrchestrator`, agent tools, and the dependency analysis subsystem this module depends on +- [Dependency Analyzer Core](../dependency-analyzer-core/dependency-analyzer-core.md) — provides `DependencyGraphBuilder` used to parse the repository +- [Tree-sitter Analyzers](../tree-sitter-analyzers/tree-sitter-analyzers.md) — language-specific parsers invoked during dependency graph construction +- [Dependency Analyzer Models](../dependency-analyzer-models/dependency-analyzer-models.md) — `Node` and related data models flowing through this pipeline +- [Agent Tools Core](../agent-tools-core/agent-tools-core.md) — tools (`EditTool`, `Filemap`, `CodeWikiDeps`) used by agents during leaf-module generation +- [LLM Services](../llm-services/llm-services.md) — `CountingFallbackModel` and `call_llm`, used both directly (for parent/overview synthesis) and indirectly (via `AgentOrchestrator`) for leaf modules +- [Config Core](../../../config-core.md) — the `Config` object supplied to `DocumentationGenerator` on construction +- [CLI Core](../../../cli-core.md) — command-line entry point that constructs `Config` and invokes `DocumentationGenerator.run()` +- [Frontend Core](../../../frontend-core.md) — web application entry point (`BackgroundWorker`) that runs the same generation pipeline for submitted repositories diff --git a/docs/reference/architecture/backend-core/llm-services/llm-services.md b/docs/reference/architecture/backend-core/llm-services/llm-services.md new file mode 100644 index 00000000..965ca114 --- /dev/null +++ b/docs/reference/architecture/backend-core/llm-services/llm-services.md @@ -0,0 +1,224 @@ +# Llm Services + +The Llm Services module is the central factory and runtime layer for all Large Language Model (LLM) interactions in CodeWiki's backend. It builds configured `pydantic-ai` model instances (main, fallback, and cluster models) from a [Config](../../config-core.md) object, wraps them with automatic failover and request-counting behavior, and exposes a low-level synchronous `call_llm()` helper for direct OpenAI-compatible API calls. Every LLM-driven component in the backend — the [Documentation Generator](../documentation-generator/documentation-generator.md) and the agent orchestration layer that powers it — depends on this module to obtain ready-to-use model clients without needing to know provider-specific details (base URLs, API keys, temperature support, or token-limit parameter naming). + +## Purpose and Scope + +CodeWiki supports multiple LLM providers (OpenAI, Anthropic, and OpenAI-compatible endpoints) across three distinct operational stages: + +- **Main/generation** — the primary model used to write module documentation. +- **Fallback** — a secondary model automatically invoked if the main model fails. +- **Cluster** — a (usually cheaper/faster) model used for the module-clustering stage of documentation generation. + +Because each stage can use a different provider with different base URLs, API keys, temperature support, and token-limit parameter names (`max_tokens` vs. `max_completion_tokens` for reasoning models like o3/o3-mini), the Llm Services module centralizes this provider-resolution logic in one place so the rest of the backend can remain provider-agnostic. + +## Core Responsibilities + +1. **Model factory functions** — build `pydantic-ai` `OpenAIModel` instances per stage (`create_main_model`, `create_fallback_model`, `create_cluster_model`), each validating that the required `base_url` and `api_key` are present in [Config](../../config-core.md) and raising descriptive `ValueError`s otherwise. +2. **Fallback chain assembly** — `create_fallback_models()` combines the main and fallback models into a single `CountingFallbackModel`, the core component of this module, which `pydantic-ai`'s `Agent` uses transparently: if the main model errors, the fallback model is tried automatically. +3. **Request counting/telemetry** — a module-level counter (`_request_counter`) tracks how many LLM requests have been made for the current module being documented, logging progress every 100 requests via `increment_request_counter()` / `reset_request_counter()`. +4. **Direct/synchronous LLM calls** — `call_llm()` provides a simpler, non-agentic path for callers (such as the parent-module overview generation flow in [Documentation Generator](../documentation-generator/documentation-generator.md)) that just need a single prompt/response round-trip using the raw `openai` Python SDK client built by `create_openai_client()`. +5. **Token-limit and temperature normalization** — `get_max_output_tokens()` and `get_model_max_token_field()` resolve environment-variable overrides (`MAX_OUTPUT_TOKENS`, `CODEWIKI_*_MAX_TOKEN_FIELD`) so reasoning models that require `max_completion_tokens` instead of `max_tokens` work without code changes. + +## Core Component + +### CountingFallbackModel + +`CountingFallbackModel` extends `pydantic-ai`'s `FallbackModel` and overrides `request()` to call `increment_request_counter()` before delegating to the parent implementation. This gives CodeWiki visibility into LLM usage volume per module without modifying `pydantic-ai` internals or scattering counting logic throughout the agent code. It is the object returned by `create_fallback_models()` and passed directly to `pydantic-ai.Agent` by the [Agent Orchestrator](../documentation-generator/documentation-generator.md) (part of `AgentOrchestrator.create_agent`), meaning **every** agent-driven LLM call in CodeWiki flows through this wrapper. + +```python +class CountingFallbackModel(FallbackModel): + """FallbackModel wrapper that counts and logs requests every 100 calls.""" + + async def request(self, *args, **kwargs): + """Wrap request to count calls.""" + increment_request_counter() + return await super().request(*args, **kwargs) +``` + +## Architecture + +```mermaid +flowchart TD + Config["Config"] --> CreateMain["create_main_model()"] + Config --> CreateFallback["create_fallback_model()"] + Config --> CreateCluster["create_cluster_model()"] + Config --> CreateClient["create_openai_client()"] + + CreateMain --> MainModel["OpenAIModel (main)"] + CreateFallback --> FallbackModel_["OpenAIModel (fallback)"] + + MainModel --> CreateChain["create_fallback_models()"] + FallbackModel_ --> CreateChain + + CreateChain --> CountingModel["CountingFallbackModel"] + CountingModel --> IncCounter["increment_request_counter()"] + IncCounter --> Counter[("_request_counter")] + + CountingModel --> Agent["pydantic-ai Agent"] + + CreateCluster --> ClusterModel["OpenAIModel (cluster)"] + ClusterModel --> ClusterAgent["Clustering LLM calls"] + + CreateClient --> OpenAIClient["openai.OpenAI client"] + OpenAIClient --> CallLLM["call_llm()"] + CallLLM --> Response["LLM response text"] +``` + +## How the Module Fits Into the Backend + +```mermaid +flowchart LR + subgraph LlmServices["Llm Services"] + Factory["Model factories
(main / fallback / cluster)"] + Counting["CountingFallbackModel"] + DirectCall["call_llm()"] + end + + subgraph DocGen["Documentation Generator"] + DG["DocumentationGenerator"] + Orchestrator["AgentOrchestrator"] + end + + ConfigCore["Config"] --> Factory + Factory --> Counting + Orchestrator -->|"create_fallback_models(config)"| Factory + Orchestrator -->|"used by pydantic-ai Agent"| Counting + DG -->|"call_llm(prompt, config)"| DirectCall + + DirectCall --> OpenAISDK["openai.OpenAI client"] + Counting --> PydanticAI["pydantic-ai FallbackModel / OpenAIProvider"] +``` + +*[Documentation Generator](../documentation-generator/documentation-generator.md) owns the `AgentOrchestrator`, which calls `create_fallback_models(config)` once per module to build the agent's LLM backend, and calls `call_llm()` directly when generating parent-module overview documentation (bypassing the agent framework for simpler prompt/response interactions).* + +## Model Creation Flow + +Each `create_*_model()` function follows the same validation-and-build pattern, differing only in which per-provider config fields it reads (`main_*`, `fallback_*`, or `cluster_*`): + +```mermaid +sequenceDiagram + participant Caller + participant Factory as "create_main_model() / create_fallback_model() / create_cluster_model()" + participant Config as "Config" + participant Provider as "OpenAIProvider" + participant Model as "OpenAIModel" + + Caller->>Factory: create_X_model(config) + Factory->>Config: getattr(config, "X_max_tokens", None) + Factory->>Config: getattr(config, "X_temperature", 0.0) + Factory->>Config: getattr(config, "X_temperature_supported", True) + Factory->>Config: getattr(config, "X_base_url", None) + alt base_url missing + Factory-->>Caller: raise ValueError("X_base_url is required...") + end + Factory->>Config: getattr(config, "X_api_key", None) + alt api_key missing + Factory-->>Caller: raise ValueError("X_api_key is required...") + end + Factory->>Provider: OpenAIProvider(base_url, api_key) + Factory->>Model: OpenAIModel(model_name, provider, settings) + Model-->>Caller: configured OpenAIModel +``` + +## Fallback Chain Construction + +```mermaid +sequenceDiagram + participant Orchestrator as "AgentOrchestrator" + participant LlmServices as "llm_services" + participant MainModel as "OpenAIModel (main)" + participant FallbackModelObj as "OpenAIModel (fallback)" + participant CFM as "CountingFallbackModel" + + Orchestrator->>LlmServices: create_fallback_models(config) + LlmServices->>LlmServices: create_main_model(config) + LlmServices-->>MainModel: main + LlmServices->>LlmServices: create_fallback_model(config) + LlmServices-->>FallbackModelObj: fallback + LlmServices->>CFM: CountingFallbackModel(main, fallback) + CFM-->>Orchestrator: fallback_models + + Note over Orchestrator,CFM: Passed as the model backend
for pydantic-ai Agent +``` + +At runtime, when `pydantic-ai`'s `Agent.run()` invokes `CountingFallbackModel.request()`: + +1. `increment_request_counter()` runs first, bumping the module-scoped counter and logging a progress line every 100 requests. +2. The call is delegated to `FallbackModel.request()` (the parent class), which attempts the **main** model first and automatically retries with the **fallback** model if the main model raises an error. + +## Direct LLM Invocation (`call_llm`) + +For flows that do not need the full `pydantic-ai` agent/tool-calling machinery — such as generating a parent-module overview from already-generated child documentation — the module exposes a synchronous `call_llm()` helper built directly on the `openai` SDK. + +```mermaid +sequenceDiagram + participant Caller + participant CallLLM as "call_llm()" + participant ClientFactory as "create_openai_client()" + participant Config as "Config" + participant OpenAISDK as "openai.OpenAI" + + Caller->>CallLLM: call_llm(prompt, config, model, temperature) + CallLLM->>CallLLM: resolve stage (main / cluster / fallback) + CallLLM->>ClientFactory: create_openai_client(config, model) + ClientFactory->>Config: resolve base_url / api_key / api_version for stage + alt base_url or api_key missing + ClientFactory-->>CallLLM: raise ValueError(...) + end + ClientFactory->>OpenAISDK: OpenAI(base_url, api_key, default_headers) + OpenAISDK-->>CallLLM: client + CallLLM->>CallLLM: get_model_max_token_field(stage) + CallLLM->>OpenAISDK: client.chat.completions.create(**kwargs) + alt success + OpenAISDK-->>CallLLM: response + CallLLM-->>Caller: response_content + else OpenAIError + CallLLM-->>Caller: raise RuntimeError(context-wrapped) + end +``` + +Key behaviors of `call_llm()`: + +- **Stage detection** — determines whether `model` matches `config.cluster_model`, `config.fallback_model`, or defaults to main/generation, so it can select the correct `base_url`, `api_key`, `api_version`, `max_tokens` value, temperature support flag, and token-field name for that specific provider. +- **Dynamic token-field naming** — uses `get_model_max_token_field(stage)` to decide whether to send `max_tokens` or `max_completion_tokens` in the request payload, accommodating reasoning models (o3, o3-mini) that reject the standard `max_tokens` parameter. +- **Conditional temperature** — only includes `temperature` in the request if the resolved `*_temperature_supported` config flag is `True`, since some reasoning models reject custom temperature values entirely. +- **Rich error context** — wraps both `OpenAIError` and generic exceptions in a `RuntimeError` that includes the stage, model, base URL, temperature, and max-token settings to aid debugging misconfigured providers. +- **Structured logging** — emits detailed, tree-formatted log lines (stage, model, base URL, prompt length/preview, temperature, token settings) both before the request and after a successful response, easing observability during large documentation runs. + +## Environment-Driven Configuration Helpers + +Two helper functions decouple token-limit behavior from hardcoded values, reading overrides from environment variables at call time: + +| Function | Environment Variable(s) | Purpose | +|---|---|---| +| `get_max_output_tokens()` | `MAX_OUTPUT_TOKENS` | Returns the default max output tokens (16384) unless overridden; used as a fallback when a stage-specific `*_max_tokens` field is not set on [Config](../../config-core.md). | +| `get_model_max_token_field(stage)` | `CODEWIKI_CLUSTER_MAX_TOKEN_FIELD`, `CODEWIKI_GENERATION_MAX_TOKEN_FIELD`, `CODEWIKI_FALLBACK_MAX_TOKEN_FIELD` | Returns `max_tokens` (default) or `max_completion_tokens` for the given stage, letting operators switch reasoning-model support without code changes. | + +## Relationship to Configuration + +All factory functions and `call_llm()` read their provider settings exclusively from the [Config](../../config-core.md) object — specifically its per-provider fields (`main_base_url`, `main_api_key`, `main_temperature`, `cluster_base_url`, `cluster_api_key`, `fallback_base_url`, `fallback_api_key`, etc.). This module performs no environment-variable parsing of its own for credentials; all credential/URL resolution responsibility lives in `Config.from_args()`, `Config.from_cli()`, and `Config.from_web_job()`. Llm Services only reads the resulting per-provider attributes via `getattr()`, defaulting gracefully where sensible (e.g., temperature defaults to `0.0`, token limits default to `get_max_output_tokens()`). + +## Error Handling Philosophy + +Every factory function fails fast with a descriptive `ValueError` when required configuration (`base_url`, `api_key`) is missing, naming the exact CLI flag or config-file field the caller should set: + +```python +raise ValueError( + "main_base_url is required in configuration for main/generation model.\n" + f"Model: {config.main_model}\n" + "Please set via CLI: --main-base-url \n" + "Or in config file: main_base_url = ''" +) +``` + +This "actionable error message" pattern is applied consistently across `create_main_model`, `create_fallback_model`, `create_cluster_model`, and `create_openai_client`, minimizing debugging time when a user misconfigures a provider for one of the three stages. + +## Summary + +The Llm Services module is a thin but critical abstraction layer that: + +- Converts a single [Config](../../config-core.md) object into fully-configured, provider-agnostic `pydantic-ai` models for three distinct LLM stages (main, fallback, cluster). +- Provides `CountingFallbackModel` as the automatic-failover, telemetry-instrumented model backend used by every `pydantic-ai` `Agent` created by the [Documentation Generator](../documentation-generator/documentation-generator.md)'s agent orchestration layer. +- Offers a simpler synchronous `call_llm()` path for non-agentic prompt/response use cases, with the same provider-resolution and error-handling guarantees. +- Normalizes cross-provider quirks (reasoning-model token-field naming, optional temperature support) so callers never need provider-specific branching logic. diff --git a/docs/reference/architecture/backend-core/logging-config/logging-config.md b/docs/reference/architecture/backend-core/logging-config/logging-config.md new file mode 100644 index 00000000..d6a398f1 --- /dev/null +++ b/docs/reference/architecture/backend-core/logging-config/logging-config.md @@ -0,0 +1,181 @@ +# Logging Config + +The Logging Config module provides a colorized console logging formatter and helper functions used across the CodeWiki backend to produce readable, severity-differentiated log output. It centers on the `ColoredFormatter` class, a custom subclass of Python's standard `logging.Formatter` that decorates log records with ANSI colors based on log level, and two convenience setup functions (`setup_logging` and `setup_module_logging`) that wire the formatter into console handlers for the root logger or a specific named logger. + +This module is a small, focused utility that other backend components depend on for consistent, human-friendly terminal logging during dependency analysis, documentation generation, and related long-running backend operations. + +## Purpose and Scope + +Backend processes such as repository analysis, call graph construction, and documentation generation emit a large volume of log messages while processing potentially large codebases. Plain, uncolored log output makes it difficult to visually scan for warnings and errors in a busy terminal. The Logging Config module solves this by: + +- Coloring log messages according to severity (DEBUG, INFO, WARNING, ERROR, CRITICAL) +- Coloring timestamps and other structural elements distinctly from the message body +- Providing simple, one-call setup functions so any part of the backend can enable colored logging without repeating formatter/handler boilerplate +- Ensuring cross-platform compatibility (including Windows terminals) via the `colorama` library + +## Core Component + +### ColoredFormatter + +`ColoredFormatter` extends `logging.Formatter` and overrides the `format(record)` method to inject ANSI color codes into the rendered log line. + +**Color scheme:** + +| Log Level / Element | Color | +|---|---| +| DEBUG | Blue | +| INFO | Cyan | +| WARNING | Yellow | +| ERROR | Red | +| CRITICAL | Red + Bright | +| Timestamp | Blue | +| Reset | Style reset (no color) | + +Two internal class-level dictionaries drive this behavior: + +- `COLORS`: maps `record.levelname` (e.g. `"INFO"`, `"ERROR"`) to a `colorama.Fore` color code +- `COMPONENT_COLORS`: maps structural elements (`timestamp`, `module`, `reset`) to their respective colors + +The `format()` method builds the final log line by: +1. Looking up the color for the record's level (falling back to no color if the level is unrecognized) +2. Formatting the timestamp (`HH:MM:SS`) and wrapping it in the timestamp color +3. Formatting the level name (left-padded to 8 characters) in the level's color +4. Formatting the message text in the same color as the level, for visual consistency +5. Concatenating timestamp, level, and message into a single colored line +6. Appending formatted exception traceback text (uncolored) if the record carries exception info + +```mermaid +flowchart TD + A["logging.LogRecord"] --> B["ColoredFormatter.format(record)"] + B --> C["Look up level color in COLORS"] + B --> D["Format timestamp HH:MM:SS"] + D --> E["Wrap timestamp in blue"] + C --> F["Wrap levelname in level color"] + C --> G["Wrap message in level color"] + E --> H["Concatenate: timestamp + level + message"] + F --> H + G --> H + H --> I{"record.exc_info present?"} + I -->|"Yes"| J["Append formatException() output"] + I -->|"No"| K["Return colored log line"] + J --> K +``` + +## Setup Functions + +Alongside `ColoredFormatter`, the module exposes two helper functions that configure logging handlers using the formatter. + +### setup_logging(level=logging.INFO) + +Configures the **root logger** for the entire application: +1. Creates a `logging.StreamHandler` writing to `sys.stdout` +2. Attaches a `ColoredFormatter` instance to the handler +3. Clears any existing handlers on the root logger (to avoid duplicate output when called more than once) +4. Sets the root logger's level and attaches the new handler + +This is intended to be called once, early in a process's lifecycle (e.g., at the start of a CLI or backend service run), to enable colored output application-wide. + +### setup_module_logging(module_name, level=logging.INFO) + +Configures a **named logger** for a specific module rather than the root logger: +1. Retrieves (or creates) a logger via `logging.getLogger(module_name)` +2. Creates a `StreamHandler` to `stdout` with a `ColoredFormatter` +3. Clears existing handlers on that logger +4. Sets `logger.propagate = False` to prevent messages from bubbling up to the root logger (which would otherwise cause duplicate log lines) +5. Returns the configured logger for direct use + +This allows individual backend components — for example an analyzer or the analysis service — to have isolated, independently configured colored logging without interfering with (or being interfered by) the root logger's configuration. + +```mermaid +sequenceDiagram + participant Caller as "Backend Component" + participant Setup as "setup_module_logging()" + participant Logger as "logging.Logger" + participant Handler as "StreamHandler(stdout)" + participant Fmt as "ColoredFormatter" + + Caller->>Setup: setup_module_logging("my_module", level) + Setup->>Logger: logging.getLogger("my_module") + Setup->>Handler: create StreamHandler(sys.stdout) + Setup->>Fmt: create ColoredFormatter() + Setup->>Handler: setFormatter(Fmt) + Setup->>Logger: handlers.clear() + Setup->>Logger: addHandler(Handler) + Setup->>Logger: propagate = False + Setup-->>Caller: return configured logger + Caller->>Logger: logger.info("message") + Logger->>Handler: emit(record) + Handler->>Fmt: format(record) + Fmt-->>Handler: colored log line + Handler-->>Caller: printed to stdout +``` + +## Class Structure + +```mermaid +classDiagram + class Formatter { + <> + +format(record) + +formatTime(record, datefmt) + +formatException(exc_info) + } + class ColoredFormatter { + +COLORS : dict + +COMPONENT_COLORS : dict + +format(record) str + } + Formatter <|-- ColoredFormatter +``` + +## Integration with the Backend + +The Logging Config module is a leaf utility within the broader backend codebase. It is imported wherever colored console output is desired, most notably by components that perform long-running, verbose operations such as dependency graph analysis and repository scanning. Because it only depends on the Python standard library `logging` module and `colorama`, it can be adopted independently by any backend component without introducing coupling to other subsystems. + +Within the backend hierarchy, this module sits alongside sibling utility and domain modules such as [Dependency Analyzer Core](dependency-analyzer-core/dependency-analyzer-core.md), [Tree-Sitter Analyzers](tree-sitter-analyzers/tree-sitter-analyzers.md), [Dependency Analyzer Models](dependency-analyzer-models/dependency-analyzer-models.md), [Documentation Generator](documentation-generator/documentation-generator.md), and [LLM Services](llm-services/llm-services.md), all of which are children of the top-level backend module. + +Note that the CLI package maintains its own separate logging utility (`CLILogger`) for command-line output; the Logging Config module described here is specific to the backend's internal logging needs and is not shared code with the CLI's logging utilities. + +```mermaid +graph TD + Backend["Backend Core"] --> LoggingConfig["Logging Config"] + Backend --> DepAnalyzer["Dependency Analyzer Core"] + Backend --> TreeSitter["Tree-Sitter Analyzers"] + Backend --> DepModels["Dependency Analyzer Models"] + Backend --> DocGen["Documentation Generator"] + Backend --> LLMServices["LLM Services"] + Backend --> AgentTools["Agent Tools Core"] + + DepAnalyzer -.->|"colored console output"| LoggingConfig + TreeSitter -.->|"colored console output"| LoggingConfig + DocGen -.->|"colored console output"| LoggingConfig +``` + +## Usage Pattern + +Typical usage within a backend component follows one of two patterns: + +**Application-wide setup** (once, at process start): + +```python +from codewiki.src.be.dependency_analyzer.utils.logging_config import setup_logging +import logging + +setup_logging(level=logging.INFO) +``` + +**Per-module isolated setup:** + +```python +from codewiki.src.be.dependency_analyzer.utils.logging_config import setup_module_logging +import logging + +logger = setup_module_logging(__name__, level=logging.DEBUG) +logger.debug("Starting analysis...") +``` + +Both patterns rely on the same underlying `ColoredFormatter` to ensure consistent color-coded output regardless of which setup function is used. + +## Relationship to the Parent Module + +This module is part of the backend codebase and is documented as a child of [Backend Core](../backend-core.md). It has no further child modules of its own, as its single component (`ColoredFormatter`) and the accompanying setup functions form a complete, self-contained unit of functionality. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md new file mode 100644 index 00000000..290b37c3 --- /dev/null +++ b/docs/reference/architecture/backend-core/tree-sitter-analyzers/c_family_analyzers.md @@ -0,0 +1,260 @@ +# C Family Analyzers + +## Introduction + +The C Family Analyzers module provides static-analysis front-ends for three closely related C-style languages: **C**, **C++**, and **C#**. Each analyzer parses a single source file with a language-specific [tree-sitter](https://tree-sitter.github.io/tree-sitter/) grammar, walks the resulting concrete syntax tree, and produces two normalized outputs: + +- **Nodes** — structural components such as functions, classes, structs, methods, namespaces, and global variables +- **Call Relationships** — edges describing how those components interact (calls, inheritance, instantiation, field/property usage) + +These outputs conform to the shared `Node` and `CallRelationship` data models used across the entire dependency-analysis pipeline, allowing the C-family analyzers to be interchangeable with analyzers for other languages such as Java, Python, PHP, and JavaScript/TypeScript. + +The module is a leaf-level component in the analyzer layer: it has no knowledge of cross-file resolution, graph construction, or clustering — that responsibility belongs to higher-level orchestration components described later in this document. + +## Purpose and Scope + +| Analyzer | Component | Source Extensions | Grammar Package | +|---|---|---|---| +| C | `TreeSitterCAnalyzer` | `.c`, `.h` | `tree_sitter_c` | +| C++ | `TreeSitterCppAnalyzer` | `.cpp`, `.cc`, `.cxx`, `.hpp`, `.h` | `tree_sitter_cpp` | +| C# | `TreeSitterCSharpAnalyzer` | `.cs` | `tree_sitter_c_sharp` | + +Each analyzer is instantiated per-file with the file path, raw source content, and (optionally) the repository root path. Construction is eager: the constructor immediately runs the full analysis pipeline (`_analyze()`), after which the caller reads the populated `nodes` and `call_relationships` lists. + +```python +analyzer = TreeSitterCAnalyzer(file_path, content, repo_path) +functions_and_structs = analyzer.nodes +relationships = analyzer.call_relationships +``` + +Each analyzer file also exposes a thin module-level convenience function (`analyze_c_file`, `analyze_cpp_file`, `analyze_csharp_file`) that wraps this construction pattern and returns a `(nodes, call_relationships)` tuple — this is the entry point used by upstream orchestration. + +## Architecture + +All three analyzers share an identical structural pattern despite differences in the underlying grammars: a two-pass traversal (node extraction, then relationship extraction) built on top of tree-sitter's parse tree. + +```mermaid +flowchart TD + Source["Source File Content"] --> Parser["tree-sitter Parser"] + Parser --> Tree["Concrete Syntax Tree"] + Tree --> Extract["Node Extraction Pass"] + Extract --> TopLevel["Top-Level Node Registry"] + TopLevel --> Rel["Relationship Extraction Pass"] + Rel --> Nodes["List of Node objects"] + Rel --> Calls["List of CallRelationship objects"] +``` + +### Common Construction Flow + +1. **Language binding** — Each analyzer loads its grammar via the corresponding `tree_sitter_*` package and wraps it in a `Language`/`Parser` pair. +2. **Parsing** — The raw UTF-8 file content is parsed once into a `tree_sitter` syntax tree. +3. **Node extraction (`_extract_nodes`)** — A recursive depth-first traversal identifies structurally significant syntax node types (e.g. `function_definition`, `class_specifier`, `struct_specifier`) and converts them into a shared `Node` model, while also registering them in a local `top_level_nodes` dictionary keyed by name (or qualified name for C++ methods). +4. **Relationship extraction (`_extract_relationships`)** — A second recursive traversal inspects call expressions, inheritance clauses, instantiation expressions, and identifier usages, cross-referencing them against `top_level_nodes` to emit `CallRelationship` edges. + +### Component ID Convention + +Every analyzer derives a fully-qualified component identifier from the file's path relative to the repository root, converting path separators to dots and stripping the language-specific file extension: + +```text +path/to/file.cpp -> path.to.file +component_id = "path.to.file::FunctionName" +``` + +For C++ methods, the identifier additionally embeds the containing class using dot notation: `path.to.file::ClassName.methodName`. This mirrors the identifier scheme used by the other language analyzers so that identifiers remain comparable across the whole codebase graph. + +## Component Details + +### TreeSitterCAnalyzer + +Parses C source using the `tree_sitter_c` grammar. It recognizes: + +- **Functions** (`function_definition`) — extracted via the `function_declarator` → `identifier` chain +- **Structs** (`struct_specifier`, and `typedef struct { ... } Name;` via `type_definition`) +- **Global variables** (`declaration` nodes not nested inside a function body) + +Relationship extraction covers two cases: +- **Function calls** (`call_expression`) — the callee is recorded by its *simple name only* (`is_resolved=False`), deferring cross-file resolution to the call-graph analyzer described below. A hardcoded set of common C-standard-library and SDL functions (`printf`, `malloc`, `SDL_Init`, etc.) is filtered out to avoid noise. +- **Global variable usage** — identifiers inside a function body that match a known global variable are recorded as resolved, same-file relationships (`is_resolved=True`). + +Only `function` and `struct` node types are appended to the public `nodes` list; global `variable` nodes are tracked internally (in `top_level_nodes`) purely to support usage-relationship detection. + +### TreeSitterCppAnalyzer + +Parses C++ using the `tree_sitter_cpp` grammar. It extends the C model with object-oriented and namespace constructs: + +- **Classes and structs** (`class_specifier`, `struct_specifier`) +- **Functions and methods** (`function_definition`) — distinguished by walking up the parent chain to detect an enclosing `class_specifier`/`struct_specifier`; methods are keyed in `top_level_nodes` by `ClassName.methodName` +- **Namespaces** (`namespace_definition`) +- **Global variables** (`declaration` nodes outside any function/class/struct body) + +Relationship extraction is the richest of the three analyzers, detecting: + +| Relationship Type | Triggering Syntax | Notes | +|---|---|---| +| `calls` | `call_expression` | Resolves plain function calls and, for method calls via `field_expression`, attempts to locate the owning class through `_find_class_containing_method` | +| `inherits` | `base_class_clause` | Extracts base `type_identifier` names | +| `creates` | `new_expression` | Detects object instantiation (`new ClassName(...)`) | +| `uses` | bare `identifier` | Detects references to global variables from within functions/methods | + +A hardcoded system-function filter (`printf`, `cout`, `new`, `delete`, etc.) suppresses standard-library noise, mirroring the C analyzer. + +### TreeSitterCSharpAnalyzer + +Parses C# using the `tree_sitter_c_sharp` grammar. Its node vocabulary is the broadest of the three, reflecting C#'s richer type-declaration surface: + +- **Classes** — further classified as `class`, `abstract class`, or `static class` based on detected `modifier` nodes +- **Interfaces** (`interface_declaration`) +- **Structs** (`struct_declaration`) +- **Enums** (`enum_declaration`) +- **Records** (`record_declaration`) +- **Delegates** (`delegate_declaration`) + +Unlike the C and C++ analyzers, the C# analyzer does **not** attempt call-expression resolution. Instead, it focuses on **type-usage relationships** that reflect C#'s declarative, strongly-typed structure: + +| Relationship Type | Triggering Syntax | Notes | +|---|---|---| +| inheritance/implementation | `class_declaration` → `base_list` | Emits a resolved relationship (`is_resolved=True`) only when the base name matches another top-level node in the same file | +| property type usage | `property_declaration` | Emits an unresolved relationship (`is_resolved=False`) when the property type is non-primitive | +| field type usage | `field_declaration` | Same pattern as properties | +| parameter type usage | `method_declaration` → `parameter_list` | Emits an unresolved relationship per non-primitive parameter type | + +A `_is_primitive_type` allow-list (C# built-ins like `int`, `string`, `List`, `Dictionary`, `Task`, `DateTime`, etc.) prevents common framework types from polluting the relationship graph. + +## Data Model + +All three analyzers populate the shared `Node` and `CallRelationship` structures defined in the dependency-analyzer models layer. Key fields populated by every C-family analyzer include: + +- `Node`: `id`, `name`, `component_type`/`node_type`, `file_path`, `relative_path`, `source_code`, `start_line`, `end_line`, `display_name`, `component_id`, and (for C++ methods) `class_name` +- `CallRelationship`: `caller`, `callee`, `call_line`, `is_resolved`, and optionally `relationship_type` (C++ only: `calls`, `inherits`, `creates`, `uses`) + +Notably, the three analyzers differ in how aggressively they resolve `is_resolved`: +- The **C** analyzer defers function-call resolution entirely (`is_resolved=False`), but resolves same-file variable usage. +- The **C++** analyzer resolves relationships whenever a callee/base-class name matches a `top_level_nodes` entry in the same file. +- The **C#** analyzer resolves inheritance within the same file but leaves type-usage relationships unresolved, since types may be declared elsewhere. + +For full schema definitions, see the dependency analyzer models. + +## Process Flow: Node and Relationship Extraction + +The following sequence illustrates the two-pass traversal shared by all three analyzers, using the C++ analyzer as a representative example: + +```mermaid +sequenceDiagram + participant Caller as "analyze_cpp_file()" + participant Analyzer as "TreeSitterCppAnalyzer" + participant TS as "tree-sitter Parser" + participant Pass1 as "_extract_nodes()" + participant Pass2 as "_extract_relationships()" + + Caller->>Analyzer: "__init__(file_path, content, repo_path)" + Analyzer->>TS: "parse(content)" + TS-->>Analyzer: "syntax tree" + Analyzer->>Pass1: "traverse(root, top_level_nodes, lines)" + Pass1->>Pass1: "detect class/struct/function/namespace/variable" + Pass1-->>Analyzer: "populated top_level_nodes + nodes list" + Analyzer->>Pass2: "traverse(root, top_level_nodes)" + Pass2->>Pass2: "detect calls, inheritance, instantiation, usage" + Pass2-->>Analyzer: "call_relationships list" + Analyzer-->>Caller: "nodes, call_relationships" +``` + +## Component Relationships + +```mermaid +classDiagram + class TreeSitterCAnalyzer { + +file_path + +content + +repo_path + +nodes + +call_relationships + -_analyze() + -_extract_nodes() + -_extract_relationships() + } + class TreeSitterCppAnalyzer { + +file_path + +content + +repo_path + +nodes + +call_relationships + -_analyze() + -_extract_nodes() + -_extract_relationships() + } + class TreeSitterCSharpAnalyzer { + +file_path + +content + +repo_path + +nodes + +call_relationships + -_analyze() + -_extract_nodes() + -_extract_relationships() + } + class Node { + +id + +name + +component_type + +file_path + +source_code + } + class CallRelationship { + +caller + +callee + +call_line + +is_resolved + } + TreeSitterCAnalyzer --> Node : produces + TreeSitterCAnalyzer --> CallRelationship : produces + TreeSitterCppAnalyzer --> Node : produces + TreeSitterCppAnalyzer --> CallRelationship : produces + TreeSitterCSharpAnalyzer --> Node : produces + TreeSitterCSharpAnalyzer --> CallRelationship : produces +``` + +## Integration with the Dependency Analysis Pipeline + +The C Family Analyzers do not run in isolation. They are invoked per-file by the language-dispatching parser, and their raw (sometimes unresolved) output is later consolidated by the call-graph resolution stage: + +```mermaid +flowchart LR + Repo["Repository Source Files"] --> Parser["Dependency Parser"] + Parser -->|".c / .h"| CAnalyzer["TreeSitterCAnalyzer"] + Parser -->|".cpp / .hpp"| CppAnalyzer["TreeSitterCppAnalyzer"] + Parser -->|".cs"| CSharpAnalyzer["TreeSitterCSharpAnalyzer"] + CAnalyzer --> RawNodes["Raw Nodes + Relationships"] + CppAnalyzer --> RawNodes + CSharpAnalyzer --> RawNodes + RawNodes --> GraphBuilder["Dependency Graph Builder"] + RawNodes --> CallGraph["Call Graph Analyzer"] + CallGraph --> Resolved["Resolved Cross-File Relationships"] + GraphBuilder --> FinalGraph["Repository Dependency Graph"] + Resolved --> FinalGraph +``` + +- The **Dependency Parser** (`DependencyParser`) selects the appropriate analyzer based on file extension and orchestrates per-file invocation across the whole repository. +- The **Call Graph Analyzer** (`CallGraphAnalyzer`) consumes the unresolved `callee` names emitted by these analyzers (particularly from the C and C++ analyzers) and resolves them into fully-qualified cross-file identifiers. +- The **Dependency Graph Builder** (`DependencyGraphBuilder`) assembles the final `Node`/`CallRelationship` collections from all language analyzers — including this module — into the unified repository dependency graph. + +These orchestration components live in the backend's dependency analyzer core, which handles graph construction and the analysis pipeline for details on how per-file analyzer output is consolidated and resolved. + +The shared `Node`, `CallRelationship`, and related schema types used by all three analyzers are defined in the dependency analyzer models. + +## Relationship to Other Language Analyzers + +The C Family Analyzers module is one of several sibling analyzer groups under [Tree Sitter Analyzers](tree-sitter-analyzers.md), each following the same node/relationship extraction contract but tailored to a different language's grammar and idioms: + +- [Java Analyzer](java_analyzer.md) +- [JavaScript/TypeScript Analyzers](javascript_typescript_analyzers.md) +- [PHP Analyzer](php_analyzer.md) +- [Python Analyzer](python_analyzer.md) + +Because all analyzers emit the same `Node`/`CallRelationship` shapes, the downstream graph-construction and call-resolution logic remains language-agnostic — new language support can be added by implementing a new analyzer following this same pattern without modifying the orchestration layer. + +## Design Notes and Limitations + +- **Per-file, single-pass scope**: Each analyzer only sees one file at a time and has no visibility into other files in the repository. Cross-file symbol resolution (e.g. resolving a C function call to its definition in another translation unit) is intentionally deferred to the call-graph resolution stage. +- **Heuristic-based class/method matching**: The C++ analyzer's `_class_has_method` and `_find_class_containing_method` use lightweight source-text heuristics (substring matching on method signatures) rather than full semantic type resolution, since tree-sitter provides only syntactic — not semantic — information. +- **Standard-library filtering is hardcoded**: Both the C and C++ analyzers maintain fixed allow-lists of common standard-library/system function names to exclude from the relationship graph. This is a pragmatic simplification rather than a complete standard-library model. +- **C# favors type usage over call resolution**: Unlike C and C++, the C# analyzer does not attempt to trace method call expressions; it instead surfaces structural type dependencies (inheritance, field/property/parameter types), which are typically more informative for understanding C# codebases dominated by object-oriented composition. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md new file mode 100644 index 00000000..663d7f3e --- /dev/null +++ b/docs/reference/architecture/backend-core/tree-sitter-analyzers/java_analyzer.md @@ -0,0 +1,178 @@ +# Java Analyzer + +The Java Analyzer module is a language-specific static analysis component within CodeWiki's dependency analysis pipeline. It uses [tree-sitter](https://tree-sitter.github.io/tree-sitter/) to parse Java source files into an Abstract Syntax Tree (AST) and extracts both **structural components** (classes, interfaces, enums, records, annotations, and methods) and **relationships** between them (inheritance, interface implementation, field type usage, method invocations, and object instantiation). The resulting `Node` and `CallRelationship` objects feed into the broader dependency graph that CodeWiki uses to power documentation generation and repository visualization. + +This module is one of several language-specific analyzers registered under the [Tree Sitter Analyzers](tree-sitter-analyzers.md) module, sitting alongside analyzers for C/C++/C#, JavaScript/TypeScript, PHP, and Python. + +## Purpose and Scope + +The sole responsibility of the Java Analyzer is to convert raw Java source text into a set of normalized, language-agnostic data structures that downstream components (such as the dependency analyzer core and its graph construction pipeline) can consume without needing to know anything about Java syntax. + +Specifically, the analyzer: + +1. Parses a single `.java` file's source text using the `tree-sitter-java` grammar. +2. Walks the resulting AST to identify top-level and nested structural declarations. +3. Assigns each declaration a fully-qualified component ID in the `module.path::ClassName` (or `module.path::ClassName.methodName`) format expected by the rest of the system. +4. Walks the AST a second time to detect relationships between declarations (inheritance, interface implementation, field types, method calls, and object creation). +5. Exposes the extracted `Node` and `CallRelationship` lists for consumption by the parsing pipeline. + +## Core Component + +### TreeSitterJavaAnalyzer + +`TreeSitterJavaAnalyzer` (defined in `codewiki/src/be/dependency_analyzer/analyzers/java.py`) is the single class that implements the entire analysis workflow for one Java file. It is instantiated with the file's path, its content, and (optionally) the repository root path used to compute relative/module paths. + +```python +class TreeSitterJavaAnalyzer: + def __init__(self, file_path: str, content: str, repo_path: str = None): + ... + self._analyze() +``` + +On construction, the analyzer immediately runs its full pipeline (`_analyze()`), populating two public attributes: + +- `self.nodes: List[Node]` — the structural components discovered in the file. +- `self.call_relationships: List[CallRelationship]` — the relationships discovered between those components (and to external/unresolved types). + +A module-level convenience function is also provided: + +```python +def analyze_java_file(file_path: str, content: str, repo_path: str = None) -> Tuple[List[Node], List[CallRelationship]]: + analyzer = TreeSitterJavaAnalyzer(file_path, content, repo_path) + return analyzer.nodes, analyzer.call_relationships +``` + +This function is the primary entry point that callers in the dependency analysis pipeline are expected to use, mirroring the calling convention of the other tree-sitter based analyzers (C, C++, C#, JavaScript, TypeScript, PHP). + +## Architecture Overview + +The diagram below shows how `TreeSitterJavaAnalyzer` fits into the surrounding analysis pipeline and which data models it depends on. + +```mermaid +flowchart TD + subgraph Pipeline["Dependency Analysis Pipeline"] + CGA["CallGraphAnalyzer"] -->|"dispatches .java files"| JA["TreeSitterJavaAnalyzer"] + end + + subgraph JavaAnalyzerModule["Java Analyzer"] + JA -->|"1: parse source"| TS["tree-sitter-java grammar"] + TS -->|"AST"| EN["_extract_nodes()"] + TS -->|"AST"| ER["_extract_relationships()"] + EN -->|"appends"| Nodes["self.nodes: List[Node]"] + ER -->|"appends"| Rels["self.call_relationships: List[CallRelationship]"] + end + + Nodes -->|"consumed by"| DP["DependencyParser"] + Rels -->|"consumed by"| DP + DP -->|"builds"| Graph["Dependency Graph"] + + Node["Node model"] -.->|"defines schema for"| Nodes + CallRel["CallRelationship model"] -.->|"defines schema for"| Rels +``` + +- **CallGraphAnalyzer**, part of the dependency analyzer's analysis pipeline (which lives under the dependency analyzer core), dispatches individual source files to the appropriate language analyzer based on file extension. For `.java` files, it invokes `TreeSitterJavaAnalyzer`/`analyze_java_file`. +- **DependencyParser**, part of the dependency analyzer core's graph construction stage, aggregates the `Node` and `CallRelationship` objects produced by all analyzers (Java and otherwise) into a unified, namespaced dependency graph. +- **Node** and **CallRelationship** are shared Pydantic models defined in the dependency analyzer models. The Java Analyzer produces instances of these models but does not define them itself. + +## Internal Processing Flow + +`_analyze()` performs two independent, full-tree traversals over the same parsed AST: one for structural nodes, and one for relationships. This two-pass design ensures that all top-level declarations are known (via `top_level_nodes`) before relationships that reference them (e.g., method calls, field types) are resolved. + +```mermaid +sequenceDiagram + participant Caller as "analyze_java_file()" + participant Analyzer as "TreeSitterJavaAnalyzer" + participant Parser as "tree_sitter.Parser" + participant NodesPass as "_extract_nodes()" + participant RelsPass as "_extract_relationships()" + + Caller->>Analyzer: "__init__(file_path, content, repo_path)" + Analyzer->>Analyzer: "_analyze()" + Analyzer->>Parser: "parse(content)" + Parser-->>Analyzer: "AST root node" + Analyzer->>NodesPass: "_extract_nodes(root, top_level_nodes, lines)" + NodesPass-->>Analyzer: "self.nodes populated" + Analyzer->>RelsPass: "_extract_relationships(root, top_level_nodes)" + RelsPass-->>Analyzer: "self.call_relationships populated" + Analyzer-->>Caller: "nodes, call_relationships" +``` + +### Pass 1: Structural Node Extraction (`_extract_nodes`) + +This recursive method walks every node in the AST looking for the following Java declaration types: + +| AST Node Type | Resulting `component_type` | +|---|---| +| `class_declaration` (with `abstract` modifier) | `abstract class` | +| `class_declaration` (without `abstract` modifier) | `class` | +| `interface_declaration` | `interface` | +| `enum_declaration` | `enum` | +| `record_declaration` | `record` | +| `annotation_type_declaration` | `annotation` | +| `method_declaration` | `method` | + +For each match, the analyzer builds a fully-qualified `component_id` via `_get_component_id()`, which combines the file's module path (derived from its path relative to the repository root, with `/` replaced by `.`) and the declaration name using the `module.path::Name` convention (or `module.path::ClassName.methodName` for methods, where the containing class is resolved via `_find_containing_class_name`). + +A `Node` instance is constructed with the extracted source snippet, line ranges, and a human-readable `display_name` (e.g., `"class UserService"`), then appended to `self.nodes` and registered in the `top_level_nodes` dictionary for use during relationship extraction. + +### Pass 2: Relationship Extraction (`_extract_relationships`) + +This second recursive traversal identifies five categories of relationships, each producing one or more `CallRelationship` entries: + +```mermaid +flowchart LR + A["class_declaration
with superclass"] -->|"1: Inheritance"| R1["CallRelationship
class extends BaseClass"] + B["class/enum/record
with super_interfaces"] -->|"2: Interface Implementation"| R2["CallRelationship
class implements Interface"] + C["field_declaration"] -->|"3: Field Type Use"| R3["CallRelationship
class has field of Type"] + D["method_invocation"] -->|"4: Method Call"| R4["CallRelationship
caller calls object.method()"] + E["object_creation_expression"] -->|"5: Object Creation"| R5["CallRelationship
class creates new Type()"] +``` + +1. **Inheritance** — For `class_declaration` nodes with a `superclass` child, a relationship is created from the class to its base class (unless the base class is a Java primitive/built-in type). +2. **Interface Implementation** — For classes, enums, or records with a `super_interfaces` clause, a relationship is created to each implemented interface. +3. **Field Type Use** — For each `field_declaration`, if the containing class is resolvable and the field's type is a non-primitive type, a relationship is recorded from the class to the field's type. +4. **Method Calls** — For each `method_invocation`, the analyzer attempts to resolve the invoked object's declared type: first by checking if the object name matches a known top-level declaration, then by searching local variable declarations (`_search_variable_declaration`) and field declarations (`_find_variable_type`) within the enclosing method/class. If a type is resolved, a relationship is recorded from the calling method (or containing class, if outside a method) to the resolved type. +5. **Object Creation** — For each `object_creation_expression`, a relationship is recorded from the containing class to the instantiated type. + +All relationships are created with `is_resolved=False`, since the Java Analyzer only performs local, per-file resolution — full cross-file/cross-module resolution is handled later by the dependency analyzer core when building the complete dependency graph. + +### Type Filtering + +`_is_primitive_type()` filters out Java primitives (`int`, `boolean`, `char`, etc.), their boxed equivalents (`Integer`, `Boolean`, `Character`, etc.), and common JDK built-ins (`String`, `Object`, `List`, `Set`, `Map`, `Collection`, `Optional`, `void`, `Void`) so that relationships are only recorded for application-relevant types, keeping the dependency graph focused on meaningful code relationships rather than noise from standard library usage. + +## Data Model Reference + +The Java Analyzer produces instances of two shared Pydantic models defined outside this module: + +- **`Node`** — represents a structural component (class, interface, method, etc.) with fields such as `id`, `name`, `component_type`, `file_path`, `source_code`, `start_line`/`end_line`, and `display_name`. The `id` and `component_id` fields hold the same value: the analyzer's locally-computed `module.path::Name` identifier, which is later re-namespaced by the `DependencyParser` into a full FQDN. +- **`CallRelationship`** — represents a directed edge between a `caller` and `callee` component ID, with an optional `call_line` and an `is_resolved` flag. + +For full field definitions and usage across other analyzers, see the dependency analyzer models module. + +## Component ID Convention + +A key contract that `TreeSitterJavaAnalyzer` must honor is the component ID format expected by the rest of the system: `module.path::ClassName` (and `module.path::ClassName.methodName` for methods). This is implemented via two helper methods: + +- `_get_module_path()` — converts the file's path (relative to `repo_path`, if provided) into a dotted module path, stripping the `.java` extension and replacing path separators with dots. +- `_get_component_id(name, parent_class=None)` — combines the module path with the declaration name (optionally qualified by a parent class name) using the `::` separator. + +This convention ensures that IDs produced by the Java Analyzer are structurally consistent with those produced by the other analyzers under [Tree Sitter Analyzers](tree-sitter-analyzers.md) (C/C++/C#, JavaScript/TypeScript, PHP, Python), allowing the Dependency Parser to merge components from multiple languages and repositories into a single, namespaced dependency graph. + +## Integration Points + +| Consumer | Relationship | +|---|---| +| Analysis Pipeline (`CallGraphAnalyzer`) | Invokes `analyze_java_file()` for each `.java` file discovered during repository structure analysis. | +| Graph Construction (`DependencyParser`) | Consumes the raw `Node`/`CallRelationship` dictionaries (converted to `functions`/`relationships` lists) to build namespaced components and resolve dependencies, including cross-namespace resolution for multi-repository analysis. | +| Dependency Analyzer Models | Supplies the `Node` and `CallRelationship` schemas that this analyzer instantiates. | + +## Related Analyzers + +The Java Analyzer is one of several sibling language analyzers grouped under [Tree Sitter Analyzers](tree-sitter-analyzers.md): + +- [C Family Analyzers](c_family_analyzers.md) — C, C++, and C# analysis +- [JavaScript/TypeScript Analyzers](javascript_typescript_analyzers.md) +- [PHP Analyzer](php_analyzer.md) +- [Python Analyzer](python_analyzer.md) + +All of these analyzers follow the same general contract — accept a file path and content, return `Node` and `CallRelationship` lists — even though each implements language-specific AST traversal logic suited to its grammar. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md new file mode 100644 index 00000000..da8498a6 --- /dev/null +++ b/docs/reference/architecture/backend-core/tree-sitter-analyzers/javascript_typescript_analyzers.md @@ -0,0 +1,229 @@ +# JavaScript & TypeScript Analyzers + +The JavaScript & TypeScript Analyzers module implements the tree-sitter based source-code analyzers responsible for parsing JavaScript and TypeScript files and extracting the structural entities (functions, classes, interfaces, methods, type aliases, enums, etc.) and the relationships between them (calls, inheritance, type usage) that feed the dependency graph built by the backend analysis pipeline. + +It provides two closely related analyzer implementations: + +- `TreeSitterJSAnalyzer` — parses plain JavaScript (and JSX/mixed) source using the `tree_sitter_javascript` grammar. +- `TreeSitterTSAnalyzer` — parses TypeScript (and its embedded JavaScript constructs) using the `tree_sitter_typescript` grammar, adding support for TypeScript-only constructs such as interfaces, type aliases, enums, and type annotations. + +Both analyzers produce the same output contract — a list of `Node` objects and a list of `CallRelationship` objects — defined in the dependency analyzer models, so that downstream components can treat all language analyzers uniformly. + +## Role in the Analysis Pipeline + +This module is one of several language-specific analyzer implementations grouped under the [Tree Sitter Analyzers](tree-sitter-analyzers.md) parent module, alongside the [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), [PHP Analyzer](php_analyzer.md), and [Python Analyzer](python_analyzer.md). + +The analyzers in this module are invoked by the `DependencyParser` in the dependency analyzer core's graph construction stage, which dispatches source files to the appropriate analyzer based on file extension. The resulting `Node` and `CallRelationship` objects are then consumed by the `DependencyGraphBuilder` to assemble the full repository dependency graph, and by the analysis pipeline (`AnalysisService`, `RepoAnalyzer`, `CallGraphAnalyzer`) for higher-level repository analysis. + +```mermaid +flowchart LR + DP["DependencyParser"] -->|"dispatches .js/.jsx/.mjs/.cjs files"| JSA["TreeSitterJSAnalyzer"] + DP -->|"dispatches .ts/.tsx files"| TSA["TreeSitterTSAnalyzer"] + JSA -->|"produces"| Nodes["Node objects"] + JSA -->|"produces"| Rels["CallRelationship objects"] + TSA -->|"produces"| Nodes + TSA -->|"produces"| Rels + Nodes --> DGB["DependencyGraphBuilder"] + Rels --> DGB + DGB --> Graph["Repository Dependency Graph"] +``` + +## Component Overview + +| Component | File | Responsibility | +|---|---|---| +| `TreeSitterJSAnalyzer` | `codewiki/src/be/dependency_analyzer/analyzers/javascript.py` | Parses JavaScript source with tree-sitter, extracts functions/classes/methods and call/inheritance/JSDoc-type relationships | +| `analyze_javascript_file_treesitter` | same file | Module-level convenience function that instantiates `TreeSitterJSAnalyzer`, runs `analyze()`, and returns `(nodes, relationships)` | +| `TreeSitterTSAnalyzer` | `codewiki/src/be/dependency_analyzer/analyzers/typescript.py` | Parses TypeScript source with tree-sitter, extracts a broader set of entities (interfaces, type aliases, enums, exports) and relationships (calls, `new`, member access, type annotations, inheritance) | +| `analyze_typescript_file_treesitter` | same file | Module-level convenience function that instantiates `TreeSitterTSAnalyzer`, runs `analyze()`, and returns `(nodes, relationships)` | + +Both classes share the same overall shape: they accept `file_path`, `content`, and an optional `repo_path`, build a `tree_sitter.Parser` bound to the relevant `Language`, walk the resulting AST, and accumulate `Node` and `CallRelationship` instances (as defined in the dependency analyzer models). + +```mermaid +classDiagram + class TreeSitterJSAnalyzer { + +file_path: Path + +content: str + +repo_path: str + +nodes: List~Node~ + +call_relationships: List~CallRelationship~ + +top_level_nodes: dict + +analyze() None + -_extract_functions(node) None + -_extract_call_relationships(node) None + -_extract_jsdoc_type_dependencies(node, caller) None + } + class TreeSitterTSAnalyzer { + +file_path: Path + +content: str + +repo_path: str + +nodes: List~Node~ + +call_relationships: List~CallRelationship~ + +top_level_nodes: dict + +analyze() None + -_extract_all_entities(node, all_entities, depth) None + -_filter_top_level_declarations(all_entities) None + -_extract_all_relationships(node, all_entities) None + } + class Node { + +id: str + +name: str + +component_type: str + +source_code: str + +start_line: int + +end_line: int + +node_type: str + +base_classes: List~str~ + } + class CallRelationship { + +caller: str + +callee: str + +call_line: int + +is_resolved: bool + } + TreeSitterJSAnalyzer --> Node : produces + TreeSitterJSAnalyzer --> CallRelationship : produces + TreeSitterTSAnalyzer --> Node : produces + TreeSitterTSAnalyzer --> CallRelationship : produces +``` + +## TreeSitterJSAnalyzer + +`TreeSitterJSAnalyzer` initializes a tree-sitter `Parser` bound to the `tree_sitter_javascript` grammar. If grammar/language initialization fails, the parser is left `None` and `analyze()` becomes a no-op (logged as a warning), which keeps the pipeline resilient to environment issues. + +### Extraction flow + +1. **`analyze()`** parses the source into an AST and calls `_extract_functions()` followed by `_extract_call_relationships()`. +2. **`_extract_functions()` / `_traverse_for_functions()`** walk the AST recursively, recognizing: + - `class_declaration`, `abstract_class_declaration`, `interface_declaration` → extracted via `_extract_class_declaration()`, with their `class_body` scanned for `method_definition` and arrow-function `field_definition` members via `_extract_methods_from_class()`. + - `function_declaration` / `generator_function_declaration` (only when not nested in a class) → `_extract_function_declaration()`. + - `export_statement` → `_extract_exported_function()` (handles `export function` and `export default function`). + - `lexical_declaration` (`const`/`let`) → `_extract_arrow_function_from_declaration()` for arrow functions and function expressions assigned to a variable. +3. **`_extract_call_relationships()` / `_traverse_for_calls()`** re-walks the tree tracking the "current top-level" entity (the enclosing function or class), and records: + - `call_expression` and `await_expression`-wrapped calls → resolved via `_extract_call_from_node()`, checking `this.`/`super.` member calls against known methods to avoid duplicate self-references. + - `new_expression` → constructor instantiation relationships. + - Class `class_heritage` (`extends`) → inheritance `CallRelationship` entries. + - JSDoc comments (`@param {Type}`, `@returns {Type}`, `@type {Type}`, `@typedef`, `@interface`) → parsed by `_parse_jsdoc_types()` / `_extract_base_types_from_jsdoc()` into type-dependency relationships, filtered against a built-in type list via `_is_builtin_type_js()`. + +All relationships are de-duplicated through `_add_relationship()`, which tracks a `(caller, callee, call_line)` key in `seen_relationships`. + +```mermaid +flowchart TD + Start["analyze()"] --> Parse["Parser.parse(content)"] + Parse --> ExtractFn["_extract_functions(root_node)"] + ExtractFn --> TraverseFn["_traverse_for_functions(node)"] + TraverseFn -->|"class/abstract/interface"| ExtractClass["_extract_class_declaration()"] + ExtractClass --> ExtractMethods["_extract_methods_from_class()"] + TraverseFn -->|"function_declaration"| ExtractFunc["_extract_function_declaration()"] + TraverseFn -->|"export_statement"| ExtractExport["_extract_exported_function()"] + TraverseFn -->|"lexical_declaration"| ExtractArrow["_extract_arrow_function_from_declaration()"] + ExtractFn --> ExtractCalls["_extract_call_relationships(root_node)"] + ExtractCalls --> TraverseCalls["_traverse_for_calls(node, current_top_level)"] + TraverseCalls -->|"call_expression"| CallRel["_extract_call_from_node()"] + TraverseCalls -->|"new_expression"| NewRel["Constructor relationship"] + TraverseCalls -->|"class_heritage"| InheritRel["Inheritance relationship"] + TraverseCalls -->|"JSDoc comment"| JsdocRel["_parse_jsdoc_types()"] + CallRel --> AddRel["_add_relationship() (dedup)"] + NewRel --> AddRel + InheritRel --> AddRel + JsdocRel --> AddRel +``` + +### Component identity + +Component IDs are built by `_get_component_id()` as `::` for top-level entities and `::.` for methods, where `_get_module_path()` derives a dotted module path from the file's path relative to `repo_path`, stripping `.js`/`.ts`/`.jsx`/`.tsx`/`.mjs`/`.cjs` extensions. Note that internally, call/inheritance relationships built during traversal use a distinct dotted convention (`.`) rather than the `::` component-id separator; graph construction downstream reconciles these against the canonical `Node.id` values. + +## TreeSitterTSAnalyzer + +`TreeSitterTSAnalyzer` follows a two-pass design that is more elaborate than the JS analyzer, reflecting TypeScript's richer set of top-level declaration forms. + +### Pass 1 — Entity collection (`_extract_all_entities`) + +A single recursive traversal builds a flat `all_entities` dictionary keyed by entity name, capturing one of the following node types at any depth: +`function_declaration`, `generator_function_declaration`, `arrow_function`, `method_definition`, `class_declaration`, `abstract_class_declaration`, `interface_declaration`, `type_alias_declaration`, `enum_declaration`, `variable_declarator`, `export_statement`, `lexical_declaration`, `variable_declaration`, `ambient_declaration`. + +Each captured entity dict stores `name`, `type`/`subtype`, `code_snippet`, `display_name`, line range, and (for functions) `parameters`, plus bookkeeping fields `depth`, `node` (the raw tree-sitter node), and `parent_context` (via `_get_parent_context()`). + +### Pass 2 — Top-level filtering (`_filter_top_level_declarations`) + +For every collected entity, `_is_actually_top_level()` walks up the parent chain to decide whether the entity is genuinely a module-level declaration (directly under `program`, `export_statement`, `ambient_declaration`, `module`, or a `statement_block` that is itself inside a `module`/`ambient_declaration`) versus nested inside a function body (checked via `_is_inside_function_body()`). Only entities classified as top-level are converted into `Node` objects (via `_create_node_from_entity()`) and filtered further by `_should_include_node()`, which excludes bare `variable` entities and reserved names (`constructor`, `__proto__`, `prototype`). + +For class/abstract-class entities, `_extract_constructor_dependencies()` additionally inspects the constructor's `formal_parameters` for TypeScript type annotations and records a dependency relationship per typed parameter via `_extract_parameter_dependencies()`. + +### Pass 3 — Relationship extraction (`_extract_all_relationships`) + +A third traversal (`_traverse_for_relationships`) tracks the current enclosing top-level entity (updated whenever a "new top-level" node type is encountered per `_is_new_top_level()`/`_get_top_level_name()`) and, for each descendant node, extracts: + +| Node type | Handler | Relationship captured | +|---|---|---| +| `call_expression` | `_extract_call_relationship()` | Function/method calls, filtering `this.`/`super.` calls to nested methods of the same class | +| `new_expression` | `_extract_new_relationship()` | Constructor instantiation | +| `member_expression` | `_extract_member_relationship()` | Property/member access | +| `type_annotation` | `_extract_type_relationship()` | Parameter/return type usage, filtered by `_is_builtin_type()` | +| `type_arguments` | `_extract_type_arguments_relationship()` | Generic type parameters | +| `extends_clause` / `implements_clause` | `_extract_inheritance_relationship()` | Class/interface inheritance and interface implementation | + +All relationships are appended via `_add_relationship()`, which builds `caller`/`callee` ids as `.` and marks them `is_resolved=False` (resolution against the full repository graph happens downstream in the dependency analyzer core's graph construction stage). + +```mermaid +flowchart TD + A["analyze()"] --> B["Pass 1: _extract_all_entities()
builds all_entities dict"] + B --> C["Pass 2: _filter_top_level_declarations()
_is_actually_top_level() check"] + C --> D["Node objects appended to self.nodes
and self.top_level_nodes"] + D --> E["_extract_constructor_dependencies()
(classes only)"] + B --> F["Pass 3: _extract_all_relationships()"] + F --> G["_traverse_for_relationships()
tracks current_top_level"] + G -->|"call_expression"| H["_extract_call_relationship()"] + G -->|"new_expression"| I["_extract_new_relationship()"] + G -->|"member_expression"| J["_extract_member_relationship()"] + G -->|"type_annotation / type_arguments"| K["_extract_type_relationship()"] + G -->|"extends_clause / implements_clause"| L["_extract_inheritance_relationship()"] + H --> M["_add_relationship()"] + I --> M + J --> M + K --> M + L --> M +``` + +## JS vs TS Analyzer: Key Differences + +| Aspect | `TreeSitterJSAnalyzer` | `TreeSitterTSAnalyzer` | +|---|---|---| +| Grammar | `tree_sitter_javascript` | `tree_sitter_typescript` (`language_typescript()`) | +| Traversal design | Two combined traversals (functions, then calls) inline while walking | Three explicit passes: entity collection, top-level filtering, relationship extraction | +| TypeScript-only entities | Not supported | `interface_declaration`, `type_alias_declaration`, `enum_declaration`, `ambient_declaration` | +| Type-dependency source | JSDoc comments (`@param`, `@returns`, `@type`, `@typedef`, `@interface`) parsed with regex | Native TS syntax: `type_annotation`, `type_arguments`, `extends_clause`/`implements_clause` | +| Constructor parameter typing | Not applicable (no static types) | `_extract_constructor_dependencies()` reads `type_annotation` on constructor parameters | +| Variable/export handling | Arrow functions from `const`/`let` handled inline in traversal | Dedicated entity extractors: `_extract_lexical_declaration_entity()`, `_extract_variable_declaration_entity()`, `_extract_export_statement_entity()` | +| Built-in type filtering | `_is_builtin_type_js()` (broad DOM + JS globals list) | `_is_builtin_type()` (primitive TS types only) + `_is_builtin_function()` (currently empty) | + +## Data Flow: From Source File to Dependency Graph + +```mermaid +sequenceDiagram + participant DP as "DependencyParser" + participant JS as "TreeSitterJSAnalyzer" + participant TS as "TreeSitterTSAnalyzer" + participant DGB as "DependencyGraphBuilder" + + DP->>JS: analyze_javascript_file_treesitter(path, content, repo_path) + JS->>JS: Parser.parse(content) + JS->>JS: _extract_functions() / _extract_call_relationships() + JS-->>DP: (nodes, call_relationships) + + DP->>TS: analyze_typescript_file_treesitter(path, content, repo_path) + TS->>TS: Parser.parse(content) + TS->>TS: _extract_all_entities() / _filter_top_level_declarations() / _extract_all_relationships() + TS-->>DP: (nodes, call_relationships) + + DP->>DGB: aggregated nodes + relationships + DGB->>DGB: resolve callee ids against repository Node index + DGB-->>DP: Repository dependency graph +``` + +## Integration Points + +- **Input contract**: Both analyzers are constructed with `(file_path, content, repo_path)` and expose `analyze()` plus `nodes: List[Node]` and `call_relationships: List[CallRelationship]`, matching the shared interface used across all analyzers in [Tree Sitter Analyzers](tree-sitter-analyzers.md) (see also [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), [PHP Analyzer](php_analyzer.md), [Python Analyzer](python_analyzer.md)). +- **Output models**: `Node` and `CallRelationship` are defined in the dependency analyzer models. +- **Consumers**: The `DependencyParser` and `DependencyGraphBuilder` in the dependency analyzer core's graph construction stage invoke these analyzers per-file and merge their output into the overall repository graph; the analysis pipeline then operates on the resulting graph for call-graph and repository-level analysis. +- **Module-level entry points**: `analyze_javascript_file_treesitter()` and `analyze_typescript_file_treesitter()` are the primary functions called by upstream dispatch logic; both wrap analyzer instantiation and `analyze()` in a try/except that returns empty lists on failure, ensuring a single malformed file does not abort the overall analysis run. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md new file mode 100644 index 00000000..0019a4d7 --- /dev/null +++ b/docs/reference/architecture/backend-core/tree-sitter-analyzers/php_analyzer.md @@ -0,0 +1,225 @@ +# PHP Analyzer + +The PHP Analyzer module is a language-specific static analysis component of CodeWiki's dependency analysis pipeline. It parses PHP source files using `tree-sitter-php` to extract structural code entities (classes, interfaces, traits, enums, functions, and methods) and the relationships between them (inheritance, interface implementation, object instantiation, static calls, constructor property promotion, and namespace imports). Its output feeds directly into the shared dependency analyzer models (`Node`, `CallRelationship`) that the rest of the CodeWiki dependency graph pipeline consumes. + +This module is one of several per-language analyzers registered under the [Tree Sitter Analyzers](tree-sitter-analyzers.md) module, alongside analyzers for [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), [JavaScript/TypeScript Analyzers](javascript_typescript_analyzers.md), and [Python Analyzer](python_analyzer.md). + +## Purpose and Scope + +PHP has language features that make dependency extraction non-trivial compared to simpler languages: + +- **Namespaces and `use` statements**, including aliasing (`use App\User as U;`) and grouped imports (`use App\{User, Post};`) +- **Fully-qualified vs. relative class names** that must be resolved relative to the current namespace and import table +- **Multiple declaration kinds** in a single file: classes (including abstract classes), interfaces, traits, enums, free functions, and methods +- **Template files** (e.g. Blade `.blade.php`, `.phtml`) that mix PHP with HTML/markup and should typically be excluded from dependency analysis +- **PHP 8+ constructor property promotion**, which implicitly creates typed dependencies on constructor parameters + +The PHP Analyzer module addresses all of these concerns through two cooperating classes: + +- **`NamespaceResolver`** — resolves short/aliased PHP type names to fully-qualified names (FQNs) +- **`TreeSitterPHPAnalyzer`** — walks the tree-sitter AST for a PHP file, using the resolver to produce `Node` and `CallRelationship` objects + +A convenience function, `analyze_php_file`, wraps analyzer construction and returns the extracted `(nodes, call_relationships)` tuple, matching the calling convention expected by the broader dependency-analysis pipeline. + +## Architecture + +```mermaid +flowchart TD + Caller["Dependency Analysis Pipeline"] -->|"analyze_php_file(path, content, repo_path)"| Analyze["analyze_php_file()"] + Analyze --> Analyzer["TreeSitterPHPAnalyzer"] + Analyzer -->|"uses"| Resolver["NamespaceResolver"] + Analyzer -->|"parses with"| TreeSitter["tree_sitter_php / Parser"] + Analyzer -->|"produces"| Nodes["List of Node"] + Analyzer -->|"produces"| Rels["List of CallRelationship"] + Nodes --> Models["Dependency Analyzer Models"] + Rels --> Models +``` + +`Node` and `CallRelationship` are defined outside this module in the shared dependency analyzer models component group, so the PHP Analyzer module only produces instances of them — it does not own their schema. + +### Class Relationships + +```mermaid +classDiagram + class NamespaceResolver { + +string current_namespace + +dict use_map + +register_namespace(ns) + +register_use(fqn, alias) + +resolve(name) string + } + class TreeSitterPHPAnalyzer { + +Path file_path + +string content + +string repo_path + +List nodes + +List call_relationships + +NamespaceResolver namespace_resolver + +_analyze() + +_extract_namespace_info(node, depth) + +_extract_nodes(node, lines, depth, parent_class) + +_extract_relationships(node, depth) + +_is_template_file() bool + +_get_component_id(name, parent_class) string + } + class Node { + +string id + +string name + +string component_type + +string file_path + +Set depends_on + +string source_code + +int start_line + +int end_line + } + class CallRelationship { + +string caller + +string callee + +int call_line + +bool is_resolved + } + TreeSitterPHPAnalyzer --> NamespaceResolver : "delegates name resolution" + TreeSitterPHPAnalyzer --> Node : "creates" + TreeSitterPHPAnalyzer --> CallRelationship : "creates" +``` + +## Core Components + +### NamespaceResolver + +`NamespaceResolver` maintains the state needed to translate PHP type names as they appear in source code into fully-qualified names (FQNs) suitable for use as dependency identifiers. + +Internal state: +- `current_namespace: str` — the namespace declared for the file (set via `register_namespace`) +- `use_map: Dict[str, str]` — maps a local alias (or the short class name if no alias) to its FQN, populated via `register_use` + +Resolution logic in `resolve(name)`: +1. If `name` already starts with `\`, it is already fully qualified — the leading backslash is stripped and it is returned as-is. +2. If `name` matches an entry in `use_map` exactly, the mapped FQN is returned. +3. If the first segment of a dotted/qualified name matches an alias in `use_map`, the base is substituted and the remainder appended. +4. Otherwise, the current namespace is prepended to `name`. +5. If there is no current namespace and no match, the name is returned unchanged. + +This mirrors PHP's own name-resolution rules for `use App\Foo as Bar; new Bar();` style code, ensuring dependency edges point at real, de-aliased classes. + +### TreeSitterPHPAnalyzer + +`TreeSitterPHPAnalyzer` is the primary entry point for analyzing a single PHP file. It is constructed with the file path, its content, and (optionally) the repository root path used to compute relative/module paths. + +**Initialization flow:** + +```mermaid +sequenceDiagram + participant Caller as "Caller" + participant Analyzer as "TreeSitterPHPAnalyzer" + participant Parser as "tree_sitter Parser" + participant Resolver as "NamespaceResolver" + + Caller->>Analyzer: "__init__(file_path, content, repo_path)" + Analyzer->>Analyzer: "_is_template_file()" + alt is template file + Analyzer-->>Caller: "return (skip analysis)" + else not a template + Analyzer->>Parser: "parse(content)" + Parser-->>Analyzer: "AST root node" + Analyzer->>Analyzer: "_extract_namespace_info(root)" + Analyzer->>Resolver: "register_namespace() / register_use()" + Analyzer->>Analyzer: "_extract_nodes(root, lines)" + Analyzer->>Analyzer: "_extract_relationships(root)" + Analyzer->>Resolver: "resolve(type_name)" + Analyzer-->>Caller: "nodes, call_relationships populated" + end +``` + +#### Template File Skipping + +Before any parsing occurs, `_is_template_file()` checks the file path against: +- Extension patterns: `.blade.php`, `.phtml`, `.twig.php` +- Directory patterns: `views`, `templates`, `resources/views` + +If matched, analysis is skipped entirely (`nodes` and `call_relationships` remain empty), avoiding noisy or irrelevant dependency data from view/template files. + +#### Three-Pass Analysis + +Once a file is confirmed to be analyzable PHP, `_analyze()` parses it with the `tree_sitter_php` grammar and runs three sequential passes over the AST: + +1. **Namespace/use extraction** (`_extract_namespace_info`) — walks the whole tree looking for `namespace_definition` and `namespace_use_declaration` nodes, registering the namespace and populating the `NamespaceResolver`'s `use_map` (including grouped `use App\{User, Post};` syntax via `_extract_use_statement`). +2. **Node extraction** (`_extract_nodes`) — walks the tree again, this time creating a `Node` for each recognized declaration type: + - `class_declaration` → `"class"` or `"abstract class"` (detected via `abstract_modifier`) + - `interface_declaration` → `"interface"` + - `trait_declaration` → `"trait"` + - `enum_declaration` → `"enum"` + - `function_definition` → `"function"` + - `method_declaration` → `"method"` (name is composed as `ContainingClass.methodName`) + + For classes/methods, extra metadata is captured: `parameters` (via `_extract_parameters`) and `base_classes` (via `_extract_base_classes`). Every node also gets a preceding PHPDoc comment attached as its `docstring`, discovered by `_get_preceding_docstring`. +3. **Relationship extraction** (`_extract_relationships`) — a further tree walk that emits `CallRelationship` objects for: + - `namespace_use_declaration` → import-based dependency from the file module to the imported FQN + - `class_declaration` with a `base_clause` → inheritance (`extends`) + - `class_declaration`/`enum_declaration` with a `class_interface_clause` → interface implementation + - `object_creation_expression` (`new Foo()`) → instantiation dependency + - `scoped_call_expression` (`Foo::bar()`) → static-call dependency + - `property_promotion_parameter` (PHP 8 constructor promotion) → typed dependency on the promoted parameter's type + +All PHP primitive/built-in types (see `PHP_PRIMITIVES`, e.g. `string`, `int`, `self`, `static`, `Exception`, `DateTime`) are filtered out via `_is_primitive` so that dependency edges reflect meaningful application code rather than language or standard-library noise. + +A `MAX_RECURSION_DEPTH` guard (100) is applied to every recursive tree walk to protect against pathologically deep ASTs causing a Python `RecursionError`. + +#### Component ID Generation + +`_get_component_id(name, parent_class)` produces the identifier used as a `Node.id` / `CallRelationship.caller`/`callee` value. If a namespace was detected for the file, the ID is built as: + +```text +Namespace.With.Dots::ParentClass.memberName +``` + +Otherwise it falls back to a module path derived from the file's location relative to the repository root (via `_get_module_path`), using `::` to separate the module path from the local declaration name — consistent with the ID convention expected by the graph construction stage's `DependencyParser`, which splits on `::` when deriving module groupings. + +### analyze_php_file + +```python +def analyze_php_file(file_path: str, content: str, repo_path: str = None) -> Tuple[List[Node], List[CallRelationship]] +``` + +This module-level function is the simple functional interface used by callers that do not need direct access to the analyzer instance: it constructs a `TreeSitterPHPAnalyzer`, runs analysis as part of construction, and returns the resulting `nodes` and `call_relationships` lists. This mirrors the calling pattern used by sibling analyzers such as those in [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), and [Python Analyzer](python_analyzer.md), allowing the orchestrating graph construction parser to treat all language analyzers uniformly. + +## Data Flow: PHP File to Dependency Graph + +```mermaid +flowchart LR + File["PHP source file"] --> Analyzer["TreeSitterPHPAnalyzer"] + Analyzer -->|"nodes"| NodeList["List of Node objects
(classes, methods, functions, etc.)"] + Analyzer -->|"relationships"| RelList["List of CallRelationship objects
(extends, implements, new, static call)"] + NodeList --> Parser["DependencyParser
(Graph Construction)"] + RelList --> Parser + Parser --> Graph["Dependency Graph
(components + depends_on edges)"] +``` + +The `Node` objects produced here carry `depends_on` as an empty set by default; it is the surrounding graph construction logic (`DependencyParser`) that consumes the emitted `CallRelationship` list to populate each `Node.depends_on` set and assign final namespaced component IDs (FQDNs) across the whole repository — potentially spanning multiple languages. + +## Relationship Types Extracted + +| PHP Construct | AST Node Type | Relationship Produced | +|---|---|---| +| `use App\Foo;` | `namespace_use_declaration` | File-to-class import dependency | +| `class A extends B` | `class_declaration` + `base_clause` | `A` → `B` (inheritance) | +| `class A implements I` | `class_declaration`/`enum_declaration` + `class_interface_clause` | `A` → `I` (implementation) | +| `new Foo()` | `object_creation_expression` | Containing class/function → `Foo` (instantiation) | +| `Foo::bar()` | `scoped_call_expression` | Containing class/function → `Foo` (static call) | +| `public function __construct(private Foo $foo)` | `property_promotion_parameter` | Containing class → `Foo` (typed dependency) | + +Each relationship is emitted as `is_resolved=False`, meaning name resolution to a final graph-wide component ID is deferred to the shared graph-construction stage rather than being finalized within this module. + +## Integration Points + +- **Input model**: raw PHP file paths and content, plus an optional repository root, supplied by the orchestrating dependency parser as part of graph construction. +- **Output model**: `Node` and `CallRelationship` instances defined in the dependency analyzer models. +- **Sibling analyzers**: registered together under [Tree Sitter Analyzers](tree-sitter-analyzers.md) for other supported languages (C/C++/C#, Java, JavaScript/TypeScript, Python). +- **Downstream consumers**: the assembled dependency graph is used by the analysis pipeline (`AnalysisService`, `RepoAnalyzer`, `CallGraphAnalyzer`) as part of generating repository-wide documentation. + +## Design Notes + +- **Template exclusion is path/extension based**, not content based — a `.blade.php` file is skipped regardless of its actual PHP content, trading recall for precision and avoiding markup-heavy files that rarely represent meaningful code dependencies. +- **Namespace resolution happens in a dedicated pre-pass** before node/relationship extraction so that `use` aliases declared anywhere in the file are available when resolving type references later in the same file, regardless of declaration order. +- **Recursion depth protection** (`MAX_RECURSION_DEPTH = 100`) is applied uniformly across all three AST walks, with graceful degradation (a warning log and early return) rather than raising unhandled exceptions. +- **Primitive/built-in type filtering** is centralized in `PHP_PRIMITIVES`, keeping dependency graphs focused on user-defined types. diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md new file mode 100644 index 00000000..5b917ddd --- /dev/null +++ b/docs/reference/architecture/backend-core/tree-sitter-analyzers/python_analyzer.md @@ -0,0 +1,196 @@ +# Python Analyzer + +The Python Analyzer module implements an AST-based static analysis engine for Python source files. It walks the Python `ast` (Abstract Syntax Tree) produced by the standard library to extract structural information — classes, functions, and their call relationships — and converts that information into the standardized `Node` and `CallRelationship` models used throughout the dependency analysis pipeline. + +This module is one of several language-specific analyzers under the [Tree Sitter Analyzers](tree-sitter-analyzers.md) family. While its siblings (C, C++, C#, Java, JavaScript/TypeScript, PHP) rely on the `tree-sitter` parsing library, the Python analyzer uniquely leverages Python's own built-in `ast` module, since CPython ships with a fully-featured native parser for its own syntax. + +## Purpose and Role in the System + +The Python Analyzer is invoked by the dependency analyzer core during repository analysis to process individual `.py` files. Its output — lists of `Node` and `CallRelationship` objects — feeds directly into the dependency analyzer models and ultimately into the graph construction stage handled by `DependencyGraphBuilder`. + +Within the broader system, the analysis pipeline works roughly as follows: + +1. The `RepoAnalyzer` (part of the dependency analyzer core) walks the repository file tree and identifies Python source files. +2. For each Python file, `PythonASTAnalyzer` is instantiated and its `analyze()` method is invoked. +3. The analyzer parses the file content into an AST, traverses it, and produces `Node` objects (representing classes and top-level functions) plus `CallRelationship` objects (representing call-graph edges between them). +4. These results are aggregated with the output of other language analyzers and passed to `DependencyGraphBuilder` to construct the overall dependency graph. + +```mermaid +flowchart TD + RepoAnalyzer["RepoAnalyzer"] -->|"reads .py file"| PyFile["Python Source File"] + RepoAnalyzer -->|"instantiate"| PythonASTAnalyzer["PythonASTAnalyzer"] + PyFile -->|"content"| PythonASTAnalyzer + PythonASTAnalyzer -->|"ast.parse()"| AST["Python AST"] + AST -->|"NodeVisitor traversal"| PythonASTAnalyzer + PythonASTAnalyzer -->|"produces"| Nodes["List of Node"] + PythonASTAnalyzer -->|"produces"| Relationships["List of CallRelationship"] + Nodes --> DependencyGraphBuilder["DependencyGraphBuilder"] + Relationships --> DependencyGraphBuilder + DependencyGraphBuilder --> Graph["Dependency Graph"] +``` + +## Core Component + +### `PythonASTAnalyzer` + +`PythonASTAnalyzer` extends `ast.NodeVisitor` from the Python standard library, using the visitor pattern to walk the syntax tree and collect structural information as it goes. + +**Constructor parameters:** + +| Parameter | Type | Description | +|---|---|---| +| `file_path` | `str` | Absolute or relative path to the Python file being analyzed | +| `content` | `str` | Raw source code content of the file | +| `repo_path` | `Optional[str]` | Repository root path, used to compute paths relative to the repo for FQDN generation | + +**Key internal state:** + +- `nodes: List[Node]` — accumulates discovered classes and top-level functions +- `call_relationships: List[CallRelationship]` — accumulates discovered call edges +- `current_class_name` / `current_function_name` — tracks the visitor's current traversal context so calls can be attributed to the correct enclosing scope +- `top_level_nodes: dict` — maps simple names to their `Node` objects, used to resolve whether a given call target is a locally-defined class/function or an external reference + +## Component ID Generation (FQDN Format) + +A central responsibility of the analyzer is producing globally consistent, fully-qualified component IDs so that call relationships can be resolved across the whole dependency graph — not just within a single file. IDs follow the format: + +```text +:: +``` + +- The dotted module path is derived from the file's path relative to the repository root (via `_get_relative_path()` and `_get_module_path()`), stripping the `.py`/`.pyx` extension and converting path separators to dots. +- The component name is either a bare class/function name, or `ClassName.method_name` when inside a class body (tracked via `current_class_name`). + +This is implemented in `_get_module_path()` and `_get_component_id()`: + +```mermaid +flowchart LR + FilePath["file_path"] -->|"os.path.relpath"| RelPath["relative_path"] + RelPath -->|"strip .py/.pyx, replace separators with dots"| ModulePath["module.path"] + ModulePath -->|"+ '::' + ComponentName"| ComponentID["component_id"] +``` + +## AST Traversal and Extraction Logic + +### Class Extraction — `visit_ClassDef` + +When the visitor encounters a `ClassDef` node: + +1. Base classes are extracted via `_extract_base_class_name`, which handles both simple names (`class Foo(Bar)`) and dotted attribute access (`class Foo(module.Bar)`). +2. A `Node` is constructed with `component_type="class"`, capturing source code slice, docstring, line ranges, and base class names. +3. The class is registered into `top_level_nodes` for later call resolution. +4. If any base class is itself a known top-level node in the same file, an `is_resolved=True` `CallRelationship` is recorded representing the inheritance edge. +5. `current_class_name` is set before recursing into the class body (via `generic_visit`) and cleared afterward, so nested calls/methods are correctly attributed. + +### Function Extraction — `visit_FunctionDef` / `visit_AsyncFunctionDef` + +Both regular and `async def` functions are routed through `_process_function_node`: + +1. Only **top-level** functions (i.e., `current_class_name` is `None`) are added as standalone `Node` objects with `component_type="function"`. Methods defined inside classes are intentionally excluded from being top-level nodes (they are traversed for call extraction, but not separately registered as class members in this analyzer version). +2. `_should_include_function` filters out functions whose names start with `_test_` (helper filter to avoid polluting the graph with certain test scaffolding functions). +3. `current_function_name` is tracked during the recursive visit so that calls inside the function body can be attributed to it. + +### Call Relationship Extraction — `visit_Call` + +For every function/method call expression encountered during traversal: + +1. `_get_call_name` resolves the callee's name from the AST call target, handling: + - Simple name calls: `foo()` + - Attribute calls: `obj.method()` → resolved as `obj.method` + - Chained attributes: `a.b.method()` + - A hardcoded set of Python builtins (`print`, `len`, `isinstance`, etc.) is filtered out to avoid noise from standard library calls. +2. The caller ID is determined by the current traversal context — either the enclosing class or the enclosing top-level function. +3. If the resolved call name matches an entry in `top_level_nodes`, the relationship is marked `is_resolved=True` and the callee ID is fully qualified with the module path. Otherwise, the raw call name is stored as an unresolved reference (to potentially be resolved later at the cross-file graph-building stage). + +```mermaid +sequenceDiagram + participant Visitor as PythonASTAnalyzer + participant AST as ast.NodeVisitor + participant Nodes as top_level_nodes + participant Rels as call_relationships + + AST->>Visitor: visit_ClassDef(node) + Visitor->>Visitor: extract base_classes + Visitor->>Nodes: register class Node + Visitor->>Rels: append inheritance CallRelationship (if resolved) + Visitor->>Visitor: set current_class_name + Visitor->>AST: generic_visit(node) + AST->>Visitor: visit_Call(node) [inside class/function body] + Visitor->>Visitor: _get_call_name(node.func) + Visitor->>Nodes: lookup call_name + Visitor->>Rels: append CallRelationship (resolved or unresolved) + Visitor->>Visitor: clear current_class_name +``` + +## Data Model Alignment + +The `Node` and `CallRelationship` objects produced by this analyzer are defined in the dependency analyzer models. The Python Analyzer populates the following key fields: + +**`Node` fields populated:** + +| Field | Description | +|---|---| +| `id` / `component_id` | FQDN-format identifier (`module.path::Name`) | +| `name` | Simple class or function name | +| `component_type` | `"class"` or `"function"` | +| `file_path` / `relative_path` | Absolute and repo-relative file locations | +| `source_code` | Exact source lines spanning the node's definition | +| `start_line` / `end_line` | Line range from the AST node | +| `has_docstring` / `docstring` | Extracted via `ast.get_docstring()` | +| `parameters` | Argument names, for functions only | +| `node_type` | Mirrors `component_type` (`"class"` / `"function"`) | +| `base_classes` | List of resolved base class names, for classes only | +| `display_name` | Human-readable label, e.g. `"class Foo"` or `"function bar"` | + +**`CallRelationship` fields populated:** + +| Field | Description | +|---|---| +| `caller` | FQDN of the enclosing class or function | +| `callee` | FQDN (if resolved) or raw call name (if unresolved) | +| `call_line` | Line number of the call expression | +| `is_resolved` | `True` if the callee matches a locally-known top-level node | + +## Error Handling and Robustness + +The `analyze()` method wraps parsing and traversal in exception handling to ensure a single malformed file does not halt the entire repository analysis: + +- **`SyntaxError`**: Caught and logged as a warning when `ast.parse()` fails on invalid Python syntax; analysis for that file is skipped gracefully. +- **General exceptions**: Caught and logged as errors with full traceback (`exc_info=True`), keeping the pipeline resilient to unexpected AST edge cases. +- **`SyntaxWarning` suppression**: Parsing is wrapped in `warnings.catch_warnings()` to silence `SyntaxWarning`s (e.g., from invalid escape sequences in string/regex literals within analyzed source files), preventing noisy log output during large-scale repository scans. + +```mermaid +flowchart TD + Start["analyze() called"] --> Parse["ast.parse(content)"] + Parse -->|"success"| Visit["self.visit(tree)"] + Parse -->|"SyntaxError"| WarnLog["log warning, skip file"] + Visit -->|"success"| Done["nodes + call_relationships populated"] + Visit -->|"unexpected Exception"| ErrLog["log error with traceback"] +``` + +## Public Entry Point + +The module exposes a convenience function for one-shot analysis without manually managing the analyzer instance: + +```python +def analyze_python_file( + file_path: str, content: str, repo_path: Optional[str] = None +) -> Tuple[List[Node], List[CallRelationship]]: + ... +``` + +This function instantiates `PythonASTAnalyzer`, calls `.analyze()`, and returns the `(nodes, call_relationships)` tuple directly — the typical integration point used by callers such as `RepoAnalyzer` in the dependency analyzer core. + +## Design Notes and Limitations + +- **Top-level focus**: Only classes and top-level (module-scope) functions are registered as `Node` objects. Methods defined within classes are traversed for call-extraction purposes but are not independently registered as separate top-level nodes in this analyzer's current implementation — call attribution for method bodies is scoped to the enclosing class. +- **Builtin filtering**: A static allowlist of common Python builtins is excluded from call relationship extraction to reduce graph noise; this list is not exhaustive and may need periodic updates as usage patterns evolve. +- **Local-file resolution only**: Call resolution (`is_resolved`) is determined using only the nodes discovered within the same file (via `top_level_nodes`). Cross-file/cross-module call resolution is deferred to later stages of the pipeline, such as `DependencyGraphBuilder` in the dependency analyzer core. +- **Test function filtering**: Functions with names starting with `_test_` are excluded via `_should_include_function`, a convention-based filter to reduce noise from certain test helper patterns. + +## Related Modules + +- [Tree Sitter Analyzers](tree-sitter-analyzers.md) — parent grouping of all language-specific analyzers, including the [C Family Analyzers](c_family_analyzers.md), [Java Analyzer](java_analyzer.md), [JavaScript/TypeScript Analyzers](javascript_typescript_analyzers.md), and [PHP Analyzer](php_analyzer.md), which follow an analogous extraction pattern using `tree-sitter` grammars instead of Python's native `ast` module. +- Dependency Analyzer Core — hosts `RepoAnalyzer`, `AnalysisService`, `CallGraphAnalyzer`, and `DependencyGraphBuilder`, which orchestrate invocation of this analyzer and consume its output. +- Dependency Analyzer Models — defines the `Node` and `CallRelationship` Pydantic models that structure this analyzer's output. + diff --git a/docs/reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md b/docs/reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md new file mode 100644 index 00000000..632f68f5 --- /dev/null +++ b/docs/reference/architecture/backend-core/tree-sitter-analyzers/tree-sitter-analyzers.md @@ -0,0 +1,115 @@ +# Tree Sitter Analyzers + +## Purpose + +The Tree Sitter Analyzers module is the multi-language source-code parsing layer of CodeWiki's dependency analysis pipeline. For every supported programming language, it provides a dedicated analyzer that reads a single source file and produces two things: + +1. **Nodes** — structural entities such as classes, functions, methods, interfaces, structs, enums, and global variables. +2. **Call Relationships** — edges describing how those entities depend on or interact with one another (function calls, inheritance, interface implementation, object creation, type usage, and more). + +These `Node` and `CallRelationship` objects are the common currency of the entire dependency analysis system. They are defined by the shared data models used across the backend and are consumed by the graph-construction stage that assembles the full repository dependency graph and the call-graph analyzer used for clustering and documentation generation. + +Because every analyzer emits data in the exact same shape, the rest of the pipeline (dependency graph building, call-graph resolution, clustering, and documentation generation) can operate on any supported language without needing to know the language-specific parsing details. + +## Supported Languages + +| Language | Analyzer | Parsing Technology | +|---|---|---| +| C | `TreeSitterCAnalyzer` | tree-sitter (`tree_sitter_c`) | +| C++ | `TreeSitterCppAnalyzer` | tree-sitter (`tree_sitter_cpp`) | +| C# | `TreeSitterCSharpAnalyzer` | tree-sitter (`tree_sitter_c_sharp`) | +| Java | `TreeSitterJavaAnalyzer` | tree-sitter (`tree_sitter_java`) | +| JavaScript | `TreeSitterJSAnalyzer` | tree-sitter (`tree_sitter_javascript`) | +| TypeScript | `TreeSitterTSAnalyzer` | tree-sitter (`tree_sitter_typescript`) | +| PHP | `TreeSitterPHPAnalyzer` | tree-sitter (`tree_sitter_php`) | +| Python | `PythonASTAnalyzer` | Python's built-in `ast` module | + +All analyzers except the Python one are built on top of [tree-sitter](https://tree-sitter.github.io/tree-sitter/) grammars. The Python analyzer instead uses Python's native `ast` module since a fully-featured, zero-dependency parser is already available in the standard library. + +## Architecture + +Every analyzer follows the same lifecycle: + +1. **Instantiate** with `(file_path, content, repo_path)`. +2. **Parse** the file content into a syntax tree (tree-sitter grammar or Python `ast`). +3. **Extract nodes** — walk the tree once (or twice, for languages that need entity pre-collection) to identify top-level declarations and record them as `Node` objects. +4. **Extract relationships** — walk the tree again to detect calls, inheritance, instantiation, and type usage, producing `CallRelationship` objects that reference the nodes discovered in step 3. +5. **Expose results** via `self.nodes` and `self.call_relationships`, typically also through a module-level `analyze__file(...)` convenience function that returns a `(nodes, relationships)` tuple. + +```mermaid +flowchart TD + Caller["Dependency Parser"] -->|"dispatches by file extension"| Router["Language Router"] + Router -->|".c / .h"| CAnalyzer["TreeSitterCAnalyzer"] + Router -->|".cpp / .hpp / .cc"| CppAnalyzer["TreeSitterCppAnalyzer"] + Router -->|".cs"| CSharpAnalyzer["TreeSitterCSharpAnalyzer"] + Router -->|".java"| JavaAnalyzer["TreeSitterJavaAnalyzer"] + Router -->|".js / .jsx / .mjs"| JSAnalyzer["TreeSitterJSAnalyzer"] + Router -->|".ts / .tsx"| TSAnalyzer["TreeSitterTSAnalyzer"] + Router -->|".php"| PHPAnalyzer["TreeSitterPHPAnalyzer"] + Router -->|".py"| PyAnalyzer["PythonASTAnalyzer"] + + CAnalyzer --> Output["Node and CallRelationship lists"] + CppAnalyzer --> Output + CSharpAnalyzer --> Output + JavaAnalyzer --> Output + JSAnalyzer --> Output + TSAnalyzer --> Output + PHPAnalyzer --> Output + PyAnalyzer --> Output + + Output --> Builder["Dependency Graph Builder"] + Output --> CallGraph["Call Graph Analyzer"] +``` + +### Component identity (FQDN scheme) + +Every analyzer computes a **component ID** for each extracted node so that entities can be uniquely referenced across the whole repository, following the pattern: + +```text +:: +::. +``` + +For example, a method `save` inside class `UserRepository` located at `src/data/user_repository.py` becomes: + +```text +src.data.user_repository::UserRepository.save +``` + +The module path is derived by taking the file path relative to the repository root, stripping the language-specific file extension, and replacing path separators with dots. This scheme is shared by all analyzers so that downstream consumers (dependency graph construction, clustering, and documentation generation) can treat component IDs uniformly regardless of source language. + +### Two-phase extraction + +Most analyzers separate **node extraction** from **relationship extraction** into two full tree traversals: + +- The first traversal builds a lookup table of top-level names (`top_level_nodes` / `all_entities`) mapping simple names to their `Node` objects. +- The second traversal re-walks the tree, and whenever it encounters a call, instantiation, inheritance clause, or type reference, it looks up the target name in that table to decide whether the relationship is *locally resolved* (target found in the same file) or *unresolved* (left for the cross-file resolution stage to handle later, similar to how the call-graph analyzer resolves cross-file callees). + +This is important because a single file cannot know about declarations in other files; each analyzer only guarantees **intra-file** resolution and marks everything else with `is_resolved=False`, leaving further resolution to the broader dependency analysis pipeline. + +## Sub-modules + +The eight language analyzers are grouped into five documentation sub-modules based on shared structural patterns and language families: + +- [C Family Analyzers](c_family_analyzers.md) — `TreeSitterCAnalyzer`, `TreeSitterCppAnalyzer`, `TreeSitterCSharpAnalyzer`. These three analyzers share a nearly identical single-pass extraction structure with growing structural complexity (functions only → classes/namespaces → interfaces/records/delegates). +- [Java Analyzer](java_analyzer.md) — `TreeSitterJavaAnalyzer`. Handles Java's rich type system: classes, interfaces, enums, records, annotations, and detailed method/field/object-creation relationship extraction, including local-variable type inference. +- [JavaScript & TypeScript Analyzers](javascript_typescript_analyzers.md) — `TreeSitterJSAnalyzer`, `TreeSitterTSAnalyzer`. Both analyze ECMAScript-family syntax (functions, arrow functions, classes, JSDoc/TypeScript type annotations) but differ in how strictly they validate "true top-level" declarations versus nested ones. +- [PHP Analyzer](php_analyzer.md) — `TreeSitterPHPAnalyzer` (with its `NamespaceResolver` helper). Unique among the analyzers in that it must resolve PHP namespaces and `use` statement aliases to fully qualified class names before building relationships. +- [Python Analyzer](python_analyzer.md) — `PythonASTAnalyzer`. The only analyzer built on Python's native `ast` module rather than tree-sitter, using the visitor pattern (`ast.NodeVisitor`) instead of manual tree traversal. + +## Node & Relationship Type Coverage + +| Language | Node Types Extracted | Relationship Types Extracted | +|---|---|---| +| C | function, struct, variable | calls, global-variable usage | +| C++ | class, struct, function, method, namespace, variable | calls, inherits, creates (`new`), uses (variable) | +| C# | class, abstract class, static class, interface, struct, enum, record, delegate | inherits (base list), property/field/parameter type usage | +| Java | class, abstract class, interface, enum, record, annotation, method | extends, implements, field type use, method invocation, object creation | +| JavaScript | class, abstract class, interface, function, generator function, arrow function, method | calls, inheritance, JSDoc-derived type dependencies | +| TypeScript | function, class, abstract class, interface, type alias, enum, variable, export statement | calls, `new` instantiation, member access, type annotations/arguments, inheritance/implements | +| PHP | class, abstract class, interface, trait, enum, function, method | `use` imports, extends, implements, `new`, static (`::`) calls, constructor property promotion | +| Python | class, function | calls, class inheritance | + +## Relationship to the Rest of the System + +These analyzers do not run standalone — they are invoked by the dependency parsing stage of the broader dependency analyzer, which selects the appropriate analyzer per file based on its extension, aggregates the resulting nodes and relationships across the whole repository, and hands them off to graph construction and call-graph resolution. The `Node` and `CallRelationship` data structures they produce are defined in the shared dependency analyzer models, ensuring a single consistent contract across every language-specific analyzer. diff --git a/docs/reference/architecture/cli-core/cli-core.md b/docs/reference/architecture/cli-core/cli-core.md new file mode 100644 index 00000000..794ff033 --- /dev/null +++ b/docs/reference/architecture/cli-core/cli-core.md @@ -0,0 +1,103 @@ +# Cli Core + +## Overview + +The Cli Core module is the command-line interface layer of CodeWiki. It is the entry point that end users interact with when running `codewiki` from a terminal: it manages persistent user configuration, wraps git repository operations, drives the backend documentation pipeline with progress reporting, renders a static HTML viewer for GitHub Pages, and provides shared logging/progress utilities used throughout the CLI experience. + +Rather than reimplementing documentation generation, Cli Core acts as a **thin orchestration and presentation layer** on top of the backend engine (see the [Backend Core](backend-core.md) module). It translates CLI-specific concerns — credential storage, terminal progress bars, colored logging, git branch management, and static site generation — into calls against the backend's `DocumentationGenerator` and `Config` primitives. + +## Responsibilities + +- **Configuration persistence**: securely store LLM provider credentials (cluster/main/fallback API keys) in the OS keyring, and persist non-sensitive settings (models, base URLs, token limits, agent instructions) to `~/.codewiki/config.json`. +- **Job orchestration**: adapt the backend's async documentation pipeline into a CLI-friendly staged workflow (dependency analysis → module clustering → documentation generation → optional HTML generation → finalization) with verbose/non-verbose progress reporting. +- **Git integration**: detect repository state, verify a clean working tree, create timestamped documentation branches, commit generated docs, and compute GitHub PR/Pages URLs. +- **Static site generation**: render a self-contained `index.html` viewer for GitHub Pages by combining a template with the generated `module_tree.json` and `metadata.json`. +- **Job/result modeling**: define typed data models (`DocumentationJob`, `LLMConfig`, `JobStatistics`, `GenerationOptions`, `JobStatus`) that describe a documentation run's inputs, progress, and outputs, with JSON (de)serialization for the `metadata.json` artifact. +- **Terminal UX utilities**: colored logging and multi-stage progress tracking with ETA estimation, shared across all CLI commands. + +## Architecture + +```mermaid +flowchart TD + User["CLI User"] -->|"codewiki generate"| ConfigMgr["ConfigManager"] + ConfigMgr -->|"loads/saves"| ConfigFile[("~/.codewiki/config.json")] + ConfigMgr -->|"stores API keys"| Keyring[("OS Keyring")] + ConfigMgr -->|"provides"| Config["Configuration + AgentInstructions"] + + Config -->|"to_backend_config()"| CLIDocGen["CLIDocumentationGenerator"] + GitMgr["GitManager"] -->|"branch/commit info"| CLIDocGen + + CLIDocGen -->|"drives"| BackendGen["Backend DocumentationGenerator"] + CLIDocGen -->|"tracks progress via"| Progress["ProgressTracker"] + CLIDocGen -->|"builds"| Job["DocumentationJob"] + CLIDocGen -->|"optional"| HTMLGen["HTMLGenerator"] + + HTMLGen -->|"reads"| ModuleTree[("module_tree.json")] + HTMLGen -->|"reads"| Metadata[("metadata.json")] + HTMLGen -->|"writes"| IndexHTML[("index.html")] + + Job -->|"serializes to"| Metadata + + Logger["CLILogger"] -.->|"used by"| CLIDocGen + Logger -.->|"used by"| ConfigMgr + Logger -.->|"used by"| GitMgr + + BackendGen -->|"belongs to"| BackendCore["Backend Core module"] + + style BackendCore fill:#eee,stroke:#999,stroke-dasharray: 5 5 +``` + +At a high level, a CLI command (e.g. `generate`) loads user settings via `ConfigManager`, optionally inspects/manipulates the repository via `GitManager`, and then hands control to `CLIDocumentationGenerator`, which converts CLI configuration into a backend `Config` object and drives the [Backend Core](backend-core.md) pipeline stage by stage, reporting progress through `ProgressTracker`. The resulting `DocumentationJob` captures statistics and status, which are persisted as `metadata.json`. If HTML output is requested, `HTMLGenerator` renders a static viewer from the generated `module_tree.json` and `metadata.json`. + +## Sub-modules + +Cli Core is organized into the following functional areas: + +| Sub-module | Responsibility | +|---|---| +| [Generation](cli-core/generation/generation.md) | Adapts the backend documentation pipeline for CLI use with staged progress reporting and logging configuration. | +| [Configuration](cli-core/configuration/configuration.md) | Manages persistent CLI settings and secure API key storage via the OS keyring; defines the configuration data model. | +| [Job Models](cli-core/job_models/job_models.md) | Typed data models describing a documentation job's status, statistics, and LLM configuration, with JSON serialization. | +| [Git Integration](cli-core/git_integration/git_integration.md) | Wraps git operations needed for documentation branch workflows (clean-check, branch creation, commit, remote/PR URL detection). | +| [Html Generation](cli-core/html_generation/html_generation.md) | Renders a static, self-contained HTML documentation viewer for GitHub Pages. | +| [Utils](cli-core/utils/utils.md) | Shared terminal UX helpers: colored logging and multi-stage progress tracking with ETA. | + +### Generation + +The [Generation](cli-core/generation/generation.md) sub-module contains `CLIDocumentationGenerator`, the central adapter that bridges CLI configuration and the backend engine. It normalizes additional source paths, builds a backend `Config`, configures backend logging with colored output, and runs the five-stage pipeline (dependency analysis, module clustering, documentation generation, optional HTML generation, finalization), reporting progress and raising `APIError` on failure. + +### Configuration + +The [Configuration](cli-core/configuration/configuration.md) sub-module contains `ConfigManager` together with the `Configuration` and `AgentInstructions` data models. `ConfigManager` persists non-sensitive settings to `~/.codewiki/config.json` and stores per-provider API keys (cluster/main/fallback) in the OS keyring, falling back gracefully when the keyring is unavailable. `Configuration.to_backend_config()` is the bridge that converts persisted CLI settings into a backend `Config` instance for a specific run. + +### Job Models + +The [Job Models](cli-core/job_models/job_models.md) sub-module defines `DocumentationJob` and its supporting types (`JobStatus`, `JobStatistics`, `GenerationOptions`, `LLMConfig`). These models track a single documentation run's lifecycle (pending → running → completed/failed), record statistics such as files analyzed and leaf nodes, and support round-trip JSON serialization used for the `metadata.json` artifact. + +### Git Integration + +The [Git Integration](cli-core/git_integration/git_integration.md) sub-module contains `GitManager`, which wraps the `git` Python library to support the "create documentation branch and commit" workflow: checking for a clean working directory, generating timestamped branch names, committing generated docs, and deriving GitHub remote/PR URLs. + +### Html Generation + +The [Html Generation](cli-core/html_generation/html_generation.md) sub-module contains `HTMLGenerator`, which loads `module_tree.json` and `metadata.json` from a documentation output directory, populates a bundled HTML template, and writes a self-contained `index.html` suitable for GitHub Pages hosting. + +### Utils + +The [Utils](cli-core/utils/utils.md) sub-module contains `CLILogger`, `ProgressTracker`, and `ModuleProgressBar` — shared terminal presentation helpers used across the CLI for colored log output and multi-stage progress bars with ETA estimation. + +## Relationship to Other Modules + +- **Backend Core**: Cli Core does not implement documentation generation itself; it delegates to the backend's `DocumentationGenerator`, dependency analyzer, and LLM services. See the [Backend Core](backend-core.md) module for details on dependency analysis, clustering, and agent orchestration. +- **Config Core**: The backend-facing `Config` object that `Configuration.to_backend_config()` produces is defined in the shared configuration module. See [Config Core](config-core.md) for details on the runtime configuration surface consumed by the backend. + +## Error Handling + +Cli Core raises a small hierarchy of typed exceptions used to communicate failures to the CLI entry point with distinct exit codes: + +- `ConfigurationError` — raised by `ConfigManager` when configuration cannot be loaded/saved or the keychain is unavailable. +- `RepositoryError` — raised by `GitManager` when git operations fail (e.g., dirty working directory, invalid repository). +- `FileSystemError` — raised by `HTMLGenerator` and configuration I/O helpers when reading/writing files fails. +- `APIError` — raised by `CLIDocumentationGenerator` when a backend/LLM stage (dependency analysis, clustering, documentation generation) fails. + +Each exception type maps to a dedicated process exit code, allowing CLI scripts and CI pipelines to distinguish between configuration problems, repository issues, filesystem errors, and API failures. diff --git a/docs/reference/architecture/cli-core/cli-core/configuration/configuration.md b/docs/reference/architecture/cli-core/cli-core/configuration/configuration.md new file mode 100644 index 00000000..7ed25c0a --- /dev/null +++ b/docs/reference/architecture/cli-core/cli-core/configuration/configuration.md @@ -0,0 +1,249 @@ +# Configuration + +The Configuration module provides the persistent settings layer for the CodeWiki CLI. It defines the data models that describe user preferences (LLM providers, model names, token limits, temperature settings, and documentation-generation instructions) and the manager responsible for reading and writing those settings safely to disk and to the operating system's secure credential store. + +This module is a child of [Cli Core](../../cli-core.md) and works alongside sibling modules such as the job/runtime models, generation pipeline, and utility helpers to support the CLI's end-to-end workflow. + +## Purpose and Scope + +The Configuration module answers three core questions for the CLI: + +1. **What settings does the user have configured?** — captured by the `Configuration` dataclass. +2. **How should the documentation agent behave for a given run?** — captured by `AgentInstructions`. +3. **How are these settings persisted, validated, and loaded securely?** — handled by `ConfigManager`. + +Sensitive values (API keys) are never written to plaintext configuration files. Instead, they are stored using the system keyring (macOS Keychain, Windows Credential Manager, or Linux Secret Service) and only non-sensitive settings are persisted to `~/.codewiki/config.json`. + +## Core Components + +| Component | Responsibility | +|---|---| +| `Configuration` | Dataclass representing all persistent CLI settings: model names, base URLs, API versions, token/temperature limits, and clustering parameters. | +| `AgentInstructions` | Dataclass representing optional, user-customizable instructions for the documentation agent (file filters, focus modules, doc type, free-form instructions). | +| `ConfigManager` | Orchestrates loading/saving `Configuration` to `~/.codewiki/config.json` and securely storing/retrieving API keys via keyring. | + +## Architecture Overview + +```mermaid +flowchart TD + subgraph ConfigModule["Configuration Module"] + CM["ConfigManager"] + Cfg["Configuration"] + AI["AgentInstructions"] + end + + FS["~/.codewiki/config.json"] + KR["System Keyring"] + Backend["Backend Config"] + + CM -->|"load() / save()"| FS + CM -->|"get/set API keys"| KR + CM -->|"holds"| Cfg + Cfg -->|"has one"| AI + Cfg -->|"to_backend_config()"| Backend +``` + +`Configuration` is a plain, serializable dataclass with no direct dependency on keyring or the filesystem — those concerns are owned exclusively by `ConfigManager`. This separation keeps the data model easy to test and reuse, while `ConfigManager` acts as the single access point for persistence. + +## Data Model: Configuration + +`Configuration` captures all settings needed to drive a documentation-generation run, organized into three provider "roles": + +- **cluster** — model used for module clustering/decomposition +- **main** — primary model used for documentation generation +- **fallback** — fallback model used when the main model fails or is rate-limited + +For each role, the model tracks: +- Model name (`*_model`) +- Base URL (`*_base_url`, optional — for OpenAI-compatible or self-hosted endpoints) +- API version (`*_api_version`, optional) +- Max tokens (`*_max_tokens`) +- Temperature (`*_temperature`) and whether temperature is supported (`*_temperature_supported`) +- Max-token parameter field name (`*_max_token_field`, e.g., `"max_tokens"` vs. provider-specific names) + +In addition to per-provider settings, `Configuration` holds shared clustering parameters (`max_token_per_module`, `max_token_per_leaf_module`, `max_depth`) and a `default_output` directory. API keys (`cluster_api_key`, `main_api_key`, `fallback_api_key`) exist on the dataclass only as **runtime-only fields** — they are never serialized by `to_dict()` and are populated separately from the keyring by `ConfigManager`. + +```mermaid +classDiagram + class Configuration { + +str main_model + +str cluster_model + +str fallback_model + +str default_output + +str cluster_base_url + +str main_base_url + +str fallback_base_url + +int cluster_max_tokens + +int main_max_tokens + +int fallback_max_tokens + +float cluster_temperature + +float main_temperature + +float fallback_temperature + +bool cluster_temperature_supported + +bool main_temperature_supported + +bool fallback_temperature_supported + +int max_token_per_module + +int max_token_per_leaf_module + +int max_depth + +AgentInstructions agent_instructions + +validate() + +to_dict() dict + +from_dict(data) Configuration + +is_complete() bool + +to_backend_config(...) Config + } + + class AgentInstructions { + +List~str~ include_patterns + +List~str~ exclude_patterns + +List~str~ focus_modules + +str doc_type + +str custom_instructions + +to_dict() dict + +from_dict(data) AgentInstructions + +is_empty() bool + +get_prompt_addition() str + } + + Configuration "1" *-- "1" AgentInstructions : agent_instructions +``` + +### Validation + +`Configuration.validate()` performs field-level checks: +- Base URLs (when set) are validated via `validate_url`. +- Model names for all three roles are validated via `validate_model_name`. + +`from_dict()` performs defensive type coercion when loading from JSON (which may contain string-typed numbers/booleans due to manual editing), converting and range-checking integers (e.g., token limits), floats (temperature, bounded `0.0`–`2.0`), and booleans, raising `ValueError` on invalid input. + +### Serialization Rules + +- `to_dict()` only emits optional fields (`*_base_url`, `*_api_version`, `agent_instructions`) when they are set/non-empty, keeping the persisted JSON minimal. +- API keys are **excluded** from `to_dict()` entirely — they are never written to `config.json`. + +## Data Model: AgentInstructions + +`AgentInstructions` lets users customize how the documentation agent analyzes a repository and generates content: + +- `include_patterns` / `exclude_patterns` — glob-style file filters (e.g., `["*.cs"]`, `["*Tests*"]`) +- `focus_modules` — modules that should receive more detailed documentation +- `doc_type` — a preset documentation style (`api`, `architecture`, `user-guide`, `developer`) or a free-form type +- `custom_instructions` — arbitrary additional guidance passed to the LLM + +`get_prompt_addition()` translates these fields into a natural-language instruction block that is merged into the agent's prompt at generation time. `is_empty()` allows callers to distinguish "no customization" from "customization with all-default values." + +Instructions can be set persistently (stored in `config.json` as part of `Configuration`) or supplied at runtime for a single job; `Configuration.to_backend_config()` merges the two, with runtime instructions taking precedence field-by-field. + +## ConfigManager: Persistence and Secure Storage + +`ConfigManager` is the sole component responsible for reading and writing configuration state. It manages two independent storage backends: + +1. **JSON file** (`~/.codewiki/config.json`) — non-sensitive settings, versioned with a `CONFIG_VERSION` marker for future migrations. +2. **System keyring** (via the `keyring` library) — the three provider API keys (`cluster_api_key`, `main_api_key`, `fallback_api_key`), stored under a shared `codewiki` service name with distinct account identifiers. + +```mermaid +classDiagram + class ConfigManager { + -Optional~str~ _cluster_api_key + -Optional~str~ _main_api_key + -Optional~str~ _fallback_api_key + -Optional~Configuration~ _config + -bool _keyring_available + +load() bool + +save(...) void + +get_cluster_api_key() Optional~str~ + +get_main_api_key() Optional~str~ + +get_fallback_api_key() Optional~str~ + +get_config() Optional~Configuration~ + +is_configured() bool + +delete_api_keys() void + +clear() void + +keyring_available bool + +config_file_path Path + } + ConfigManager --> Configuration : loads/saves +``` + +### Load Flow + +```mermaid +sequenceDiagram + participant Caller + participant CM as ConfigManager + participant FS as "config.json" + participant KR as "System Keyring" + + Caller->>CM: load() + CM->>FS: check exists / read + alt file missing + FS-->>CM: not found + CM-->>Caller: False + else file present + FS-->>CM: JSON content + CM->>CM: Configuration.from_dict(data) + CM->>KR: get_password(cluster_api_key) + CM->>KR: get_password(main_api_key) + CM->>KR: get_password(fallback_api_key) + KR-->>CM: key values (or None) + CM-->>Caller: True + end +``` + +### Save Flow + +```mermaid +sequenceDiagram + participant Caller + participant CM as ConfigManager + participant FS as "config.json" + participant KR as "System Keyring" + + Caller->>CM: save(fields..., api_keys...) + CM->>FS: ensure_directory(CONFIG_DIR) + alt no in-memory config + CM->>CM: load() existing or create default Configuration + end + CM->>CM: apply provided field updates + CM->>CM: Configuration.validate() (if models set) + CM->>KR: set_password(...) for each provided API key + CM->>FS: write JSON (version + Configuration.to_dict()) + CM-->>Caller: done (or raises ConfigurationError) +``` + +Key behaviors: +- **Partial updates**: `save()` accepts every field as an optional keyword argument; only provided values overwrite the in-memory `Configuration`, allowing incremental configuration (e.g., `codewiki configure` sub-commands that set one field at a time). +- **Validation gate**: full validation (`Configuration.validate()`) only runs once `main_model` and `cluster_model` are both set, avoiding premature failures during multi-step setup. +- **Keyring failures surface as `ConfigurationError`**, with an actionable message when the OS keychain is unavailable or misconfigured. +- **`is_configured()`** combines two checks: all three API keys must be retrievable from keyring, and `Configuration.is_complete()` must be true (all three model names set). +- **`clear()`** performs a full reset — deleting API keys from keyring and removing `config.json` — used by commands like `codewiki configure --reset`. + +## Bridging to Runtime Execution: to_backend_config + +`Configuration.to_backend_config()` is the seam between this module's persistent settings and the runtime configuration consumed by the documentation-generation backend. It: + +1. Fetches any missing API keys from the keyring via a fresh `ConfigManager` instance (if not explicitly passed in). +2. Merges `runtime_instructions` (per-invocation `AgentInstructions`) over the persisted `agent_instructions`, with runtime values taking precedence field-by-field. +3. Constructs and returns a backend `Config` object (`Config.from_cli(...)`) populated with all model, token, temperature, and clustering settings, ready to drive a documentation job. + +```mermaid +flowchart LR + A["CLI Command"] --> B["ConfigManager.load()"] + B --> C["Configuration"] + C --> D["Configuration.to_backend_config()"] + D --> E["ConfigManager (keyring lookup for missing keys)"] + D --> F["Merge AgentInstructions (runtime over persisted)"] + D --> G["Config.from_cli(...)"] + G --> H["Backend Config"] +``` + +This bridging pattern keeps the CLI's persistent, user-facing settings model (`Configuration`) decoupled from the backend's execution-time `Config` model. The backend `Config` class itself is documented in the [Config Core](../../config-core.md) module. + +## Relationship to Other CLI Modules + +- **[Job Models](../job_models/job_models.md)** — represents the runtime state of an in-progress documentation job (`DocumentationJob`, `JobStatus`, `LLMConfig`, `GenerationOptions`, `JobStatistics`). Where `Configuration` describes *persistent user preferences*, the job models describe the *live execution* of a single generation run, often derived from a `Configuration` via `to_backend_config()`. +- **[Generation](../generation/generation.md)** — the `CLIDocumentationGenerator` adapter consumes a resolved `Configuration`/backend `Config` to drive the documentation pipeline. +- **[Utils](../utils/utils.md)** — provides shared filesystem (`ensure_directory`, `safe_read`, `safe_write`) and error types (`ConfigurationError`, `FileSystemError`) used internally by `ConfigManager`. + +## Summary + +The Configuration module is the trust boundary for user secrets and the single source of truth for persistent CLI preferences. `Configuration` and `AgentInstructions` define what can be configured; `ConfigManager` defines how those settings are safely loaded, validated, updated, and translated into a form the backend documentation pipeline can execute against. diff --git a/docs/reference/architecture/cli-core/cli-core/generation/generation.md b/docs/reference/architecture/cli-core/cli-core/generation/generation.md new file mode 100644 index 00000000..17120aeb --- /dev/null +++ b/docs/reference/architecture/cli-core/cli-core/generation/generation.md @@ -0,0 +1,167 @@ +# Generation + +The Generation module is the central orchestration layer of the CodeWiki CLI. It bridges the command-line interface with the backend documentation engine, coordinating the full end-to-end pipeline that turns a source code repository into a structured, hierarchical documentation set — including dependency analysis, LLM-driven module clustering, documentation generation, optional Mermaid diagram extraction, and optional HTML output. + +Its single core component, `CLIDocumentationGenerator`, acts as an **adapter**: it wraps the backend's `DocumentationGenerator` (from [Backend Core](../../backend-core.md)) with CLI-specific concerns such as progress reporting, colored logging, verbose diagnostics, and job lifecycle tracking via the [Job Models](../job_models/job_models.md) module. + +## Purpose and Role in the System + +As a child of [Cli Core](../../cli-core.md), the Generation module is invoked whenever a user runs a documentation generation command. It does not implement dependency analysis, clustering, or documentation writing itself — those responsibilities belong to `backend-core`. Instead, Generation is responsible for: + +1. **Translating CLI configuration** into the backend's `Config` object (model selection, API keys, base URLs, token limits, clustering depth, agent instructions, additional source paths). +2. **Orchestrating the multi-stage pipeline** (dependency analysis → module clustering → documentation generation → optional HTML generation → finalization) with a five-stage `ProgressTracker`. +3. **Tracking job state** using a `DocumentationJob` model, recording statistics, generated files, and success/failure outcomes. +4. **Handling synthetic module fallback** to avoid context-window overflows when the LLM clustering step returns an empty module tree. +5. **Optionally generating an HTML viewer** via [Html Generation](../html_generation/html_generation.md) and extracting Mermaid diagrams into a dedicated directory. + +## Architecture Overview + +```mermaid +flowchart TD + CLI["CLI Command Entry Point"] --> CDG["CLIDocumentationGenerator"] + CDG --> PT["ProgressTracker (utils)"] + CDG --> Job["DocumentationJob (job_models)"] + CDG --> BC["BackendConfig (config-core)"] + CDG --> DG["DocumentationGenerator (backend-core)"] + DG --> GB["DependencyGraphBuilder"] + DG --> AO["AgentOrchestrator"] + CDG --> CM["cluster_modules (backend-core)"] + CDG --> HG["HTMLGenerator (html_generation)"] + CDG --> LOG["ColoredFormatter (backend-core logging)"] +``` + +## Core Component + +### CLIDocumentationGenerator + +`CLIDocumentationGenerator` (`codewiki/cli/adapters/doc_generator.py`) is instantiated with: + +- `repo_path` — path to the repository being documented +- `output_dir` — target directory for generated docs +- `config` — a dictionary of LLM/model configuration (models, API keys, base URLs, token limits, temperature settings, clustering parameters, agent instructions, additional source paths) +- `verbose` — whether to emit detailed diagnostic output +- `generate_html` — whether to produce an HTML viewer (`index.html`) +- `diagrams_dir` — optional separate output directory for extracted Mermaid diagrams + +On construction, it: +- Creates a `ProgressTracker` configured for 5 pipeline stages. +- Creates a `DocumentationJob` and populates its metadata (repository path/name, output directory, `LLMConfig`). +- Configures backend logging by attaching a `ColoredFormatter`-based handler to the `codewiki.src.be` logger namespace, switching verbosity between `INFO` (verbose) and `WARNING` (quiet) levels, and disabling propagation to avoid duplicate log lines. + +#### Key Responsibilities + +| Responsibility | Method | +|---|---| +| Build backend configuration from CLI config dict | `generate()` | +| Run the full async generation pipeline | `_run_backend_generation()` | +| Generate the HTML viewer | `_run_html_generation()` | +| Ensure job metadata file exists | `_finalize_job()` | +| Configure colored backend logging | `_configure_backend_logging()` | + +## Generation Pipeline + +The `generate()` method is the public entry point. It is synchronous from the caller's perspective but internally drives an `asyncio`-based backend pipeline. It returns a completed `DocumentationJob` (see [Job Models](../job_models/job_models.md)) or raises an `APIError` on failure. + +```mermaid +sequenceDiagram + participant Caller as "CLI Command" + participant CDG as "CLIDocumentationGenerator" + participant BC as "BackendConfig" + participant DG as "DocumentationGenerator" + participant CM as "cluster_modules" + participant HG as "HTMLGenerator" + participant Job as "DocumentationJob" + + Caller->>CDG: generate() + CDG->>Job: start() + CDG->>BC: Config.from_cli(...) + CDG->>DG: _run_backend_generation(backend_config) + DG->>DG: graph_builder.build_dependency_graph() + DG->>CM: cluster_modules(leaf_nodes, components, config) + Note over DG,CM: Synthetic module fallback if tree is empty + DG->>DG: generate_module_documentation(components, leaf_nodes) + opt diagrams_dir configured + DG->>DG: extract_and_save_mermaid_diagrams() + end + opt generate_html is true + CDG->>HG: generate(output_path, ...) + end + CDG->>CDG: _finalize_job() + CDG->>Job: complete() + CDG-->>Caller: DocumentationJob +``` + +### Stage 1 — Dependency Analysis + +Instantiates the backend `DocumentationGenerator` (which internally wires up `DependencyGraphBuilder` and `AgentOrchestrator` from [Backend Core](../../backend-core.md)) and calls `doc_generator.graph_builder.build_dependency_graph()`. The result is a map of `components` and a list of `leaf_nodes`. Statistics (`total_files_analyzed`, `leaf_nodes`) are recorded on the `DocumentationJob`. Failures are wrapped as `APIError("Dependency analysis failed: ...")`. + +### Stage 2 — Module Clustering + +Loads a cached `first_module_tree.json` if present, otherwise calls `cluster_modules(leaf_nodes, components, backend_config)` to invoke the LLM clustering model. The result is cached to disk (`first_module_tree_path`) and then persisted as the working module tree (`module_tree_path`). + +**Synthetic Module Patch**: If the module tree ends up empty despite having leaf nodes (to prevent an LLM "whole-repo" fallback that could exceed API context limits), the generator batches leaf nodes into synthetic modules of a configurable size (`CODEWIKI_MAX_FILES_PER_MODULE`, default `5`) and re-persists the tree. This safeguard applies even when loading from cache, closing a previously identified cache-bypass gap. + +The final module count is stored on the `DocumentationJob.module_count` field. + +### Stage 3 — Documentation Generation + +Calls `doc_generator.generate_module_documentation(components, leaf_nodes)`, which performs the topologically-ordered (leaf-first) generation of module documentation via the `AgentOrchestrator`. After generation: + +- `doc_generator.create_documentation_metadata(...)` writes `metadata.json`. +- Generated `.md` and `.json` files in the output directory are collected into `DocumentationJob.files_generated`. +- If `diagrams_dir` was configured, Mermaid diagrams embedded in generated markdown are extracted via `extract_and_save_mermaid_diagrams` and indexed with `create_diagrams_readme`. + +### Stage 4 — HTML Generation (Optional) + +If `generate_html=True`, `_run_html_generation()` uses the [Html Generation](../html_generation/html_generation.md) module's `HTMLGenerator` to detect repository info (name, URL, GitHub Pages URL) and render `index.html`, auto-loading the module tree and metadata from the output directory. The generated file is appended to `DocumentationJob.files_generated`. + +### Stage 5 — Finalization + +`_finalize_job()` verifies that `metadata.json` exists in the output directory; if the backend did not already write it, the generator writes the job's own JSON representation (`DocumentationJob.to_json()`) as a fallback. + +## Configuration Mapping + +`generate()` builds the backend `Config` object (see [Config Core](../../config-core.md)) via `Config.from_cli(...)`, translating the CLI's flat configuration dictionary into per-provider settings: + +```mermaid +flowchart LR + subgraph CLIConfig["CLI config dict"] + A1["main_model / cluster_model / fallback_model"] + A2["*_api_key"] + A3["*_base_url"] + A4["*_api_version"] + A5["*_max_tokens / *_temperature"] + A6["max_token_per_module / max_depth"] + A7["agent_instructions"] + A8["additional_paths"] + end + CLIConfig --> FromCLI["Config.from_cli()"] + FromCLI --> BackendConfig["Backend Config instance"] + BackendConfig --> DG2["DocumentationGenerator"] +``` + +Additional source paths supplied in the CLI config are normalized to absolute paths (resolved relative to `repo_path`) before being passed through as `additional_source_paths`. + +In verbose mode, the generator prints a detailed configuration summary (model names, base URLs, token limits, module settings, additional paths, and a preview of custom agent instructions) before kicking off Stage 1. + +## Progress Tracking and Logging + +The Generation module relies on utilities from the [Utils](../utils/utils.md) module: + +- **`ProgressTracker`**: manages a 5-stage weighted progress model (Dependency Analysis 40%, Module Clustering 20%, Documentation Generation 30%, HTML Generation 5%, Finalization 5%), providing `start_stage`, `update_stage`, `complete_stage`, elapsed-time formatting, and ETA estimation. +- **`CLILogger`**: complements verbose/non-verbose console output alongside `ProgressTracker`'s stage banners. + +Backend logs are captured by attaching a handler using `ColoredFormatter` (from `backend-core`'s Logging Config child module) directly to the `codewiki.src.be` logger, ensuring consistent colored output regardless of whether the CLI or backend emitted the log line, while preventing duplicate propagation to the root logger. + +## Error Handling + +All backend-facing calls (`build_dependency_graph`, `cluster_modules`, `generate_module_documentation`) are wrapped in `try`/`except` blocks that convert unexpected exceptions into `APIError` with contextual messages (e.g., `"Dependency analysis failed: ..."`). At the top level, `generate()` catches both `APIError` and generic `Exception`, marks the `DocumentationJob` as failed via `job.fail(str(e))`, and re-raises so the CLI layer can present the error to the user. + +## Relationship to Other Modules + +- **[Cli Core](../../cli-core.md)**: parent module; Generation is one of its functional children alongside [Configuration](../configuration/configuration.md), [Job Models](../job_models/job_models.md), [Git Integration](../git_integration/git_integration.md), [Html Generation](../html_generation/html_generation.md), and [Utils](../utils/utils.md). +- **[Job Models](../job_models/job_models.md)**: supplies `DocumentationJob`, `LLMConfig`, `JobStatus`, and `JobStatistics`, which Generation populates throughout the pipeline. +- **[Html Generation](../html_generation/html_generation.md)**: invoked optionally at Stage 4 to render the documentation as a browsable HTML site. +- **[Utils](../utils/utils.md)**: provides `ProgressTracker` for stage-based progress reporting. +- **[Backend Core](../../backend-core.md)**: supplies the actual documentation engine (`DocumentationGenerator`, `AgentOrchestrator`, `DependencyGraphBuilder`) that Generation orchestrates but does not reimplement. +- **[Config Core](../../config-core.md)**: supplies the `Config` class used to translate CLI settings into the backend's runtime configuration. diff --git a/docs/reference/architecture/cli-core/cli-core/git_integration/git_integration.md b/docs/reference/architecture/cli-core/cli-core/git_integration/git_integration.md new file mode 100644 index 00000000..41acdeb2 --- /dev/null +++ b/docs/reference/architecture/cli-core/cli-core/git_integration/git_integration.md @@ -0,0 +1,147 @@ +# Git Integration + +The Git Integration module provides the `GitManager` component, which encapsulates all git repository operations required by the CodeWiki CLI's documentation workflow. It is responsible for validating that a target directory is a git repository, inspecting working-directory cleanliness, creating dedicated documentation branches, committing generated documentation, and deriving useful metadata such as remote URLs, branch names, commit hashes, and GitHub pull-request links. + +This module is a child of the [Cli Core](../../cli-core.md) module and is used primarily by the CLI's generation workflow (see [Generation](../generation/generation.md)) to safely manage version control state before, during, and after documentation is generated. + +## Purpose and Scope + +When the CLI is run with branch-creation options (e.g. `--create-branch`), it must: + +1. Confirm the target path is inside a valid git repository. +2. Ensure there are no uncommitted changes that could be silently mixed into a documentation commit (unless the user explicitly forces the operation). +3. Create a uniquely named, timestamped branch dedicated to documentation output. +4. Stage and commit the generated documentation files. +5. Surface repository metadata (current branch, commit hash, remote URL) back to the CLI so it can print helpful summaries. +6. Build a ready-to-open GitHub pull-request URL if the remote is hosted on GitHub. + +All of this logic is centralized in the `GitManager` class, keeping git operations isolated from the rest of the CLI so that other components (progress reporting, configuration, documentation generation) do not need to know about the underlying git library (`GitPython`). + +## Core Component: GitManager + +`GitManager` wraps a `git.Repo` instance (from the `GitPython` library) and exposes a small, purpose-built API surface consumed by the CLI layer. + +### Initialization + +On construction, `GitManager` resolves the given `repo_path` to an absolute path and attempts to open it as a git repository using `git.Repo(repo_path, search_parent_directories=True)`. The `search_parent_directories=True` flag allows the manager to locate a `.git` directory even if `repo_path` points to a subdirectory of the repository (similar to how the native `git` command walks upward from the current directory). + +If no valid repository is found, a `RepositoryError` is raised with actionable guidance (e.g., suggesting `git init`). `RepositoryError` is a CLI-specific exception type that carries a dedicated exit code (`EXIT_REPOSITORY_ERROR`), allowing the CLI's top-level error handler to report failures consistently and set correct process exit statuses. + +### Key Responsibilities + +#### 1. Working Directory Cleanliness Check + +`check_clean_working_directory()` inspects the repository for uncommitted modifications or untracked files using `repo.is_dirty(untracked_files=True)`. When dirty, it builds a human-readable summary listing up to three modified and three untracked files (with a "... and N more" suffix for longer lists), returning a `(is_clean, status_message)` tuple. This check underpins the safety guarantee that documentation generation will not accidentally commit unrelated in-progress changes. + +#### 2. Documentation Branch Creation + +`create_documentation_branch(force: bool = False)` creates and checks out a new branch named `docs/codewiki-` (format `%Y%m%d-%H%M%S`). Before creating the branch it re-runs the cleanliness check unless `force=True`, raising a detailed `RepositoryError` with copy-pastable remediation commands (`git status`, `git add -A && git commit`, `git stash`) if the working directory is dirty. + +To guard against timestamp collisions, it checks existing branch names and appends an incrementing numeric suffix (`-1`, `-2`, ...) if a collision would otherwise occur. Branch creation and checkout are performed via `repo.create_head(...)` followed by `.checkout()`; any underlying `GitCommandError` is translated into a `RepositoryError`. + +#### 3. Committing Documentation + +`commit_documentation(docs_path: Path, message: Optional[str] = None)` stages the documentation output directory via `repo.index.add([str(docs_path)])` and commits it with either a caller-supplied message or the default `"Add generated documentation\n\nGenerated by CodeWiki CLI"`. It returns the resulting commit's `hexsha`. Failures during add/commit are wrapped in `RepositoryError`. + +#### 4. Repository Metadata Accessors + +- `get_remote_url(remote_name="origin")` — returns the URL of the named remote, or `None` if it does not exist. +- `get_current_branch()` — returns the active branch name, or the literal string `"HEAD"` when in a detached-HEAD state (caught via `TypeError` from `GitPython`). +- `get_commit_hash()` — returns the current `HEAD` commit's hex SHA. +- `branch_exists(branch_name)` — returns whether a branch with the given name already exists locally. + +#### 5. GitHub Pull-Request URL Construction + +`get_github_pr_url(branch_name)` derives a ready-to-use GitHub "compare" URL (`/compare/`) from the `origin` remote, but only when that remote points to `github.com`. It normalizes the remote URL by: +- Stripping a trailing slash and any `.git` suffix. +- Converting SSH-style remotes (`git@github.com:org/repo`) into HTTPS form (`https://github.com/org/repo`). + +If there is no remote or it is not a GitHub remote, it returns `None`, allowing the CLI to gracefully skip printing a PR link for non-GitHub repositories. + +## Error Handling + +All git-related failures surface as `RepositoryError` (defined outside this module in the CLI's shared error utilities), rather than leaking raw `GitCommandError` or `InvalidGitRepositoryError` exceptions from `GitPython`. This gives the CLI a single, well-known exception type to catch at the top level and translate into a clean error message and a dedicated process exit code, keeping git-library implementation details out of user-facing error handling. + +## Architecture + +```mermaid +classDiagram + class GitManager { + +Path repo_path + +Repo repo + +__init__(repo_path) + +check_clean_working_directory() Tuple + +create_documentation_branch(force) str + +commit_documentation(docs_path, message) str + +get_remote_url(remote_name) str + +get_current_branch() str + +get_commit_hash() str + +branch_exists(branch_name) bool + +get_github_pr_url(branch_name) str + } + class RepositoryError { + +__init__(message) + } + GitManager ..> RepositoryError : raises +``` + +## Interaction with the CLI Generation Workflow + +`GitManager` is typically instantiated by the CLI's generation flow when branch-based documentation workflows are requested. The sequence below illustrates a typical end-to-end interaction: verifying repository state, creating a branch, running documentation generation (handled by `CLIDocumentationGenerator` in the [Generation](../generation/generation.md) module), committing the results, and reporting a PR link. + +```mermaid +sequenceDiagram + participant CLI as "CLI Entry Point" + participant GM as "GitManager" + participant DocGen as "CLIDocumentationGenerator" + participant Repo as "git.Repo" + + CLI->>GM: "GitManager(repo_path)" + GM->>Repo: "git.Repo(repo_path)" + CLI->>GM: "check_clean_working_directory()" + GM-->>CLI: "(is_clean, status_message)" + alt "Working directory dirty and not forced" + GM-->>CLI: "raise RepositoryError" + else "Clean or forced" + CLI->>GM: "create_documentation_branch(force)" + GM->>Repo: "create_head(branch_name)" + GM->>Repo: "checkout()" + GM-->>CLI: "branch_name" + CLI->>DocGen: "generate()" + DocGen-->>CLI: "DocumentationJob" + CLI->>GM: "commit_documentation(docs_path, message)" + GM->>Repo: "index.add([docs_path])" + GM->>Repo: "index.commit(message)" + GM-->>CLI: "commit_hash" + CLI->>GM: "get_github_pr_url(branch_name)" + GM-->>CLI: "pr_url or None" + end +``` + +## Process Flow: Documentation Branch Creation + +```mermaid +flowchart TD + Start["create_documentation_branch(force)"] --> CheckForce{{force?}} + CheckForce -->|"false"| CheckClean["check_clean_working_directory()"] + CheckForce -->|"true"| GenName["Generate timestamped branch name"] + CheckClean --> IsClean{{"is_clean?"}} + IsClean -->|"false"| RaiseErr["raise RepositoryError with remediation steps"] + IsClean -->|"true"| GenName + GenName --> CheckExists{{"branch name exists?"}} + CheckExists -->|"true"| AppendCounter["Append incrementing counter suffix"] + AppendCounter --> CheckExists + CheckExists -->|"false"| CreateBranch["repo.create_head(branch_name)"] + CreateBranch --> Checkout["new_branch.checkout()"] + Checkout --> ReturnName["return branch_name"] +``` + +## Relationship to Other Modules + +- **[Cli Core](../../cli-core.md)** — the parent module; `GitManager` is one of the core services orchestrated alongside configuration management, job models, and progress tracking to deliver the full `codewiki generate` experience. +- **[Generation](../generation/generation.md)** — `CLIDocumentationGenerator` performs the actual documentation generation (dependency analysis, clustering, LLM-driven writing, optional HTML output). The CLI entry point coordinates `GitManager` and `CLIDocumentationGenerator` together: creating a branch before generation, and committing the resulting files afterward. +- **[Job Models](../job_models/job_models.md)** — while `GitManager` itself does not depend on `DocumentationJob`, the surrounding CLI workflow often records git-derived metadata (branch name, commit hash) alongside job statistics for reporting purposes. + +## Dependencies + +`GitManager` depends on the third-party `GitPython` library (imported as `git`) for all low-level repository interactions, including `git.Repo`, `git.InvalidGitRepositoryError`, and `git.exc.GitCommandError`. It also depends on the CLI's shared `RepositoryError` exception type for consistent error reporting across the CLI. diff --git a/docs/reference/architecture/cli-core/cli-core/html_generation/html_generation.md b/docs/reference/architecture/cli-core/cli-core/html_generation/html_generation.md new file mode 100644 index 00000000..f11bd2a0 --- /dev/null +++ b/docs/reference/architecture/cli-core/cli-core/html_generation/html_generation.md @@ -0,0 +1,131 @@ +# Html Generation + +The Html Generation module produces a self-contained, static HTML documentation viewer suitable for GitHub Pages (or any static file host). It is the final, optional presentation layer of the CodeWiki CLI pipeline: after the [Generation](../generation/generation.md) stage has produced Markdown documentation files, a module tree, and metadata, the Html Generation module packages that output into a single browsable `index.html` file with embedded configuration, styles, and client-side rendering logic. + +## Purpose and Scope + +The `HTMLGenerator` class is the sole core component of this module. Its responsibilities are: + +- **Template loading** — reads a static HTML template (`viewer_template.html`) shipped with the package. +- **Data discovery** — auto-loads `module_tree.json` and `metadata.json` from a documentation output directory when explicit data is not supplied. +- **Placeholder substitution** — injects title, repository link, embedded JSON data (module tree, metadata, config), and an "info panel" HTML fragment into the template. +- **Repository introspection** — inspects a local git repository to derive a display name, remote URL, and a predicted GitHub Pages URL. +- **Atomic file output** — writes the final HTML file safely using the shared filesystem utilities. + +This module has no knowledge of *how* documentation content was produced; it only consumes the artifacts (`module_tree.json`, `metadata.json`, generated Markdown files) that the [Generation](../generation/generation.md) module and the backend documentation pipeline produce. This keeps Html Generation a pure "rendering/packaging" concern, decoupled from LLM orchestration, dependency analysis, and job tracking. + +## Core Component + +### HTMLGenerator + +`HTMLGenerator` (`codewiki/cli/html_generator.py`) encapsulates all HTML viewer generation logic. + +| Method | Responsibility | +|---|---| +| `__init__(template_dir)` | Resolves the template directory, defaulting to the package's `templates/github_pages` folder. | +| `load_module_tree(docs_dir)` | Reads `module_tree.json` from the docs directory; falls back to a minimal single-node structure if the file is missing. | +| `load_metadata(docs_dir)` | Reads `metadata.json`; returns `None` (non-critical) if missing or unparsable. | +| `generate(...)` | Orchestrates the full generation: auto-loads data, builds the info panel, computes paths/links, serializes JSON, performs placeholder substitution, and writes the output file. | +| `_build_info_content(metadata)` | Builds an HTML fragment (model name, generation timestamp, commit hash, component count, max depth) displayed in the viewer's info panel. | +| `_escape_html(text)` | Escapes HTML-sensitive characters to prevent malformed markup when embedding user/repo-derived strings (e.g., title). | +| `detect_repository_info(repo_path)` | Uses `GitPython` to read the repository name, normalize the remote URL (including `git@github.com:` SSH URLs), and compute the expected `https://.github.io//` Pages URL. | + +## Architecture + +```mermaid +flowchart TD + Docs["Documentation Output Directory"] --> ModuleTreeJson["module_tree.json"] + Docs --> MetadataJson["metadata.json"] + Template["viewer_template.html"] --> Generator["HTMLGenerator"] + ModuleTreeJson --> Generator + MetadataJson --> Generator + RepoPath["Repository Path"] --> DetectInfo["detect_repository_info()"] + DetectInfo --> Generator + Generator -->|"safe_write()"| IndexHtml["index.html"] + Generator -->|"on failure"| FSError["FileSystemError"] +``` + +### Dependencies + +- **Shared filesystem helpers** — `safe_read` / `safe_write` provide atomic, encoding-safe file I/O used to load the template/JSON files and write the final HTML output. See the [Utils](../utils/utils.md) module for other shared CLI utilities such as logging and progress tracking. +- **`FileSystemError`** — raised when the template file is missing or when reading/writing fails, allowing the CLI layer to surface a consistent error type. +- **`git` (GitPython)** — used only within `detect_repository_info` to introspect the local repository; failures are caught and silently ignored, degrading gracefully to a viewer without repository links. + +## Integration with the CLI Pipeline + +Html Generation is invoked as an optional, final stage by `CLIDocumentationGenerator` (from the [Generation](../generation/generation.md) module), which drives the overall CLI workflow: dependency analysis → module clustering → documentation generation → **HTML generation** → job finalization. + +```mermaid +sequenceDiagram + participant CLIGen as "CLIDocumentationGenerator" + participant HTMLGen as "HTMLGenerator" + participant FS as "safe_read / safe_write" + participant Git as "GitPython" + + CLIGen->>HTMLGen: HTMLGenerator() + CLIGen->>HTMLGen: detect_repository_info(repo_path) + HTMLGen->>Git: Repo(repo_path) + Git-->>HTMLGen: remote URL, name + HTMLGen-->>CLIGen: name, url, github_pages_url + CLIGen->>HTMLGen: generate(output_path, title, repository_url, docs_dir) + HTMLGen->>FS: safe_read(module_tree.json) + HTMLGen->>FS: safe_read(metadata.json) + HTMLGen->>FS: safe_read(viewer_template.html) + HTMLGen->>HTMLGen: _build_info_content(metadata) + HTMLGen->>HTMLGen: substitute placeholders + HTMLGen->>FS: safe_write(index.html) + HTMLGen-->>CLIGen: index.html written +``` + +This corresponds to the `_run_html_generation` step inside `CLIDocumentationGenerator.generate()`: it is only executed when the CLI was invoked with `generate_html=True`, after the backend has produced Markdown files, `module_tree.json`, and `metadata.json` in the output directory. On success, `"index.html"` is appended to `DocumentationJob.files_generated` (see the [Job Models](../job_models/job_models.md) module). + +## Generation Flow + +```mermaid +flowchart TD + Start["generate() called"] --> CheckDocsDir{{"docs_dir provided?"}} + CheckDocsDir -->|"yes"| AutoLoadTree["load_module_tree(docs_dir)"] + CheckDocsDir -->|"yes"| AutoLoadMeta["load_metadata(docs_dir)"] + CheckDocsDir -->|"no"| UseProvided["use provided module_tree/metadata"] + AutoLoadTree --> Defaults + AutoLoadMeta --> Defaults + UseProvided --> Defaults + Defaults["Apply defaults for module_tree/config"] --> LoadTemplate["Load viewer_template.html"] + LoadTemplate -->|"missing"| RaiseErr["raise FileSystemError"] + LoadTemplate -->|"found"| BuildInfo["_build_info_content(metadata)"] + BuildInfo --> BuildRepoLink["Build repository link HTML"] + BuildRepoLink --> ComputeBasePath["Compute docs_base_path"] + ComputeBasePath --> SerializeJson["Serialize config/module_tree/metadata to JSON"] + SerializeJson --> Replace["Replace template placeholders"] + Replace --> WriteOut["safe_write(output_path)"] + WriteOut --> Done["index.html generated"] +``` + +### Template Placeholders + +The generator performs a straightforward string substitution over the template file. The following placeholders are populated by `generate()`: + +| Placeholder | Source | +|---|---| +| `{{TITLE}}` | Escaped `title` argument (e.g., repository name) | +| `{{REPO_LINK}}` | HTML anchor to `repository_url`, empty if not provided | +| `{{SHOW_INFO}}` | `"block"` or `"none"` depending on whether info content was built | +| `{{INFO_CONTENT}}` | HTML fragment from `_build_info_content` (model, timestamp, commit, stats) | +| `{{CONFIG_JSON}}` | JSON-serialized `config` dictionary | +| `{{MODULE_TREE_JSON}}` | JSON-serialized module tree structure | +| `{{METADATA_JSON}}` | JSON-serialized metadata, or the literal `null` | +| `{{DOCS_BASE_PATH}}` | Relative path from the output file to the docs directory | + +## Error Handling + +- Missing `viewer_template.html` raises `FileSystemError`, propagated up to the CLI layer. +- Missing or malformed `module_tree.json` is either substituted with a minimal fallback structure (`load_module_tree`) or, on unexpected read/parse errors, raises `FileSystemError`. +- Missing or malformed `metadata.json` is treated as non-critical: `load_metadata` swallows exceptions and returns `None`, resulting in the info panel being hidden (`{{SHOW_INFO}} = "none"`). +- Git introspection failures in `detect_repository_info` are caught broadly, so the generator always returns a usable (if partially empty) info dictionary rather than failing the whole pipeline. + +## Relationship to Other Modules + +- **[Generation](../generation/generation.md)** — the `CLIDocumentationGenerator` orchestrates when and how `HTMLGenerator` is invoked as part of the overall documentation generation job. +- **[Job Models](../job_models/job_models.md)** — the generated `index.html` file path is recorded in the `DocumentationJob.files_generated` list. +- **[Utils](../utils/utils.md)** — shares low-level filesystem helpers (`safe_read`/`safe_write`) and error types used throughout the CLI. +- **[Cli Core](../../cli-core.md)** — the parent module that aggregates Html Generation alongside generation, configuration, job models, git integration, and shared utilities into the full CLI toolchain. diff --git a/docs/reference/architecture/cli-core/cli-core/job_models/job_models.md b/docs/reference/architecture/cli-core/cli-core/job_models/job_models.md new file mode 100644 index 00000000..fb94a70c --- /dev/null +++ b/docs/reference/architecture/cli-core/cli-core/job_models/job_models.md @@ -0,0 +1,244 @@ +# Job Models + +## Introduction + +The Job Models module defines the core data structures used to represent, track, and persist documentation generation jobs within the CodeWiki CLI. It is a foundational, dependency-free module that provides typed dataclasses and enums for job state, configuration, and statistics — enabling consistent serialization, deserialization, and status tracking across the CLI documentation pipeline. + +This module is a child of the [Cli Core](../../cli-core.md) module and is consumed by sibling modules such as [Generation](../generation/generation.md), [Configuration](../configuration/configuration.md), [Utils](../utils/utils.md), and other CLI orchestration components that need to create, update, or persist job records. + +## Purpose and Scope + +The Job Models module has a single, focused responsibility: **define the shape of a documentation job and its lifecycle**. It does not perform any I/O, orchestration, or business logic beyond simple state transitions and serialization helpers. This keeps the module lightweight, easily testable, and safe to import from anywhere in the CLI codebase without introducing circular dependencies. + +Key responsibilities: +- Define the `JobStatus` enum representing the lifecycle states of a job. +- Define `GenerationOptions` for user-configurable generation behavior (branching, GitHub Pages, caching, output paths). +- Define `LLMConfig` for capturing which language models and endpoint were used for a run. +- Define `JobStatistics` for capturing quantitative results of a run (files analyzed, tree depth, tokens consumed). +- Define `DocumentationJob`, the aggregate root that ties together identity, timing, status, configuration, and results for a single documentation generation run. +- Provide robust `to_dict()` / `to_json()` / `from_dict()` methods for safe persistence and rehydration, including defensive coercion of malformed or partial data. + +## Core Components + +### JobStatus + +`JobStatus` is a `str`-backed `Enum` representing the possible lifecycle states of a documentation job: + +| Value | Meaning | +|---|---| +| `pending` | Job has been created but not yet started | +| `running` | Job is actively executing | +| `completed` | Job finished successfully | +| `failed` | Job terminated with an error | + +Because it inherits from both `str` and `Enum`, instances serialize naturally to plain strings in JSON output while still supporting type-safe comparisons in code (e.g., `job.status == JobStatus.RUNNING`). + +### GenerationOptions + +`GenerationOptions` is a dataclass capturing user-facing flags that control how a documentation run behaves: + +- `create_branch` (`bool`): Whether to create a dedicated git branch for the generated docs. Consumed by the [Git Integration](../git_integration/git_integration.md) module. +- `github_pages` (`bool`): Whether to prepare output for GitHub Pages publishing. +- `no_cache` (`bool`): Whether to bypass any caching layer during generation. +- `custom_output` (`Optional[str]`): An optional override for the output directory. + +### LLMConfig + +`LLMConfig` records which language models and endpoint were used to produce a job's output: + +- `main_model` (`str`): The primary model used for documentation generation. +- `cluster_model` (`str`): The model used for clustering/module grouping decisions. +- `base_url` (`str`): The API base URL for the LLM provider. + +This is a plain, required-field dataclass (no defaults), reflecting that when an `LLMConfig` is attached to a job, all three fields are expected to be known. + +### JobStatistics + +`JobStatistics` aggregates quantitative metrics collected during a documentation run: + +- `total_files_analyzed` (`int`): Count of source files processed. +- `leaf_nodes` (`int`): Count of leaf-level components/modules identified. +- `max_depth` (`int`): Maximum depth of the module hierarchy produced. +- `total_tokens_used` (`int`): Total LLM tokens consumed for the run. + +All fields default to `0`, so a freshly created job has a valid, zeroed-out statistics object even before execution begins. + +### DocumentationJob + +`DocumentationJob` is the central aggregate of this module. It represents a single documentation generation run end-to-end, combining identity, repository context, git metadata, timing, status, and nested configuration/statistics objects. + +**Fields:** + +| Field | Type | Description | +|---|---|---| +| `job_id` | `str` | UUID4 identifier, auto-generated if not supplied | +| `repository_path` | `str` | Absolute filesystem path to the repository being documented | +| `repository_name` | `str` | Human-readable repository name | +| `output_directory` | `str` | Destination directory for generated docs | +| `commit_hash` | `str` | Git commit SHA the job was run against | +| `branch_name` | `Optional[str]` | Git branch name, if applicable | +| `timestamp_start` | `str` | ISO-format start timestamp, auto-populated | +| `timestamp_end` | `Optional[str]` | ISO-format end timestamp, set on completion/failure | +| `status` | `JobStatus` | Current lifecycle status | +| `error_message` | `Optional[str]` | Populated when the job fails | +| `files_generated` | `List[str]` | Paths of documentation files produced | +| `module_count` | `int` | Number of modules documented | +| `generation_options` | `GenerationOptions` | Options selected for this run | +| `llm_config` | `Optional[LLMConfig]` | LLM configuration used, if known | +| `statistics` | `JobStatistics` | Quantitative results of the run | + +**Lifecycle methods:** + +- `start()` — Transitions status to `RUNNING` and refreshes `timestamp_start`. +- `complete()` — Transitions status to `COMPLETED` and sets `timestamp_end`. +- `fail(error_message)` — Transitions status to `FAILED`, records the error, and sets `timestamp_end`. + +**Serialization methods:** + +- `to_dict()` — Produces a JSON-serializable `dict` representation, flattening nested dataclasses (`GenerationOptions`, `LLMConfig`, `JobStatistics`) into plain dicts and converting `JobStatus` to its string value. +- `to_json()` — Convenience wrapper around `to_dict()` that returns a pretty-printed JSON string (2-space indent). +- `from_dict(data)` — Classmethod that reconstructs a `DocumentationJob` from a raw dictionary (e.g., loaded from a JSON file), using defensive coercion helpers to tolerate missing, malformed, or partial fields. + +### Defensive Coercion Helpers + +The module defines several private module-level helper functions used exclusively by `DocumentationJob.from_dict()` to safely rebuild nested objects from untrusted or partial data: + +- `_coerce_job_status(value, default)` — Converts a raw value to a valid `JobStatus`, falling back to `JobStatus.PENDING` (or a supplied default) if the value is `None` or not a recognized status string. +- `_coerce_int(value, default)` — Safely converts a value to `int`, returning a default (`0`) on `None`, `TypeError`, or `ValueError`. +- `_coerce_generation_options(value)` — Rebuilds a `GenerationOptions` from a dict, defaulting missing keys. +- `_coerce_llm_config(value)` — Rebuilds an `LLMConfig` from a dict, or returns `None` if no data is present. +- `_coerce_statistics(value)` — Rebuilds a `JobStatistics` from a dict, defaulting missing/invalid numeric fields to `0`. + +This coercion layer makes `DocumentationJob.from_dict()` resilient to schema drift (e.g., loading job files written by an older version of the CLI) without raising exceptions during deserialization. + +## Architecture + +### Component Structure + +```mermaid +classDiagram + class JobStatus { + <> + PENDING + RUNNING + COMPLETED + FAILED + } + class GenerationOptions { + +bool create_branch + +bool github_pages + +bool no_cache + +str custom_output + } + class LLMConfig { + +str main_model + +str cluster_model + +str base_url + } + class JobStatistics { + +int total_files_analyzed + +int leaf_nodes + +int max_depth + +int total_tokens_used + } + class DocumentationJob { + +str job_id + +str repository_path + +str repository_name + +str output_directory + +str commit_hash + +str branch_name + +str timestamp_start + +str timestamp_end + +JobStatus status + +str error_message + +List files_generated + +int module_count + +GenerationOptions generation_options + +LLMConfig llm_config + +JobStatistics statistics + +start() + +complete() + +fail(error_message) + +to_dict() dict + +to_json() str + +from_dict(data) DocumentationJob + } + DocumentationJob --> JobStatus : "status" + DocumentationJob --> GenerationOptions : "generation_options" + DocumentationJob --> LLMConfig : "llm_config (optional)" + DocumentationJob --> JobStatistics : "statistics" +``` + +### Job Lifecycle State Machine + +```mermaid +stateDiagram-v2 + [*] --> PENDING : "DocumentationJob() created" + PENDING --> RUNNING : "start()" + RUNNING --> COMPLETED : "complete()" + RUNNING --> FAILED : "fail(error_message)" + COMPLETED --> [*] + FAILED --> [*] +``` + +### Serialization / Deserialization Flow + +```mermaid +sequenceDiagram + participant Caller as "CLI Component" + participant Job as "DocumentationJob" + participant Coerce as "Coercion Helpers" + + Caller->>Job: "DocumentationJob(...)" + Job-->>Caller: "job instance (status=PENDING)" + Caller->>Job: "job.start()" + Caller->>Job: "job.to_dict() / job.to_json()" + Job-->>Caller: "dict / JSON string" + Note over Caller: "Persisted to disk or transmitted" + + Caller->>Job: "DocumentationJob.from_dict(raw_data)" + Job->>Coerce: "_coerce_job_status(raw_data.status)" + Job->>Coerce: "_coerce_int(raw_data.module_count)" + Job->>Coerce: "_coerce_generation_options(raw_data.generation_options)" + Job->>Coerce: "_coerce_llm_config(raw_data.llm_config)" + Job->>Coerce: "_coerce_statistics(raw_data.statistics)" + Coerce-->>Job: "validated nested objects" + Job-->>Caller: "reconstructed DocumentationJob" +``` + +## Integration with Other Modules + +The Job Models module is intentionally dependency-free (it only imports from the Python standard library), which allows it to be safely imported by any layer of the CLI without risk of circular imports: + +- **[Cli Core](../../cli-core.md)** (parent module): Exposes `DocumentationJob`, `GenerationOptions`, `JobStatistics`, `JobStatus`, and `LLMConfig` as part of its public component surface, making them available to all CLI subsystems. +- **[Generation](../generation/generation.md)**: The `CLIDocumentationGenerator` creates and drives `DocumentationJob` instances through their lifecycle (`start()` → `complete()`/`fail()`), populating `files_generated`, `module_count`, and `statistics` as generation proceeds. +- **[Configuration](../configuration/configuration.md)**: `ConfigManager` and the `Configuration`/`AgentInstructions` models supply values (such as model names and base URLs) that are captured into a job's `LLMConfig`. +- **[Git Integration](../git_integration/git_integration.md)**: `GitManager` supplies `commit_hash` and `branch_name` values and consults `GenerationOptions.create_branch` to decide whether to create a dedicated branch for generated documentation. +- **[Utils](../utils/utils.md)**: `CLILogger` and `ProgressTracker`/`ModuleProgressBar` report on job progress using the status and statistics captured on `DocumentationJob` instances. +- **[Html Generation](../html_generation/html_generation.md)**: `HTMLGenerator` may reference `files_generated` and `output_directory` from a completed job when producing rendered output. + +```mermaid +flowchart TD + JobModels["Job Models"] + CliCore["Cli Core (parent)"] + Generation["Generation"] + Configuration["Configuration"] + GitIntegration["Git Integration"] + HtmlGeneration["Html Generation"] + Utils["Utils"] + + CliCore --> JobModels + Generation -->|"creates & drives"| JobModels + Configuration -->|"populates LLMConfig"| JobModels + GitIntegration -->|"populates commit_hash / branch_name"| JobModels + HtmlGeneration -->|"reads files_generated"| JobModels + Utils -->|"reports on status/statistics"| JobModels +``` + +## Design Notes + +- **Immutability of shape, mutability of state**: While the dataclasses themselves are mutable (no `frozen=True`), the module encourages controlled mutation through explicit lifecycle methods (`start()`, `complete()`, `fail()`) rather than direct field assignment, keeping status transitions predictable and auditable. +- **Safe round-tripping**: The combination of `to_dict()`/`to_json()` and `from_dict()` with defensive coercion helpers ensures that job records can be persisted to disk (e.g., as job history or resumable state) and reloaded reliably even if the on-disk schema is incomplete or slightly out of date. +- **String-backed enum**: `JobStatus` extending `str` means job status values serialize directly as readable strings in JSON without requiring custom encoders, simplifying integration with logging, CLI output, and external tooling. +- **Zero external dependencies**: The module relies only on `dataclasses`, `datetime`, `typing`, `enum`, `uuid`, and `json` from the standard library, making it trivially testable and reusable across CLI contexts. diff --git a/docs/reference/architecture/cli-core/cli-core/utils/utils.md b/docs/reference/architecture/cli-core/cli-core/utils/utils.md new file mode 100644 index 00000000..339a3fd3 --- /dev/null +++ b/docs/reference/architecture/cli-core/cli-core/utils/utils.md @@ -0,0 +1,168 @@ +# Utils + +The Utils module provides the foundational output and progress-reporting primitives used throughout the [Cli Core](../../cli-core.md) subsystem. It contains no business logic of its own; instead it supplies two small, dependency-light building blocks — a colored console logger and a multi-stage progress tracker — that every other CLI component (job orchestration, generation, git integration, HTML generation) relies on to communicate status to the end user. + +Because these utilities sit at the bottom of the CLI dependency graph, they are intentionally free of imports from other CLI modules. This keeps them reusable, easy to test, and safe to import from anywhere in the CLI without risking circular dependencies. + +## Purpose and Scope + +The module addresses two related but distinct concerns: + +1. **Structured, leveled console output** — via `CLILogger`, which standardizes how informational, success, warning, error, debug, and step messages are rendered to the terminal (including color and timestamps), and quiets noisy third-party HTTP/SDK loggers. +2. **Progress and ETA reporting** — via `ProgressTracker` (stage-weighted progress across the overall documentation pipeline) and `ModuleProgressBar` (per-module progress during the module-by-module generation phase). + +These two concerns are complementary: `CLILogger` handles ad-hoc textual messages, while the progress classes handle numeric/visual progress state and time estimation for long-running operations. + +## Component Overview + +```mermaid +classDiagram + class CLILogger { + +bool verbose + +datetime start_time + +debug(message) void + +info(message) void + +success(message) void + +warning(message) void + +error(message) void + +step(message, step, total) void + +elapsed_time() str + } + + class ProgressTracker { + +int total_stages + +int current_stage + +float stage_progress + +float start_time + +bool verbose + +STAGE_WEIGHTS dict + +STAGE_NAMES dict + +start_stage(stage, description) void + +update_stage(progress, message) void + +complete_stage(message) void + +get_overall_progress() float + +get_eta() str + } + + class ModuleProgressBar { + +int total_modules + +int current_module + +bool verbose + +bar + +update(module_name, cached) void + +finish() void + } +``` + +### CLILogger + +`CLILogger` (`codewiki/cli/utils/logging.py`) is the standard mechanism for writing user-facing output during a CLI run. It wraps [Click](https://click.palletsprojects.com/)'s `echo`/`secho` helpers and adds: + +- **Leveled methods**: `debug`, `info`, `success`, `warning`, `error`, and `step`, each with a distinct color and symbol (✓ for success, ⚠️ for warning, ✗ for error) so terminal output is easy to scan. +- **Verbose-only debug output**: `debug()` messages are only rendered when `verbose=True`, and are timestamped for correlation with other logs. +- **Step announcements**: `step()` renders either a `[current/total]` prefix (when both `step` and `total` are supplied) or a generic arrow (`→`) prefix, useful for announcing pipeline phases. +- **Elapsed time tracking**: `elapsed_time()` reports time since the logger was constructed, formatted as `Xm Ys` or `Ys`. + +The module also exposes two module-level helpers: + +- `quiet_third_party_loggers(level=logging.WARNING)` — caps the log level of noisy third-party libraries (`httpx`, `openai`, `openai._base_client`, `anthropic`) so that a documentation run does not produce thousands of INFO-level HTTP request lines in CI output. This is applied explicitly rather than as an import-time side effect, so its behavior is visible at the call site. +- `create_logger(verbose=False)` — the standard factory used by CLI entry points to construct a properly configured `CLILogger`, automatically invoking `quiet_third_party_loggers()` first. + +### ProgressTracker + +`ProgressTracker` (`codewiki/cli/utils/progress.py`) models the overall documentation generation pipeline as five weighted stages: + +| Stage | Name | Weight | +|-------|------|--------| +| 1 | Dependency Analysis | 40% | +| 2 | Module Clustering | 20% | +| 3 | Documentation Generation | 30% | +| 4 | HTML Generation (optional) | 5% | +| 5 | Finalization | 5% | + +These weights (`STAGE_WEIGHTS`) reflect the relative time each stage is expected to consume, and are used to compute an aggregate `get_overall_progress()` value across the whole run, independent of how much intra-stage progress has been made. + +Key behaviors: + +- `start_stage(stage, description=None)` resets stage progress to `0.0`, records the stage start time, and prints a banner (verbose mode includes elapsed time; non-verbose mode shows a compact `[stage/total]` header). +- `update_stage(progress, message=None)` clamps `progress` to `[0.0, 1.0]` and, in verbose mode, prints an indented status message. +- `complete_stage(message=None)` sets stage progress to `1.0` and, in verbose mode, prints the stage's wall-clock duration plus any completion message. +- `get_overall_progress()` sums the weights of fully completed stages plus the weighted fraction of the current stage's progress. +- `get_eta()` extrapolates total run time from elapsed time and overall progress, returning a human-readable estimate (e.g., `"2m 15s"`, `"1h 5m"`, or `"< 1 min"`), or `None` if no progress has been made yet (avoiding a divide-by-zero). + +This stage model is the shared contract that pipeline-driving code (in the [Job Models](../job_models/job_models.md) and [Generation](../generation/generation.md) modules) uses to report high-level progress consistently. + +### ModuleProgressBar + +`ModuleProgressBar` (`codewiki/cli/utils/progress.py`) is a narrower, complementary tool used specifically during the "Documentation Generation" stage, where progress is naturally expressed as "N of M modules processed" rather than a continuous percentage. + +- In **non-verbose** mode, it wraps `click.progressbar` to render a live terminal progress bar with ETA and percentage, entering the context manager on construction (`__enter__`) and exiting it on `finish()` (`__exit__`). +- In **verbose** mode, no bar is drawn; instead, `update()` prints one line per module, showing whether the module was `✓ (cached)` or `⟳ (generating)`. +- `update(module_name, cached=False)` increments the internal module counter and reports progress using whichever mode is active. +- `finish()` safely closes the underlying `click.progressbar` context if one was opened. + +Because `ModuleProgressBar` owns a `click.progressbar` context manager internally, callers should ensure `finish()` is invoked (e.g., in a `finally` block) even if module generation raises an exception, to avoid leaving the terminal progress bar in an inconsistent state. + +## Interaction with the Broader CLI Pipeline + +The Utils module is consumed — not the consumer. Higher-level orchestration code (such as the documentation generation flow) drives both `ProgressTracker` and `ModuleProgressBar` in tandem: `ProgressTracker` reports macro-level stage progress across the whole run, while `ModuleProgressBar` provides fine-grained visibility during the module generation stage specifically. `CLILogger` is used throughout for all textual status, warnings, and errors. + +```mermaid +sequenceDiagram + participant Pipeline as "CLI Pipeline" + participant Logger as "CLILogger" + participant Tracker as "ProgressTracker" + participant ModBar as "ModuleProgressBar" + + Pipeline->>Logger: create_logger(verbose) + Pipeline->>Tracker: start_stage(1, "Dependency Analysis") + Pipeline->>Logger: info("Analyzing repository...") + Tracker-->>Pipeline: update_stage(progress) + Pipeline->>Tracker: complete_stage() + + Pipeline->>Tracker: start_stage(3, "Documentation Generation") + Pipeline->>ModBar: new ModuleProgressBar(total_modules) + loop "for each module" + Pipeline->>ModBar: update(module_name, cached) + Pipeline->>Logger: debug("module details") + end + Pipeline->>ModBar: finish() + Pipeline->>Tracker: complete_stage("Generation finished") + + Pipeline->>Logger: success("Documentation generated") +``` + +## Typical Usage Flow + +```mermaid +flowchart TD + A["create_logger(verbose)"] --> B["ProgressTracker(total_stages=5)"] + B --> C["tracker.start_stage(1, 'Dependency Analysis')"] + C --> D["perform analysis; call tracker.update_stage()"] + D --> E["tracker.complete_stage()"] + E --> F["tracker.start_stage(3, 'Documentation Generation')"] + F --> G["ModuleProgressBar(total_modules)"] + G --> H["for each module: bar.update(name, cached)"] + H --> I["bar.finish()"] + I --> J["tracker.complete_stage()"] + J --> K["logger.success('Done')"] +``` + +## Design Notes + +- **No cross-module coupling**: Neither `CLILogger` nor the progress classes depend on other CLI modules such as [Configuration](../configuration/configuration.md), [Job Models](../job_models/job_models.md), or [Generation](../generation/generation.md). This keeps them trivially reusable and testable in isolation. +- **Verbose vs. non-verbose modes are first-class**: Every class exposes a `verbose` flag that changes rendering strategy (timestamped multi-line logs vs. compact single-line/progress-bar output), rather than relying on global logging configuration. +- **Explicit side-effect application**: `quiet_third_party_loggers()` is called from `create_logger()` rather than at module import time, so that importing the `utils` package never silently mutates global logging state — the effect only occurs when a caller explicitly requests a logger. +- **Time estimation is defensive**: `get_eta()` guards against division by zero when no progress has been recorded, returning `None` instead of raising or producing a nonsensical estimate. + +## Relationship to Parent and Sibling Modules + +The Utils module is a child of [Cli Core](../../cli-core.md), alongside: + +- [Generation](../generation/generation.md) — the documentation generation adapter that most heavily relies on `ProgressTracker` and `ModuleProgressBar` to report pipeline progress. +- [Configuration](../configuration/configuration.md) — CLI configuration models and the config manager. +- [Job Models](../job_models/job_models.md) — data models representing documentation jobs, statuses, and statistics that the progress-tracking stages correspond to. +- [Git Integration](../git_integration/git_integration.md) — git operations invoked during pipeline execution, whose status is typically reported via `CLILogger`. +- [Html Generation](../html_generation/html_generation.md) — the optional HTML rendering stage represented as Stage 4 in `ProgressTracker.STAGE_NAMES`. + +These sibling modules depend on Utils for consistent status reporting; Utils itself has no dependency on them. diff --git a/docs/reference/architecture/cli-core/cli_core/utils/utils.md b/docs/reference/architecture/cli-core/cli_core/utils/utils.md new file mode 100644 index 00000000..ac1b4829 --- /dev/null +++ b/docs/reference/architecture/cli-core/cli_core/utils/utils.md @@ -0,0 +1,168 @@ +# Utils + +The Utils module provides the foundational output and progress-reporting primitives used throughout the [CLI Core](../cli_core.md) subsystem. It contains no business logic of its own; instead it supplies two small, dependency-light building blocks — a colored console logger and a multi-stage progress tracker — that every other CLI component (job orchestration, generation, git integration, HTML generation) relies on to communicate status to the end user. + +Because these utilities sit at the bottom of the CLI dependency graph, they are intentionally free of imports from other CLI modules. This keeps them reusable, easy to test, and safe to import from anywhere in the CLI without risking circular dependencies. + +## Purpose and Scope + +The module addresses two related but distinct concerns: + +1. **Structured, leveled console output** — via `CLILogger`, which standardizes how informational, success, warning, error, debug, and step messages are rendered to the terminal (including color and timestamps), and quiets noisy third-party HTTP/SDK loggers. +2. **Progress and ETA reporting** — via `ProgressTracker` (stage-weighted progress across the overall documentation pipeline) and `ModuleProgressBar` (per-module progress during the module-by-module generation phase). + +These two concerns are complementary: `CLILogger` handles ad-hoc textual messages, while the progress classes handle numeric/visual progress state and time estimation for long-running operations. + +## Component Overview + +```mermaid +classDiagram + class CLILogger { + +bool verbose + +datetime start_time + +debug(message) void + +info(message) void + +success(message) void + +warning(message) void + +error(message) void + +step(message, step, total) void + +elapsed_time() str + } + + class ProgressTracker { + +int total_stages + +int current_stage + +float stage_progress + +float start_time + +bool verbose + +STAGE_WEIGHTS dict + +STAGE_NAMES dict + +start_stage(stage, description) void + +update_stage(progress, message) void + +complete_stage(message) void + +get_overall_progress() float + +get_eta() str + } + + class ModuleProgressBar { + +int total_modules + +int current_module + +bool verbose + +bar + +update(module_name, cached) void + +finish() void + } +``` + +### CLILogger + +`CLILogger` (`codewiki/cli/utils/logging.py`) is the standard mechanism for writing user-facing output during a CLI run. It wraps [Click](https://click.palletsprojects.com/)'s `echo`/`secho` helpers and adds: + +- **Leveled methods**: `debug`, `info`, `success`, `warning`, `error`, and `step`, each with a distinct color and symbol (✓ for success, ⚠️ for warning, ✗ for error) so terminal output is easy to scan. +- **Verbose-only debug output**: `debug()` messages are only rendered when `verbose=True`, and are timestamped for correlation with other logs. +- **Step announcements**: `step()` renders either a `[current/total]` prefix (when both `step` and `total` are supplied) or a generic arrow (`→`) prefix, useful for announcing pipeline phases. +- **Elapsed time tracking**: `elapsed_time()` reports time since the logger was constructed, formatted as `Xm Ys` or `Ys`. + +The module also exposes two module-level helpers: + +- `quiet_third_party_loggers(level=logging.WARNING)` — caps the log level of noisy third-party libraries (`httpx`, `openai`, `openai._base_client`, `anthropic`) so that a documentation run does not produce thousands of INFO-level HTTP request lines in CI output. This is applied explicitly rather than as an import-time side effect, so its behavior is visible at the call site. +- `create_logger(verbose=False)` — the standard factory used by CLI entry points to construct a properly configured `CLILogger`, automatically invoking `quiet_third_party_loggers()` first. + +### ProgressTracker + +`ProgressTracker` (`codewiki/cli/utils/progress.py`) models the overall documentation generation pipeline as five weighted stages: + +| Stage | Name | Weight | +|-------|------|--------| +| 1 | Dependency Analysis | 40% | +| 2 | Module Clustering | 20% | +| 3 | Documentation Generation | 30% | +| 4 | HTML Generation (optional) | 5% | +| 5 | Finalization | 5% | + +These weights (`STAGE_WEIGHTS`) reflect the relative time each stage is expected to consume, and are used to compute an aggregate `get_overall_progress()` value across the whole run, independent of how much intra-stage progress has been made. + +Key behaviors: + +- `start_stage(stage, description=None)` resets stage progress to `0.0`, records the stage start time, and prints a banner (verbose mode includes elapsed time; non-verbose mode shows a compact `[stage/total]` header). +- `update_stage(progress, message=None)` clamps `progress` to `[0.0, 1.0]` and, in verbose mode, prints an indented status message. +- `complete_stage(message=None)` sets stage progress to `1.0` and, in verbose mode, prints the stage's wall-clock duration plus any completion message. +- `get_overall_progress()` sums the weights of fully completed stages plus the weighted fraction of the current stage's progress. +- `get_eta()` extrapolates total run time from elapsed time and overall progress, returning a human-readable estimate (e.g., `"2m 15s"`, `"1h 5m"`, or `"< 1 min"`), or `None` if no progress has been made yet (avoiding a divide-by-zero). + +This stage model is the shared contract that pipeline-driving code (in the [Job Models](../job_models/job_models.md) and generation layers) uses to report high-level progress consistently. + +### ModuleProgressBar + +`ModuleProgressBar` (`codewiki/cli/utils/progress.py`) is a narrower, complementary tool used specifically during the "Documentation Generation" stage, where progress is naturally expressed as "N of M modules processed" rather than a continuous percentage. + +- In **non-verbose** mode, it wraps `click.progressbar` to render a live terminal progress bar with ETA and percentage, entering the context manager on construction (`__enter__`) and exiting it on `finish()` (`__exit__`). +- In **verbose** mode, no bar is drawn; instead, `update()` prints one line per module, showing whether the module was `✓ (cached)` or `⟳ (generating)`. +- `update(module_name, cached=False)` increments the internal module counter and reports progress using whichever mode is active. +- `finish()` safely closes the underlying `click.progressbar` context if one was opened. + +Because `ModuleProgressBar` owns a `click.progressbar` context manager internally, callers should ensure `finish()` is invoked (e.g., in a `finally` block) even if module generation raises an exception, to avoid leaving the terminal progress bar in an inconsistent state. + +## Interaction with the Broader CLI Pipeline + +The Utils module is consumed — not the consumer. Higher-level orchestration code (such as the documentation generation flow) drives both `ProgressTracker` and `ModuleProgressBar` in tandem: `ProgressTracker` reports macro-level stage progress across the whole run, while `ModuleProgressBar` provides fine-grained visibility during the module generation stage specifically. `CLILogger` is used throughout for all textual status, warnings, and errors. + +```mermaid +sequenceDiagram + participant Pipeline as "CLI Pipeline" + participant Logger as "CLILogger" + participant Tracker as "ProgressTracker" + participant ModBar as "ModuleProgressBar" + + Pipeline->>Logger: create_logger(verbose) + Pipeline->>Tracker: start_stage(1, "Dependency Analysis") + Pipeline->>Logger: info("Analyzing repository...") + Tracker-->>Pipeline: update_stage(progress) + Pipeline->>Tracker: complete_stage() + + Pipeline->>Tracker: start_stage(3, "Documentation Generation") + Pipeline->>ModBar: new ModuleProgressBar(total_modules) + loop "for each module" + Pipeline->>ModBar: update(module_name, cached) + Pipeline->>Logger: debug("module details") + end + Pipeline->>ModBar: finish() + Pipeline->>Tracker: complete_stage("Generation finished") + + Pipeline->>Logger: success("Documentation generated") +``` + +## Typical Usage Flow + +```mermaid +flowchart TD + A["create_logger(verbose)"] --> B["ProgressTracker(total_stages=5)"] + B --> C["tracker.start_stage(1, 'Dependency Analysis')"] + C --> D["perform analysis; call tracker.update_stage()"] + D --> E["tracker.complete_stage()"] + E --> F["tracker.start_stage(3, 'Documentation Generation')"] + F --> G["ModuleProgressBar(total_modules)"] + G --> H["for each module: bar.update(name, cached)"] + H --> I["bar.finish()"] + I --> J["tracker.complete_stage()"] + J --> K["logger.success('Done')"] +``` + +## Design Notes + +- **No cross-module coupling**: Neither `CLILogger` nor the progress classes depend on other CLI modules such as [Configuration](../configuration/configuration.md), [Job Models](../job_models/job_models.md), or [Generation](../generation/generation.md). This keeps them trivially reusable and testable in isolation. +- **Verbose vs. non-verbose modes are first-class**: Every class exposes a `verbose` flag that changes rendering strategy (timestamped multi-line logs vs. compact single-line/progress-bar output), rather than relying on global logging configuration. +- **Explicit side-effect application**: `quiet_third_party_loggers()` is called from `create_logger()` rather than at module import time, so that importing the `utils` package never silently mutates global logging state — the effect only occurs when a caller explicitly requests a logger. +- **Time estimation is defensive**: `get_eta()` guards against division by zero when no progress has been recorded, returning `None` instead of raising or producing a nonsensical estimate. + +## Relationship to Parent and Sibling Modules + +The Utils module is a child of [CLI Core](../cli_core.md), alongside: + +- [Generation](../generation/generation.md) — the documentation generation adapter that most heavily relies on `ProgressTracker` and `ModuleProgressBar` to report pipeline progress. +- [Configuration](../configuration/configuration.md) — CLI configuration models and the config manager. +- [Job Models](../job_models/job_models.md) — data models representing documentation jobs, statuses, and statistics that the progress-tracking stages correspond to. +- [Git Integration](../git_integration/git_integration.md) — git operations invoked during pipeline execution, whose status is typically reported via `CLILogger`. +- [HTML Generation](../html_generation/html_generation.md) — the optional HTML rendering stage represented as Stage 4 in `ProgressTracker.STAGE_NAMES`. + +These sibling modules depend on Utils for consistent status reporting; Utils itself has no dependency on them. diff --git a/docs/reference/architecture/cli-core/configuration.md b/docs/reference/architecture/cli-core/configuration.md new file mode 100644 index 00000000..dda61e21 --- /dev/null +++ b/docs/reference/architecture/cli-core/configuration.md @@ -0,0 +1,249 @@ +# Configuration + +The Configuration module provides the persistent settings layer for the CodeWiki CLI. It defines the data models that describe user preferences (LLM providers, model names, token limits, temperature settings, and documentation-generation instructions) and the manager responsible for reading and writing those settings safely to disk and to the operating system's secure credential store. + +This module is a child of [Cli Core](../cli-core.md) and works alongside sibling modules such as the job/runtime models, generation pipeline, and utility helpers to support the CLI's end-to-end workflow. + +## Purpose and Scope + +The Configuration module answers three core questions for the CLI: + +1. **What settings does the user have configured?** — captured by the `Configuration` dataclass. +2. **How should the documentation agent behave for a given run?** — captured by `AgentInstructions`. +3. **How are these settings persisted, validated, and loaded securely?** — handled by `ConfigManager`. + +Sensitive values (API keys) are never written to plaintext configuration files. Instead, they are stored using the system keyring (macOS Keychain, Windows Credential Manager, or Linux Secret Service) and only non-sensitive settings are persisted to `~/.codewiki/config.json`. + +## Core Components + +| Component | Responsibility | +|---|---| +| `Configuration` | Dataclass representing all persistent CLI settings: model names, base URLs, API versions, token/temperature limits, and clustering parameters. | +| `AgentInstructions` | Dataclass representing optional, user-customizable instructions for the documentation agent (file filters, focus modules, doc type, free-form instructions). | +| `ConfigManager` | Orchestrates loading/saving `Configuration` to `~/.codewiki/config.json` and securely storing/retrieving API keys via keyring. | + +## Architecture Overview + +```mermaid +flowchart TD + subgraph ConfigModule["Configuration Module"] + CM["ConfigManager"] + Cfg["Configuration"] + AI["AgentInstructions"] + end + + FS["~/.codewiki/config.json"] + KR["System Keyring"] + Backend["Backend Config"] + + CM -->|"load() / save()"| FS + CM -->|"get/set API keys"| KR + CM -->|"holds"| Cfg + Cfg -->|"has one"| AI + Cfg -->|"to_backend_config()"| Backend +``` + +`Configuration` is a plain, serializable dataclass with no direct dependency on keyring or the filesystem — those concerns are owned exclusively by `ConfigManager`. This separation keeps the data model easy to test and reuse, while `ConfigManager` acts as the single access point for persistence. + +## Data Model: Configuration + +`Configuration` captures all settings needed to drive a documentation-generation run, organized into three provider "roles": + +- **cluster** — model used for module clustering/decomposition +- **main** — primary model used for documentation generation +- **fallback** — fallback model used when the main model fails or is rate-limited + +For each role, the model tracks: +- Model name (`*_model`) +- Base URL (`*_base_url`, optional — for OpenAI-compatible or self-hosted endpoints) +- API version (`*_api_version`, optional) +- Max tokens (`*_max_tokens`) +- Temperature (`*_temperature`) and whether temperature is supported (`*_temperature_supported`) +- Max-token parameter field name (`*_max_token_field`, e.g., `"max_tokens"` vs. provider-specific names) + +In addition to per-provider settings, `Configuration` holds shared clustering parameters (`max_token_per_module`, `max_token_per_leaf_module`, `max_depth`) and a `default_output` directory. API keys (`cluster_api_key`, `main_api_key`, `fallback_api_key`) exist on the dataclass only as **runtime-only fields** — they are never serialized by `to_dict()` and are populated separately from the keyring by `ConfigManager`. + +```mermaid +classDiagram + class Configuration { + +str main_model + +str cluster_model + +str fallback_model + +str default_output + +str cluster_base_url + +str main_base_url + +str fallback_base_url + +int cluster_max_tokens + +int main_max_tokens + +int fallback_max_tokens + +float cluster_temperature + +float main_temperature + +float fallback_temperature + +bool cluster_temperature_supported + +bool main_temperature_supported + +bool fallback_temperature_supported + +int max_token_per_module + +int max_token_per_leaf_module + +int max_depth + +AgentInstructions agent_instructions + +validate() + +to_dict() dict + +from_dict(data) Configuration + +is_complete() bool + +to_backend_config(...) Config + } + + class AgentInstructions { + +List~str~ include_patterns + +List~str~ exclude_patterns + +List~str~ focus_modules + +str doc_type + +str custom_instructions + +to_dict() dict + +from_dict(data) AgentInstructions + +is_empty() bool + +get_prompt_addition() str + } + + Configuration "1" *-- "1" AgentInstructions : agent_instructions +``` + +### Validation + +`Configuration.validate()` performs field-level checks: +- Base URLs (when set) are validated via `validate_url`. +- Model names for all three roles are validated via `validate_model_name`. + +`from_dict()` performs defensive type coercion when loading from JSON (which may contain string-typed numbers/booleans due to manual editing), converting and range-checking integers (e.g., token limits), floats (temperature, bounded `0.0`–`2.0`), and booleans, raising `ValueError` on invalid input. + +### Serialization Rules + +- `to_dict()` only emits optional fields (`*_base_url`, `*_api_version`, `agent_instructions`) when they are set/non-empty, keeping the persisted JSON minimal. +- API keys are **excluded** from `to_dict()` entirely — they are never written to `config.json`. + +## Data Model: AgentInstructions + +`AgentInstructions` lets users customize how the documentation agent analyzes a repository and generates content: + +- `include_patterns` / `exclude_patterns` — glob-style file filters (e.g., `["*.cs"]`, `["*Tests*"]`) +- `focus_modules` — modules that should receive more detailed documentation +- `doc_type` — a preset documentation style (`api`, `architecture`, `user-guide`, `developer`) or a free-form type +- `custom_instructions` — arbitrary additional guidance passed to the LLM + +`get_prompt_addition()` translates these fields into a natural-language instruction block that is merged into the agent's prompt at generation time. `is_empty()` allows callers to distinguish "no customization" from "customization with all-default values." + +Instructions can be set persistently (stored in `config.json` as part of `Configuration`) or supplied at runtime for a single job; `Configuration.to_backend_config()` merges the two, with runtime instructions taking precedence field-by-field. + +## ConfigManager: Persistence and Secure Storage + +`ConfigManager` is the sole component responsible for reading and writing configuration state. It manages two independent storage backends: + +1. **JSON file** (`~/.codewiki/config.json`) — non-sensitive settings, versioned with a `CONFIG_VERSION` marker for future migrations. +2. **System keyring** (via the `keyring` library) — the three provider API keys (`cluster_api_key`, `main_api_key`, `fallback_api_key`), stored under a shared `codewiki` service name with distinct account identifiers. + +```mermaid +classDiagram + class ConfigManager { + -Optional~str~ _cluster_api_key + -Optional~str~ _main_api_key + -Optional~str~ _fallback_api_key + -Optional~Configuration~ _config + -bool _keyring_available + +load() bool + +save(...) void + +get_cluster_api_key() Optional~str~ + +get_main_api_key() Optional~str~ + +get_fallback_api_key() Optional~str~ + +get_config() Optional~Configuration~ + +is_configured() bool + +delete_api_keys() void + +clear() void + +keyring_available bool + +config_file_path Path + } + ConfigManager --> Configuration : loads/saves +``` + +### Load Flow + +```mermaid +sequenceDiagram + participant Caller + participant CM as ConfigManager + participant FS as "config.json" + participant KR as "System Keyring" + + Caller->>CM: load() + CM->>FS: check exists / read + alt file missing + FS-->>CM: not found + CM-->>Caller: False + else file present + FS-->>CM: JSON content + CM->>CM: Configuration.from_dict(data) + CM->>KR: get_password(cluster_api_key) + CM->>KR: get_password(main_api_key) + CM->>KR: get_password(fallback_api_key) + KR-->>CM: key values (or None) + CM-->>Caller: True + end +``` + +### Save Flow + +```mermaid +sequenceDiagram + participant Caller + participant CM as ConfigManager + participant FS as "config.json" + participant KR as "System Keyring" + + Caller->>CM: save(fields..., api_keys...) + CM->>FS: ensure_directory(CONFIG_DIR) + alt no in-memory config + CM->>CM: load() existing or create default Configuration + end + CM->>CM: apply provided field updates + CM->>CM: Configuration.validate() (if models set) + CM->>KR: set_password(...) for each provided API key + CM->>FS: write JSON (version + Configuration.to_dict()) + CM-->>Caller: done (or raises ConfigurationError) +``` + +Key behaviors: +- **Partial updates**: `save()` accepts every field as an optional keyword argument; only provided values overwrite the in-memory `Configuration`, allowing incremental configuration (e.g., `codewiki configure` sub-commands that set one field at a time). +- **Validation gate**: full validation (`Configuration.validate()`) only runs once `main_model` and `cluster_model` are both set, avoiding premature failures during multi-step setup. +- **Keyring failures surface as `ConfigurationError`**, with an actionable message when the OS keychain is unavailable or misconfigured. +- **`is_configured()`** combines two checks: all three API keys must be retrievable from keyring, and `Configuration.is_complete()` must be true (all three model names set). +- **`clear()`** performs a full reset — deleting API keys from keyring and removing `config.json` — used by commands like `codewiki configure --reset`. + +## Bridging to Runtime Execution: to_backend_config + +`Configuration.to_backend_config()` is the seam between this module's persistent settings and the runtime configuration consumed by the documentation-generation backend. It: + +1. Fetches any missing API keys from the keyring via a fresh `ConfigManager` instance (if not explicitly passed in). +2. Merges `runtime_instructions` (per-invocation `AgentInstructions`) over the persisted `agent_instructions`, with runtime values taking precedence field-by-field. +3. Constructs and returns a backend `Config` object (`Config.from_cli(...)`) populated with all model, token, temperature, and clustering settings, ready to drive a documentation job. + +```mermaid +flowchart LR + A["CLI Command"] --> B["ConfigManager.load()"] + B --> C["Configuration"] + C --> D["Configuration.to_backend_config()"] + D --> E["ConfigManager (keyring lookup for missing keys)"] + D --> F["Merge AgentInstructions (runtime over persisted)"] + D --> G["Config.from_cli(...)"] + G --> H["Backend Config"] +``` + +This bridging pattern keeps the CLI's persistent, user-facing settings model (`Configuration`) decoupled from the backend's execution-time `Config` model. The backend `Config` class itself is documented in the [Config Core](../config-core.md) module. + +## Relationship to Other CLI Modules + +- **[Job Models](../job_models/job_models.md)** — represents the runtime state of an in-progress documentation job (`DocumentationJob`, `JobStatus`, `LLMConfig`, `GenerationOptions`, `JobStatistics`). Where `Configuration` describes *persistent user preferences*, the job models describe the *live execution* of a single generation run, often derived from a `Configuration` via `to_backend_config()`. +- **[Generation](../generation/generation.md)** — the `CLIDocumentationGenerator` adapter consumes a resolved `Configuration`/backend `Config` to drive the documentation pipeline. +- **[Utils](../utils/utils.md)** — provides shared filesystem (`ensure_directory`, `safe_read`, `safe_write`) and error types (`ConfigurationError`, `FileSystemError`) used internally by `ConfigManager`. + +## Summary + +The Configuration module is the trust boundary for user secrets and the single source of truth for persistent CLI preferences. `Configuration` and `AgentInstructions` define what can be configured; `ConfigManager` defines how those settings are safely loaded, validated, updated, and translated into a form the backend documentation pipeline can execute against. diff --git a/docs/reference/architecture/cli-core/generation.md b/docs/reference/architecture/cli-core/generation.md new file mode 100644 index 00000000..fb1d79d9 --- /dev/null +++ b/docs/reference/architecture/cli-core/generation.md @@ -0,0 +1,167 @@ +# Generation + +The Generation module is the central orchestration layer of the CodeWiki CLI. It bridges the command-line interface with the backend documentation engine, coordinating the full end-to-end pipeline that turns a source code repository into a structured, hierarchical documentation set — including dependency analysis, LLM-driven module clustering, documentation generation, optional Mermaid diagram extraction, and optional HTML output. + +Its single core component, `CLIDocumentationGenerator`, acts as an **adapter**: it wraps the backend's `DocumentationGenerator` (from [Backend Core](../backend-core.md)) with CLI-specific concerns such as progress reporting, colored logging, verbose diagnostics, and job lifecycle tracking via the [Job Models](../job_models/job_models.md) module. + +## Purpose and Role in the System + +As a child of [CLI Core](../cli-core.md), the Generation module is invoked whenever a user runs a documentation generation command. It does not implement dependency analysis, clustering, or documentation writing itself — those responsibilities belong to `backend-core`. Instead, Generation is responsible for: + +1. **Translating CLI configuration** into the backend's `Config` object (model selection, API keys, base URLs, token limits, clustering depth, agent instructions, additional source paths). +2. **Orchestrating the multi-stage pipeline** (dependency analysis → module clustering → documentation generation → optional HTML generation → finalization) with a five-stage `ProgressTracker`. +3. **Tracking job state** using a `DocumentationJob` model, recording statistics, generated files, and success/failure outcomes. +4. **Handling synthetic module fallback** to avoid context-window overflows when the LLM clustering step returns an empty module tree. +5. **Optionally generating an HTML viewer** via [HTML Generation](../html_generation/html_generation.md) and extracting Mermaid diagrams into a dedicated directory. + +## Architecture Overview + +```mermaid +flowchart TD + CLI["CLI Command Entry Point"] --> CDG["CLIDocumentationGenerator"] + CDG --> PT["ProgressTracker (utils)"] + CDG --> Job["DocumentationJob (job_models)"] + CDG --> BC["BackendConfig (config-core)"] + CDG --> DG["DocumentationGenerator (backend-core)"] + DG --> GB["DependencyGraphBuilder"] + DG --> AO["AgentOrchestrator"] + CDG --> CM["cluster_modules (backend-core)"] + CDG --> HG["HTMLGenerator (html_generation)"] + CDG --> LOG["ColoredFormatter (backend-core logging)"] +``` + +## Core Component + +### CLIDocumentationGenerator + +`CLIDocumentationGenerator` (`codewiki/cli/adapters/doc_generator.py`) is instantiated with: + +- `repo_path` — path to the repository being documented +- `output_dir` — target directory for generated docs +- `config` — a dictionary of LLM/model configuration (models, API keys, base URLs, token limits, temperature settings, clustering parameters, agent instructions, additional source paths) +- `verbose` — whether to emit detailed diagnostic output +- `generate_html` — whether to produce an HTML viewer (`index.html`) +- `diagrams_dir` — optional separate output directory for extracted Mermaid diagrams + +On construction, it: +- Creates a `ProgressTracker` configured for 5 pipeline stages. +- Creates a `DocumentationJob` and populates its metadata (repository path/name, output directory, `LLMConfig`). +- Configures backend logging by attaching a `ColoredFormatter`-based handler to the `codewiki.src.be` logger namespace, switching verbosity between `INFO` (verbose) and `WARNING` (quiet) levels, and disabling propagation to avoid duplicate log lines. + +#### Key Responsibilities + +| Responsibility | Method | +|---|---| +| Build backend configuration from CLI config dict | `generate()` | +| Run the full async generation pipeline | `_run_backend_generation()` | +| Generate the HTML viewer | `_run_html_generation()` | +| Ensure job metadata file exists | `_finalize_job()` | +| Configure colored backend logging | `_configure_backend_logging()` | + +## Generation Pipeline + +The `generate()` method is the public entry point. It is synchronous from the caller's perspective but internally drives an `asyncio`-based backend pipeline. It returns a completed `DocumentationJob` (see [Job Models](../job_models/job_models.md)) or raises an `APIError` on failure. + +```mermaid +sequenceDiagram + participant Caller as "CLI Command" + participant CDG as "CLIDocumentationGenerator" + participant BC as "BackendConfig" + participant DG as "DocumentationGenerator" + participant CM as "cluster_modules" + participant HG as "HTMLGenerator" + participant Job as "DocumentationJob" + + Caller->>CDG: generate() + CDG->>Job: start() + CDG->>BC: Config.from_cli(...) + CDG->>DG: _run_backend_generation(backend_config) + DG->>DG: graph_builder.build_dependency_graph() + DG->>CM: cluster_modules(leaf_nodes, components, config) + Note over DG,CM: Synthetic module fallback if tree is empty + DG->>DG: generate_module_documentation(components, leaf_nodes) + opt diagrams_dir configured + DG->>DG: extract_and_save_mermaid_diagrams() + end + opt generate_html is true + CDG->>HG: generate(output_path, ...) + end + CDG->>CDG: _finalize_job() + CDG->>Job: complete() + CDG-->>Caller: DocumentationJob +``` + +### Stage 1 — Dependency Analysis + +Instantiates the backend `DocumentationGenerator` from [Documentation Generator](../documentation-generator/documentation-generator.md) (which internally wires up `DependencyGraphBuilder` and `AgentOrchestrator` from [Backend Core](../backend-core.md)) and calls `doc_generator.graph_builder.build_dependency_graph()`. The result is a map of `components` and a list of `leaf_nodes`. Statistics (`total_files_analyzed`, `leaf_nodes`) are recorded on the `DocumentationJob`. Failures are wrapped as `APIError("Dependency analysis failed: ...")`. + +### Stage 2 — Module Clustering + +Loads a cached `first_module_tree.json` if present, otherwise calls `cluster_modules(leaf_nodes, components, backend_config)` to invoke the LLM clustering model. The result is cached to disk (`first_module_tree_path`) and then persisted as the working module tree (`module_tree_path`). + +**Synthetic Module Patch**: If the module tree ends up empty despite having leaf nodes (to prevent an LLM "whole-repo" fallback that could exceed API context limits), the generator batches leaf nodes into synthetic modules of a configurable size (`CODEWIKI_MAX_FILES_PER_MODULE`, default `5`) and re-persists the tree. This safeguard applies even when loading from cache, closing a previously identified cache-bypass gap. + +The final module count is stored on the `DocumentationJob.module_count` field. + +### Stage 3 — Documentation Generation + +Calls `doc_generator.generate_module_documentation(components, leaf_nodes)`, which performs the topologically-ordered (leaf-first) generation of module documentation via the `AgentOrchestrator`. After generation: + +- `doc_generator.create_documentation_metadata(...)` writes `metadata.json`. +- Generated `.md` and `.json` files in the output directory are collected into `DocumentationJob.files_generated`. +- If `diagrams_dir` was configured, Mermaid diagrams embedded in generated markdown are extracted via `extract_and_save_mermaid_diagrams` and indexed with `create_diagrams_readme`. + +### Stage 4 — HTML Generation (Optional) + +If `generate_html=True`, `_run_html_generation()` uses the [HTML Generation](../html_generation/html_generation.md) module's `HTMLGenerator` to detect repository info (name, URL, GitHub Pages URL) and render `index.html`, auto-loading the module tree and metadata from the output directory. The generated file is appended to `DocumentationJob.files_generated`. + +### Stage 5 — Finalization + +`_finalize_job()` verifies that `metadata.json` exists in the output directory; if the backend did not already write it, the generator writes the job's own JSON representation (`DocumentationJob.to_json()`) as a fallback. + +## Configuration Mapping + +`generate()` builds the backend `Config` object (see [Config Core](../config-core.md)) via `Config.from_cli(...)`, translating the CLI's flat configuration dictionary into per-provider settings: + +```mermaid +flowchart LR + subgraph CLIConfig["CLI config dict"] + A1["main_model / cluster_model / fallback_model"] + A2["*_api_key"] + A3["*_base_url"] + A4["*_api_version"] + A5["*_max_tokens / *_temperature"] + A6["max_token_per_module / max_depth"] + A7["agent_instructions"] + A8["additional_paths"] + end + CLIConfig --> FromCLI["Config.from_cli()"] + FromCLI --> BackendConfig["Backend Config instance"] + BackendConfig --> DG2["DocumentationGenerator"] +``` + +Additional source paths supplied in the CLI config are normalized to absolute paths (resolved relative to `repo_path`) before being passed through as `additional_source_paths`. + +In verbose mode, the generator prints a detailed configuration summary (model names, base URLs, token limits, module settings, additional paths, and a preview of custom agent instructions) before kicking off Stage 1. + +## Progress Tracking and Logging + +The Generation module relies on utilities from the [Utils](../utils/utils.md) module: + +- **`ProgressTracker`**: manages a 5-stage weighted progress model (Dependency Analysis 40%, Module Clustering 20%, Documentation Generation 30%, HTML Generation 5%, Finalization 5%), providing `start_stage`, `update_stage`, `complete_stage`, elapsed-time formatting, and ETA estimation. +- **`CLILogger`**: complements verbose/non-verbose console output alongside `ProgressTracker`'s stage banners. + +Backend logs are captured by attaching a handler using `ColoredFormatter` (from `backend-core`'s Logging Config child module) directly to the `codewiki.src.be` logger, ensuring consistent colored output regardless of whether the CLI or backend emitted the log line, while preventing duplicate propagation to the root logger. + +## Error Handling + +All backend-facing calls (`build_dependency_graph`, `cluster_modules`, `generate_module_documentation`) are wrapped in `try`/`except` blocks that convert unexpected exceptions into `APIError` with contextual messages (e.g., `"Dependency analysis failed: ..."`). At the top level, `generate()` catches both `APIError` and generic `Exception`, marks the `DocumentationJob` as failed via `job.fail(str(e))`, and re-raises so the CLI layer can present the error to the user. + +## Relationship to Other Modules + +- **[CLI Core](../cli-core.md)**: parent module; Generation is one of its functional children alongside [Configuration](../configuration/configuration.md), [Job Models](../job_models/job_models.md), [Git Integration](../git_integration/git_integration.md), [HTML Generation](../html_generation/html_generation.md), and [Utils](../utils/utils.md). +- **[Job Models](../job_models/job_models.md)**: supplies `DocumentationJob`, `LLMConfig`, `JobStatus`, and `JobStatistics`, which Generation populates throughout the pipeline. +- **[HTML Generation](../html_generation/html_generation.md)**: invoked optionally at Stage 4 to render the documentation as a browsable HTML site. +- **[Utils](../utils/utils.md)**: provides `ProgressTracker` for stage-based progress reporting. +- **[Backend Core](../backend-core.md)**: supplies the actual documentation engine (`DocumentationGenerator`, `AgentOrchestrator`, `DependencyGraphBuilder`) that Generation orchestrates but does not reimplement. +- **[Config Core](../config-core.md)**: supplies the `Config` class used to translate CLI settings into the backend's runtime configuration. diff --git a/docs/reference/architecture/cli-core/git_integration.md b/docs/reference/architecture/cli-core/git_integration.md new file mode 100644 index 00000000..ee7b3ac3 --- /dev/null +++ b/docs/reference/architecture/cli-core/git_integration.md @@ -0,0 +1,147 @@ +# Git Integration + +The Git Integration module provides the `GitManager` component, which encapsulates all git repository operations required by the CodeWiki CLI's documentation workflow. It is responsible for validating that a target directory is a git repository, inspecting working-directory cleanliness, creating dedicated documentation branches, committing generated documentation, and deriving useful metadata such as remote URLs, branch names, commit hashes, and GitHub pull-request links. + +This module is a child of the [CLI Core](../cli_core.md) module and is used primarily by the CLI's generation workflow (see [Generation](../generation/generation.md)) to safely manage version control state before, during, and after documentation is generated. + +## Purpose and Scope + +When the CLI is run with branch-creation options (e.g. `--create-branch`), it must: + +1. Confirm the target path is inside a valid git repository. +2. Ensure there are no uncommitted changes that could be silently mixed into a documentation commit (unless the user explicitly forces the operation). +3. Create a uniquely named, timestamped branch dedicated to documentation output. +4. Stage and commit the generated documentation files. +5. Surface repository metadata (current branch, commit hash, remote URL) back to the CLI so it can print helpful summaries. +6. Build a ready-to-open GitHub pull-request URL if the remote is hosted on GitHub. + +All of this logic is centralized in the `GitManager` class, keeping git operations isolated from the rest of the CLI so that other components (progress reporting, configuration, documentation generation) do not need to know about the underlying git library (`GitPython`). + +## Core Component: GitManager + +`GitManager` (`CodeWiki.codewiki.cli.git_manager::GitManager`) wraps a `git.Repo` instance (from the `GitPython` library) and exposes a small, purpose-built API surface consumed by the CLI layer. + +### Initialization + +On construction, `GitManager` resolves the given `repo_path` to an absolute path and attempts to open it as a git repository using `git.Repo(repo_path, search_parent_directories=True)`. The `search_parent_directories=True` flag allows the manager to locate a `.git` directory even if `repo_path` points to a subdirectory of the repository (similar to how the native `git` command walks upward from the current directory). + +If no valid repository is found, a `RepositoryError` is raised with actionable guidance (e.g., suggesting `git init`). `RepositoryError` is a CLI-specific exception type that carries a dedicated exit code (`EXIT_REPOSITORY_ERROR`), allowing the CLI's top-level error handler to report failures consistently and set correct process exit statuses. + +### Key Responsibilities + +#### 1. Working Directory Cleanliness Check + +`check_clean_working_directory()` inspects the repository for uncommitted modifications or untracked files using `repo.is_dirty(untracked_files=True)`. When dirty, it builds a human-readable summary listing up to three modified and three untracked files (with a "... and N more" suffix for longer lists), returning a `(is_clean, status_message)` tuple. This check underpins the safety guarantee that documentation generation will not accidentally commit unrelated in-progress changes. + +#### 2. Documentation Branch Creation + +`create_documentation_branch(force: bool = False)` creates and checks out a new branch named `docs/codewiki-` (format `%Y%m%d-%H%M%S`). Before creating the branch it re-runs the cleanliness check unless `force=True`, raising a detailed `RepositoryError` with copy-pastable remediation commands (`git status`, `git add -A && git commit`, `git stash`) if the working directory is dirty. + +To guard against timestamp collisions, it checks existing branch names and appends an incrementing numeric suffix (`-1`, `-2`, ...) if a collision would otherwise occur. Branch creation and checkout are performed via `repo.create_head(...)` followed by `.checkout()`; any underlying `GitCommandError` is translated into a `RepositoryError`. + +#### 3. Committing Documentation + +`commit_documentation(docs_path: Path, message: Optional[str] = None)` stages the documentation output directory via `repo.index.add([str(docs_path)])` and commits it with either a caller-supplied message or the default `"Add generated documentation\n\nGenerated by CodeWiki CLI"`. It returns the resulting commit's `hexsha`. Failures during add/commit are wrapped in `RepositoryError`. + +#### 4. Repository Metadata Accessors + +- `get_remote_url(remote_name="origin")` — returns the URL of the named remote, or `None` if it does not exist. +- `get_current_branch()` — returns the active branch name, or the literal string `"HEAD"` when in a detached-HEAD state (caught via `TypeError` from `GitPython`). +- `get_commit_hash()` — returns the current `HEAD` commit's hex SHA. +- `branch_exists(branch_name)` — returns whether a branch with the given name already exists locally. + +#### 5. GitHub Pull-Request URL Construction + +`get_github_pr_url(branch_name)` derives a ready-to-use GitHub "compare" URL (`/compare/`) from the `origin` remote, but only when that remote points to `github.com`. It normalizes the remote URL by: +- Stripping a trailing slash and any `.git` suffix. +- Converting SSH-style remotes (`git@github.com:org/repo`) into HTTPS form (`https://github.com/org/repo`). + +If there is no remote or it is not a GitHub remote, it returns `None`, allowing the CLI to gracefully skip printing a PR link for non-GitHub repositories. + +## Error Handling + +All git-related failures surface as `RepositoryError` (defined outside this module in the CLI's shared error utilities), rather than leaking raw `GitCommandError` or `InvalidGitRepositoryError` exceptions from `GitPython`. This gives the CLI a single, well-known exception type to catch at the top level and translate into a clean error message and a dedicated process exit code, keeping git-library implementation details out of user-facing error handling. + +## Architecture + +```mermaid +classDiagram + class GitManager { + +Path repo_path + +Repo repo + +__init__(repo_path) + +check_clean_working_directory() Tuple + +create_documentation_branch(force) str + +commit_documentation(docs_path, message) str + +get_remote_url(remote_name) str + +get_current_branch() str + +get_commit_hash() str + +branch_exists(branch_name) bool + +get_github_pr_url(branch_name) str + } + class RepositoryError { + +__init__(message) + } + GitManager ..> RepositoryError : raises +``` + +## Interaction with the CLI Generation Workflow + +`GitManager` is typically instantiated by the CLI's generation flow when branch-based documentation workflows are requested. The sequence below illustrates a typical end-to-end interaction: verifying repository state, creating a branch, running documentation generation (handled by `CLIDocumentationGenerator` in the [Generation](../generation/generation.md) module), committing the results, and reporting a PR link. + +```mermaid +sequenceDiagram + participant CLI as "CLI Entry Point" + participant GM as "GitManager" + participant DocGen as "CLIDocumentationGenerator" + participant Repo as "git.Repo" + + CLI->>GM: "GitManager(repo_path)" + GM->>Repo: "git.Repo(repo_path)" + CLI->>GM: "check_clean_working_directory()" + GM-->>CLI: "(is_clean, status_message)" + alt "Working directory dirty and not forced" + GM-->>CLI: "raise RepositoryError" + else "Clean or forced" + CLI->>GM: "create_documentation_branch(force)" + GM->>Repo: "create_head(branch_name)" + GM->>Repo: "checkout()" + GM-->>CLI: "branch_name" + CLI->>DocGen: "generate()" + DocGen-->>CLI: "DocumentationJob" + CLI->>GM: "commit_documentation(docs_path, message)" + GM->>Repo: "index.add([docs_path])" + GM->>Repo: "index.commit(message)" + GM-->>CLI: "commit_hash" + CLI->>GM: "get_github_pr_url(branch_name)" + GM-->>CLI: "pr_url or None" + end +``` + +## Process Flow: Documentation Branch Creation + +```mermaid +flowchart TD + Start["create_documentation_branch(force)"] --> CheckForce{{force?}} + CheckForce -->|"false"| CheckClean["check_clean_working_directory()"] + CheckForce -->|"true"| GenName["Generate timestamped branch name"] + CheckClean --> IsClean{{"is_clean?"}} + IsClean -->|"false"| RaiseErr["raise RepositoryError with remediation steps"] + IsClean -->|"true"| GenName + GenName --> CheckExists{{"branch name exists?"}} + CheckExists -->|"true"| AppendCounter["Append incrementing counter suffix"] + AppendCounter --> CheckExists + CheckExists -->|"false"| CreateBranch["repo.create_head(branch_name)"] + CreateBranch --> Checkout["new_branch.checkout()"] + Checkout --> ReturnName["return branch_name"] +``` + +## Relationship to Other Modules + +- **[CLI Core](../cli_core.md)** — the parent module; `GitManager` is one of the core services orchestrated alongside configuration management, job models, and progress tracking to deliver the full `codewiki generate` experience. +- **[Generation](../generation/generation.md)** — `CLIDocumentationGenerator` performs the actual documentation generation (dependency analysis, clustering, LLM-driven writing, optional HTML output). The CLI entry point coordinates `GitManager` and `CLIDocumentationGenerator` together: creating a branch before generation, and committing the resulting files afterward. +- **[Job Models](../job_models/job_models.md)** — while `GitManager` itself does not depend on `DocumentationJob`, the surrounding CLI workflow often records git-derived metadata (branch name, commit hash) alongside job statistics for reporting purposes. + +## Dependencies + +`GitManager` depends on the third-party [`GitPython`](https://gitpython.readthedocs.io/) library (imported as `git`) for all low-level repository interactions, including `git.Repo`, `git.InvalidGitRepositoryError`, and `git.exc.GitCommandError`. It also depends on the CLI's shared `RepositoryError` exception type for consistent error reporting across the CLI. diff --git a/docs/reference/architecture/cli-core/html_generation.md b/docs/reference/architecture/cli-core/html_generation.md new file mode 100644 index 00000000..40f0eb9d --- /dev/null +++ b/docs/reference/architecture/cli-core/html_generation.md @@ -0,0 +1,131 @@ +# Html Generation + +The Html Generation module produces a self-contained, static HTML documentation viewer suitable for GitHub Pages (or any static file host). It is the final, optional presentation layer of the CodeWiki CLI pipeline: after the [Generation](../generation/generation.md) stage has produced Markdown documentation files, a module tree, and metadata, the Html Generation module packages that output into a single browsable `index.html` file with embedded configuration, styles, and client-side rendering logic. + +## Purpose and Scope + +The `HTMLGenerator` class is the sole core component of this module. Its responsibilities are: + +- **Template loading** — reads a static HTML template (`viewer_template.html`) shipped with the package. +- **Data discovery** — auto-loads `module_tree.json` and `metadata.json` from a documentation output directory when explicit data is not supplied. +- **Placeholder substitution** — injects title, repository link, embedded JSON data (module tree, metadata, config), and an "info panel" HTML fragment into the template. +- **Repository introspection** — inspects a local git repository to derive a display name, remote URL, and a predicted GitHub Pages URL. +- **Atomic file output** — writes the final HTML file safely using the shared filesystem utilities. + +This module has no knowledge of *how* documentation content was produced; it only consumes the artifacts (`module_tree.json`, `metadata.json`, generated Markdown files) that the [Generation](../generation/generation.md) module and the backend documentation pipeline produce. This keeps Html Generation a pure "rendering/packaging" concern, decoupled from LLM orchestration, dependency analysis, and job tracking. + +## Core Component + +### HTMLGenerator + +`HTMLGenerator` (`codewiki/cli/html_generator.py`) encapsulates all HTML viewer generation logic. + +| Method | Responsibility | +|---|---| +| `__init__(template_dir)` | Resolves the template directory, defaulting to the package's `templates/github_pages` folder. | +| `load_module_tree(docs_dir)` | Reads `module_tree.json` from the docs directory; falls back to a minimal single-node structure if the file is missing. | +| `load_metadata(docs_dir)` | Reads `metadata.json`; returns `None` (non-critical) if missing or unparsable. | +| `generate(...)` | Orchestrates the full generation: auto-loads data, builds the info panel, computes paths/links, serializes JSON, performs placeholder substitution, and writes the output file. | +| `_build_info_content(metadata)` | Builds an HTML fragment (model name, generation timestamp, commit hash, component count, max depth) displayed in the viewer's info panel. | +| `_escape_html(text)` | Escapes HTML-sensitive characters to prevent malformed markup when embedding user/repo-derived strings (e.g., title). | +| `detect_repository_info(repo_path)` | Uses `GitPython` to read the repository name, normalize the remote URL (including `git@github.com:` SSH URLs), and compute the expected `https://.github.io//` Pages URL. | + +## Architecture + +```mermaid +flowchart TD + Docs["Documentation Output Directory"] --> ModuleTreeJson["module_tree.json"] + Docs --> MetadataJson["metadata.json"] + Template["viewer_template.html"] --> Generator["HTMLGenerator"] + ModuleTreeJson --> Generator + MetadataJson --> Generator + RepoPath["Repository Path"] --> DetectInfo["detect_repository_info()"] + DetectInfo --> Generator + Generator -->|"safe_write()"| IndexHtml["index.html"] + Generator -->|"on failure"| FSError["FileSystemError"] +``` + +### Dependencies + +- **`codewiki.cli.utils.fs`** — `safe_read` / `safe_write` provide atomic, encoding-safe file I/O used to load the template/JSON files and write the final HTML output. See the [Utils](../utils/utils.md) module for other shared CLI utilities such as logging and progress tracking. +- **`codewiki.cli.utils.errors::FileSystemError`** — raised when the template file is missing or when reading/writing fails, allowing the CLI layer to surface a consistent error type. +- **`git` (GitPython)** — used only within `detect_repository_info` to introspect the local repository; failures are caught and silently ignored, degrading gracefully to a viewer without repository links. + +## Integration with the CLI Pipeline + +Html Generation is invoked as an optional, final stage by [`CLIDocumentationGenerator`](../generation/generation.md) (from the [Generation](../generation/generation.md) module), which drives the overall CLI workflow: dependency analysis → module clustering → documentation generation → **HTML generation** → job finalization. + +```mermaid +sequenceDiagram + participant CLIGen as "CLIDocumentationGenerator" + participant HTMLGen as "HTMLGenerator" + participant FS as "safe_read / safe_write" + participant Git as "GitPython" + + CLIGen->>HTMLGen: HTMLGenerator() + CLIGen->>HTMLGen: detect_repository_info(repo_path) + HTMLGen->>Git: Repo(repo_path) + Git-->>HTMLGen: remote URL, name + HTMLGen-->>CLIGen: name, url, github_pages_url + CLIGen->>HTMLGen: generate(output_path, title, repository_url, docs_dir) + HTMLGen->>FS: safe_read(module_tree.json) + HTMLGen->>FS: safe_read(metadata.json) + HTMLGen->>FS: safe_read(viewer_template.html) + HTMLGen->>HTMLGen: _build_info_content(metadata) + HTMLGen->>HTMLGen: substitute placeholders + HTMLGen->>FS: safe_write(index.html) + HTMLGen-->>CLIGen: index.html written +``` + +This corresponds to the `_run_html_generation` step inside `CLIDocumentationGenerator.generate()`: it is only executed when the CLI was invoked with `generate_html=True`, after the backend has produced Markdown files, `module_tree.json`, and `metadata.json` in the output directory. On success, `"index.html"` is appended to `DocumentationJob.files_generated` (see the [Job Models](../job_models/job_models.md) module). + +## Generation Flow + +```mermaid +flowchart TD + Start["generate() called"] --> CheckDocsDir{{"docs_dir provided?"}} + CheckDocsDir -->|"yes"| AutoLoadTree["load_module_tree(docs_dir)"] + CheckDocsDir -->|"yes"| AutoLoadMeta["load_metadata(docs_dir)"] + CheckDocsDir -->|"no"| UseProvided["use provided module_tree/metadata"] + AutoLoadTree --> Defaults + AutoLoadMeta --> Defaults + UseProvided --> Defaults + Defaults["Apply defaults for module_tree/config"] --> LoadTemplate["Load viewer_template.html"] + LoadTemplate -->|"missing"| RaiseErr["raise FileSystemError"] + LoadTemplate -->|"found"| BuildInfo["_build_info_content(metadata)"] + BuildInfo --> BuildRepoLink["Build repository link HTML"] + BuildRepoLink --> ComputeBasePath["Compute docs_base_path"] + ComputeBasePath --> SerializeJson["Serialize config/module_tree/metadata to JSON"] + SerializeJson --> Replace["Replace template placeholders"] + Replace --> WriteOut["safe_write(output_path)"] + WriteOut --> Done["index.html generated"] +``` + +### Template Placeholders + +The generator performs a straightforward string substitution over the template file. The following placeholders are populated by `generate()`: + +| Placeholder | Source | +|---|---| +| `{{TITLE}}` | Escaped `title` argument (e.g., repository name) | +| `{{REPO_LINK}}` | HTML anchor to `repository_url`, empty if not provided | +| `{{SHOW_INFO}}` | `"block"` or `"none"` depending on whether info content was built | +| `{{INFO_CONTENT}}` | HTML fragment from `_build_info_content` (model, timestamp, commit, stats) | +| `{{CONFIG_JSON}}` | JSON-serialized `config` dictionary | +| `{{MODULE_TREE_JSON}}` | JSON-serialized module tree structure | +| `{{METADATA_JSON}}` | JSON-serialized metadata, or the literal `null` | +| `{{DOCS_BASE_PATH}}` | Relative path from the output file to the docs directory | + +## Error Handling + +- Missing `viewer_template.html` raises `FileSystemError`, propagated up to the CLI layer. +- Missing or malformed `module_tree.json` is either substituted with a minimal fallback structure (`load_module_tree`) or, on unexpected read/parse errors, raises `FileSystemError`. +- Missing or malformed `metadata.json` is treated as non-critical: `load_metadata` swallows exceptions and returns `None`, resulting in the info panel being hidden (`{{SHOW_INFO}} = "none"`). +- Git introspection failures in `detect_repository_info` are caught broadly, so the generator always returns a usable (if partially empty) info dictionary rather than failing the whole pipeline. + +## Relationship to Other Modules + +- **[Generation](../generation/generation.md)** — the `CLIDocumentationGenerator` orchestrates when and how `HTMLGenerator` is invoked as part of the overall documentation generation job. +- **[Job Models](../job_models/job_models.md)** — the generated `index.html` file path is recorded in the `DocumentationJob.files_generated` list. +- **[Utils](../utils/utils.md)** — shares low-level filesystem helpers (`safe_read`/`safe_write`) and error types used throughout the CLI. +- **[cli-core](../cli-core.md)** — the parent module that aggregates Html Generation alongside generation, configuration, job models, git integration, and shared utilities into the full CLI toolchain. diff --git a/docs/reference/architecture/cli-core/job_models.md b/docs/reference/architecture/cli-core/job_models.md new file mode 100644 index 00000000..7b3e33bd --- /dev/null +++ b/docs/reference/architecture/cli-core/job_models.md @@ -0,0 +1,244 @@ +# Job Models + +## Introduction + +The Job Models module defines the core data structures used to represent, track, and persist documentation generation jobs within the CodeWiki CLI. It is a foundational, dependency-free module that provides typed dataclasses and enums for job state, configuration, and statistics — enabling consistent serialization, deserialization, and status tracking across the CLI documentation pipeline. + +This module is a child of the [CLI Core](../cli_core.md) module and is consumed by sibling modules such as [Generation](../generation/generation.md), [Configuration](../configuration/configuration.md), [Utils](../utils/utils.md), and other CLI orchestration components that need to create, update, or persist job records. + +## Purpose and Scope + +The Job Models module has a single, focused responsibility: **define the shape of a documentation job and its lifecycle**. It does not perform any I/O, orchestration, or business logic beyond simple state transitions and serialization helpers. This keeps the module lightweight, easily testable, and safe to import from anywhere in the CLI codebase without introducing circular dependencies. + +Key responsibilities: +- Define the `JobStatus` enum representing the lifecycle states of a job. +- Define `GenerationOptions` for user-configurable generation behavior (branching, GitHub Pages, caching, output paths). +- Define `LLMConfig` for capturing which language models and endpoint were used for a run. +- Define `JobStatistics` for capturing quantitative results of a run (files analyzed, tree depth, tokens consumed). +- Define `DocumentationJob`, the aggregate root that ties together identity, timing, status, configuration, and results for a single documentation generation run. +- Provide robust `to_dict()` / `to_json()` / `from_dict()` methods for safe persistence and rehydration, including defensive coercion of malformed or partial data. + +## Core Components + +### JobStatus + +`JobStatus` is a `str`-backed `Enum` representing the possible lifecycle states of a documentation job: + +| Value | Meaning | +|---|---| +| `pending` | Job has been created but not yet started | +| `running` | Job is actively executing | +| `completed` | Job finished successfully | +| `failed` | Job terminated with an error | + +Because it inherits from both `str` and `Enum`, instances serialize naturally to plain strings in JSON output while still supporting type-safe comparisons in code (e.g., `job.status == JobStatus.RUNNING`). + +### GenerationOptions + +`GenerationOptions` is a dataclass capturing user-facing flags that control how a documentation run behaves: + +- `create_branch` (`bool`): Whether to create a dedicated git branch for the generated docs. Consumed by the [Git Integration](../git_integration/git_integration.md) module. +- `github_pages` (`bool`): Whether to prepare output for GitHub Pages publishing. +- `no_cache` (`bool`): Whether to bypass any caching layer during generation. +- `custom_output` (`Optional[str]`): An optional override for the output directory. + +### LLMConfig + +`LLMConfig` records which language models and endpoint were used to produce a job's output: + +- `main_model` (`str`): The primary model used for documentation generation. +- `cluster_model` (`str`): The model used for clustering/module grouping decisions. +- `base_url` (`str`): The API base URL for the LLM provider. + +This is a plain, required-field dataclass (no defaults), reflecting that when an `LLMConfig` is attached to a job, all three fields are expected to be known. + +### JobStatistics + +`JobStatistics` aggregates quantitative metrics collected during a documentation run: + +- `total_files_analyzed` (`int`): Count of source files processed. +- `leaf_nodes` (`int`): Count of leaf-level components/modules identified. +- `max_depth` (`int`): Maximum depth of the module hierarchy produced. +- `total_tokens_used` (`int`): Total LLM tokens consumed for the run. + +All fields default to `0`, so a freshly created job has a valid, zeroed-out statistics object even before execution begins. + +### DocumentationJob + +`DocumentationJob` is the central aggregate of this module. It represents a single documentation generation run end-to-end, combining identity, repository context, git metadata, timing, status, and nested configuration/statistics objects. + +**Fields:** + +| Field | Type | Description | +|---|---|---| +| `job_id` | `str` | UUID4 identifier, auto-generated if not supplied | +| `repository_path` | `str` | Absolute filesystem path to the repository being documented | +| `repository_name` | `str` | Human-readable repository name | +| `output_directory` | `str` | Destination directory for generated docs | +| `commit_hash` | `str` | Git commit SHA the job was run against | +| `branch_name` | `Optional[str]` | Git branch name, if applicable | +| `timestamp_start` | `str` | ISO-format start timestamp, auto-populated | +| `timestamp_end` | `Optional[str]` | ISO-format end timestamp, set on completion/failure | +| `status` | `JobStatus` | Current lifecycle status | +| `error_message` | `Optional[str]` | Populated when the job fails | +| `files_generated` | `List[str]` | Paths of documentation files produced | +| `module_count` | `int` | Number of modules documented | +| `generation_options` | `GenerationOptions` | Options selected for this run | +| `llm_config` | `Optional[LLMConfig]` | LLM configuration used, if known | +| `statistics` | `JobStatistics` | Quantitative results of the run | + +**Lifecycle methods:** + +- `start()` — Transitions status to `RUNNING` and refreshes `timestamp_start`. +- `complete()` — Transitions status to `COMPLETED` and sets `timestamp_end`. +- `fail(error_message)` — Transitions status to `FAILED`, records the error, and sets `timestamp_end`. + +**Serialization methods:** + +- `to_dict()` — Produces a JSON-serializable `dict` representation, flattening nested dataclasses (`GenerationOptions`, `LLMConfig`, `JobStatistics`) into plain dicts and converting `JobStatus` to its string value. +- `to_json()` — Convenience wrapper around `to_dict()` that returns a pretty-printed JSON string (2-space indent). +- `from_dict(data)` — Classmethod that reconstructs a `DocumentationJob` from a raw dictionary (e.g., loaded from a JSON file), using defensive coercion helpers to tolerate missing, malformed, or partial fields. + +### Defensive Coercion Helpers + +The module defines several private module-level helper functions used exclusively by `DocumentationJob.from_dict()` to safely rebuild nested objects from untrusted or partial data: + +- `_coerce_job_status(value, default)` — Converts a raw value to a valid `JobStatus`, falling back to `JobStatus.PENDING` (or a supplied default) if the value is `None` or not a recognized status string. +- `_coerce_int(value, default)` — Safely converts a value to `int`, returning a default (`0`) on `None`, `TypeError`, or `ValueError`. +- `_coerce_generation_options(value)` — Rebuilds a `GenerationOptions` from a dict, defaulting missing keys. +- `_coerce_llm_config(value)` — Rebuilds an `LLMConfig` from a dict, or returns `None` if no data is present. +- `_coerce_statistics(value)` — Rebuilds a `JobStatistics` from a dict, defaulting missing/invalid numeric fields to `0`. + +This coercion layer makes `DocumentationJob.from_dict()` resilient to schema drift (e.g., loading job files written by an older version of the CLI) without raising exceptions during deserialization. + +## Architecture + +### Component Structure + +```mermaid +classDiagram + class JobStatus { + <> + PENDING + RUNNING + COMPLETED + FAILED + } + class GenerationOptions { + +bool create_branch + +bool github_pages + +bool no_cache + +str custom_output + } + class LLMConfig { + +str main_model + +str cluster_model + +str base_url + } + class JobStatistics { + +int total_files_analyzed + +int leaf_nodes + +int max_depth + +int total_tokens_used + } + class DocumentationJob { + +str job_id + +str repository_path + +str repository_name + +str output_directory + +str commit_hash + +str branch_name + +str timestamp_start + +str timestamp_end + +JobStatus status + +str error_message + +List files_generated + +int module_count + +GenerationOptions generation_options + +LLMConfig llm_config + +JobStatistics statistics + +start() + +complete() + +fail(error_message) + +to_dict() dict + +to_json() str + +from_dict(data) DocumentationJob + } + DocumentationJob --> JobStatus : "status" + DocumentationJob --> GenerationOptions : "generation_options" + DocumentationJob --> LLMConfig : "llm_config (optional)" + DocumentationJob --> JobStatistics : "statistics" +``` + +### Job Lifecycle State Machine + +```mermaid +stateDiagram-v2 + [*] --> PENDING : "DocumentationJob() created" + PENDING --> RUNNING : "start()" + RUNNING --> COMPLETED : "complete()" + RUNNING --> FAILED : "fail(error_message)" + COMPLETED --> [*] + FAILED --> [*] +``` + +### Serialization / Deserialization Flow + +```mermaid +sequenceDiagram + participant Caller as "CLI Component" + participant Job as "DocumentationJob" + participant Coerce as "Coercion Helpers" + + Caller->>Job: "DocumentationJob(...)" + Job-->>Caller: "job instance (status=PENDING)" + Caller->>Job: "job.start()" + Caller->>Job: "job.to_dict() / job.to_json()" + Job-->>Caller: "dict / JSON string" + Note over Caller: "Persisted to disk or transmitted" + + Caller->>Job: "DocumentationJob.from_dict(raw_data)" + Job->>Coerce: "_coerce_job_status(raw_data.status)" + Job->>Coerce: "_coerce_int(raw_data.module_count)" + Job->>Coerce: "_coerce_generation_options(raw_data.generation_options)" + Job->>Coerce: "_coerce_llm_config(raw_data.llm_config)" + Job->>Coerce: "_coerce_statistics(raw_data.statistics)" + Coerce-->>Job: "validated nested objects" + Job-->>Caller: "reconstructed DocumentationJob" +``` + +## Integration with Other Modules + +The Job Models module is intentionally dependency-free (it only imports from the Python standard library), which allows it to be safely imported by any layer of the CLI without risk of circular imports: + +- **[CLI Core](../cli_core.md)** (parent module): Exposes `DocumentationJob`, `GenerationOptions`, `JobStatistics`, `JobStatus`, and `LLMConfig` as part of its public component surface, making them available to all CLI subsystems. +- **[Generation](../generation/generation.md)**: The `CLIDocumentationGenerator` creates and drives `DocumentationJob` instances through their lifecycle (`start()` → `complete()`/`fail()`), populating `files_generated`, `module_count`, and `statistics` as generation proceeds. +- **[Configuration](../configuration/configuration.md)**: `ConfigManager` and the `Configuration`/`AgentInstructions` models supply values (such as model names and base URLs) that are captured into a job's `LLMConfig`. +- **[Git Integration](../git_integration/git_integration.md)**: `GitManager` supplies `commit_hash` and `branch_name` values and consults `GenerationOptions.create_branch` to decide whether to create a dedicated branch for generated documentation. +- **[Utils](../utils/utils.md)**: `CLILogger` and `ProgressTracker`/`ModuleProgressBar` report on job progress using the status and statistics captured on `DocumentationJob` instances. +- **[HTML Generation](../html_generation/html_generation.md)**: `HTMLGenerator` may reference `files_generated` and `output_directory` from a completed job when producing rendered output. + +```mermaid +flowchart TD + JobModels["Job Models"] + CliCore["CLI Core (parent)"] + Generation["Generation"] + Configuration["Configuration"] + GitIntegration["Git Integration"] + HtmlGeneration["HTML Generation"] + Utils["Utils"] + + CliCore --> JobModels + Generation -->|"creates & drives"| JobModels + Configuration -->|"populates LLMConfig"| JobModels + GitIntegration -->|"populates commit_hash / branch_name"| JobModels + HtmlGeneration -->|"reads files_generated"| JobModels + Utils -->|"reports on status/statistics"| JobModels +``` + +## Design Notes + +- **Immutability of shape, mutability of state**: While the dataclasses themselves are mutable (no `frozen=True`), the module encourages controlled mutation through explicit lifecycle methods (`start()`, `complete()`, `fail()`) rather than direct field assignment, keeping status transitions predictable and auditable. +- **Safe round-tripping**: The combination of `to_dict()`/`to_json()` and `from_dict()` with defensive coercion helpers ensures that job records can be persisted to disk (e.g., as job history or resumable state) and reloaded reliably even if the on-disk schema is incomplete or slightly out of date. +- **String-backed enum**: `JobStatus` extending `str` means job status values serialize directly as readable strings in JSON without requiring custom encoders, simplifying integration with logging, CLI output, and external tooling. +- **Zero external dependencies**: The module relies only on `dataclasses`, `datetime`, `typing`, `enum`, `uuid`, and `json` from the standard library, making it trivially testable and reusable across CLI contexts. diff --git a/docs/reference/architecture/config-core/config-core.md b/docs/reference/architecture/config-core/config-core.md new file mode 100644 index 00000000..8b26d412 --- /dev/null +++ b/docs/reference/architecture/config-core/config-core.md @@ -0,0 +1,220 @@ +# Config Core + +The Config Core module defines the central `Config` dataclass that CodeWiki uses to drive every stage of the documentation-generation pipeline. It is the single source of truth for repository paths, output directories, LLM provider settings (models, API keys, base URLs, temperatures, and token limits), and agent-instruction customization such as include/exclude file patterns, focus modules, documentation type, and free-form custom instructions. + +Config Core is intentionally small and dependency-light: it contains one component, `Config`, but it is consumed heavily by [Backend Core](backend-core.md) (which executes the actual analysis and generation pipeline), [CLI Core](cli-core.md) (which builds a `Config` from a locally persisted configuration file), and [Frontend Core](frontend-core.md) (which builds a `Config` for each web-submitted documentation job). Because nearly every other module depends on `Config`, understanding its shape and construction paths is a prerequisite for understanding the rest of the system. + +## Purpose and Responsibilities + +`Config` serves three distinct responsibilities: + +1. **Data container** — holds all repository, output, and per-provider LLM settings (cluster/main/fallback models, keys, base URLs, API versions, max tokens, temperatures, and the field name used to send max-tokens to the provider). +2. **Construction factory** — offers four classmethod entry points (`from_args`, `from_web_job`, `from_cli`, `from_config_manager`) that adapt different callers' inputs (argparse namespaces, web job parameters, explicit CLI parameters, or a `ConfigManager`) into a fully validated `Config` instance. +3. **Derived-behavior provider** — exposes properties and helper methods (`include_patterns`, `exclude_patterns`, `focus_modules`, `doc_type`, `custom_instructions`, `all_source_paths`, `is_multi_path_mode`, `get_prompt_addition`, `validate_source_paths`) that downstream consumers use instead of re-implementing agent-instruction parsing or multi-path logic themselves. + +## Core Component + +### Config (`codewiki/src/config.py`) + +`Config` is a `@dataclass` with required fields (`repo_path`, `output_dir`, `dependency_graph_dir`, `docs_dir`, `max_depth`, `main_model`, `cluster_model`, `fallback_model`, `cluster_api_key`, `main_api_key`, `fallback_api_key`) and a large set of optional fields with sensible defaults for base URLs, API versions, max-tokens, temperatures, and multi-path/diagram directory support. + +```mermaid +classDiagram + class Config { + +str repo_path + +str output_dir + +str dependency_graph_dir + +str docs_dir + +int max_depth + +str main_model + +str cluster_model + +str fallback_model + +str cluster_api_key + +str main_api_key + +str fallback_api_key + +Optional~str~ cluster_base_url + +Optional~str~ main_base_url + +Optional~str~ fallback_base_url + +int cluster_max_tokens + +int main_max_tokens + +int fallback_max_tokens + +int max_token_per_module + +int max_token_per_leaf_module + +float cluster_temperature + +float main_temperature + +float fallback_temperature + +Optional~Dict~ agent_instructions + +Optional~str~ diagrams_dir + +Optional~List~str~~ additional_source_paths + +to_dict(include_secrets) Dict + +from_dict(data)$ Config + +include_patterns() Optional~List~ + +exclude_patterns() Optional~List~ + +focus_modules() Optional~List~ + +doc_type() Optional~str~ + +custom_instructions() Optional~str~ + +all_source_paths() List~str~ + +validate_source_paths() void + +is_multi_path_mode() bool + +get_prompt_addition() str + +from_args(args)$ Config + +from_web_job(repo_path, docs_dir)$ Config + +from_cli(...)$ Config + +from_config_manager(manager, repo_path, output_dir)$ Config + } +``` + +## Field Groups + +| Group | Fields | Purpose | +|---|---|---| +| Paths | `repo_path`, `output_dir`, `dependency_graph_dir`, `docs_dir`, `diagrams_dir` | Where source code is read from and where analysis artifacts and docs are written | +| Multi-path support | `additional_source_paths` | Allows analyzing multiple source directories as one unified documentation set | +| Model selection | `main_model`, `cluster_model`, `fallback_model` | Which LLM is used for generation, clustering, and fallback | +| Per-provider credentials | `cluster_api_key`, `main_api_key`, `fallback_api_key` | Required, runtime-only secrets — never serialized by `to_dict()` unless explicitly requested | +| Per-provider connectivity | `*_base_url`, `*_api_version` | Endpoint configuration for each provider | +| Per-provider limits | `*_max_tokens`, `*_max_token_field`, `max_token_per_module`, `max_token_per_leaf_module` | Token budgeting for generation and clustering | +| Per-provider sampling | `*_temperature`, `*_temperature_supported` | Controls determinism/creativity per provider, with a flag for providers that reject custom temperatures | +| Agent customization | `agent_instructions` | Dict-based include/exclude patterns, focus modules, doc type, and custom instructions consumed via properties | + +## Secret Handling: `to_dict` / `from_dict` + +`Config` maintains a frozen set of runtime-only secret field names (`cluster_api_key`, `main_api_key`, `fallback_api_key`). `to_dict()` strips these by default so that serialized configuration (e.g., cached to disk or logged) never leaks API keys; callers must pass `include_secrets=True` to get a fully round-trippable dict. `from_dict()` reconstructs a `Config` by filtering the input to only known dataclass fields, so extra keys are safely ignored, but if secrets were stripped they must be supplied separately or construction raises a `TypeError` (missing required field). + +```mermaid +flowchart TD + A["Config instance"] --> B["to_dict(include_secrets=False)"] + B --> C["Plain dict, no API keys"] + A --> D["to_dict(include_secrets=True)"] + D --> E["Plain dict, includes API keys"] + C --> F["from_dict(data)"] + E --> F + F -->|"missing secrets"| G["TypeError: missing required field"] + F -->|"secrets present"| H["Reconstructed Config"] +``` + +## Agent Instructions Properties + +`agent_instructions` is an optional dict (or an object exposing `to_dict()`, e.g. `AgentInstructions` from [CLI Core](cli-core.md)). Five read-only properties expose its contents without requiring callers to know its internal shape: + +- `include_patterns` — file glob patterns to include in analysis +- `exclude_patterns` — file glob patterns to exclude +- `focus_modules` — module names that should receive more detailed documentation +- `doc_type` — one of `api`, `architecture`, `user-guide`, `developer`, or a free-form string +- `custom_instructions` — free-form additional guidance text + +`get_prompt_addition()` combines `doc_type`, `focus_modules`, and `custom_instructions` into a single prompt-ready string, escaping curly braces in `custom_instructions` (via `escape_format_braces`) so JSON-like content does not break downstream `.format()` calls in the generation pipeline consumed by [Backend Core](backend-core.md). + +```mermaid +flowchart TD + AI["agent_instructions dict"] --> IP["include_patterns"] + AI --> EP["exclude_patterns"] + AI --> FM["focus_modules"] + AI --> DT["doc_type"] + AI --> CI["custom_instructions"] + DT --> GPA["get_prompt_addition()"] + FM --> GPA + CI -->|"escape_format_braces"| GPA + GPA --> Prompt["Combined prompt-addition string"] +``` + +## Multi-Path Source Support + +`additional_source_paths` enables analyzing more than one directory as a single logical repository. `all_source_paths` always returns `repo_path` as the first absolute path, followed by any additional paths. `validate_source_paths()` raises `ValueError`/`OSError` if any path is missing, not a directory, or unreadable. `is_multi_path_mode()` is a simple boolean check used by the analysis pipeline in [Backend Core](backend-core.md) to decide whether to merge multiple source trees. + +```mermaid +flowchart TD + Start["Config.validate_source_paths()"] --> CheckPrimary{{"repo_path exists and is dir?"}} + CheckPrimary -->|"no"| Err1["raise ValueError"] + CheckPrimary -->|"yes"| HasAdditional{{"additional_source_paths set?"}} + HasAdditional -->|"no"| Done["Validation OK single-path mode"] + HasAdditional -->|"yes"| Loop["For each additional path"] + Loop --> CheckExists{{"path exists and is dir?"}} + CheckExists -->|"no"| Err2["raise ValueError"] + CheckExists -->|"yes"| CheckRead{{"path readable?"}} + CheckRead -->|"no"| Err3["raise OSError"] + CheckRead -->|"yes"| Loop + Loop --> Done2["Validation OK multi-path mode"] +``` + +## Construction Paths + +`Config` provides four classmethods for building an instance, each tailored to a different caller in the system. + +```mermaid +flowchart TD + subgraph CLIFlow["CLI Core entry point"] + CM["ConfigManager: persisted JSON plus keyring"] + CM -->|"from_config_manager"| FC1["Config.from_cli(...)"] + end + subgraph WebFlow["Frontend Core entry point"] + BW["BackgroundWorker for web job"] + BW -->|"from_web_job"| FA1["Config.from_args wrapping Namespace"] + end + subgraph EnvFlow["Environment-driven CLI entry point"] + ArgParse["argparse.Namespace from CLI arguments"] + ArgParse -->|"from_args"| FA2["Reads MAIN_MODEL, FALLBACK_MODEL, CLUSTER_API_KEY, MAIN_API_KEY, FALLBACK_API_KEY from environment"] + end + subgraph DirectFlow["Direct parameter entry point"] + Caller["Any caller with explicit parameters"] + Caller -->|"from_cli"| Validate["Validation block: keys, urls, types, ranges, max_token_field enum"] + Validate --> VSP["validate_source_paths()"] + VSP --> Instance["Config instance"] + end + FA1 --> FA2 + FC1 --> Validate +``` + +### `from_args(args)` + +Builds a `Config` purely from environment variables (`MAIN_MODEL`, `CLUSTER_MODEL`, `LLM_BASE_URL`, `FALLBACK_MODEL`, `CLUSTER_API_KEY`, `MAIN_API_KEY`, `FALLBACK_API_KEY`) plus the `repo_path` supplied on `args`. It computes a sanitized repo name for the docs output directory and raises `ValueError` if `FALLBACK_MODEL` or any per-provider API key is missing. This is the lowest-level, environment-driven constructor. + +### `from_web_job(repo_path, docs_dir)` + +A thin wrapper used by [Frontend Core](frontend-core.md)'s background worker. It delegates to `from_args` (wrapping `repo_path` in a synthetic `argparse.Namespace`) and then overrides `docs_dir` with the job-specific output directory, avoiding the need to fabricate a fake CLI namespace at the call site. + +### `from_cli(...)` + +The most comprehensive constructor, accepting every field explicitly (models, keys, base URLs, API versions, token limits, temperatures, max-token field names, `agent_instructions`, `diagrams_dir`, `additional_source_paths`). It performs an extensive validation block before construction: + +- Required, non-empty API keys and base URLs for all three providers +- Type coercion and validation for token limits (`int`) and temperatures (`float`) +- Range checks: token limits must be positive, temperatures must be within `0.0`–`2.0` +- Enum checks: `*_max_token_field` must be `max_tokens` or `max_completion_tokens` + +After construction, it calls `validate_source_paths()` to ensure the repository and any additional paths actually exist and are accessible before returning the instance. + +### `from_config_manager(manager, repo_path, output_dir)` + +Used by [CLI Core](cli-core.md) to bridge its persisted `ConfigManager`/`Configuration` model into a runtime `Config`. It pulls the loaded `Configuration` object and per-provider API keys from the `ConfigManager`, validates that models and keys are present (raising actionable `ValueError`s referencing the `codewiki config set` command), extracts `additional_source_paths` from `agent_instructions` if present, and finally delegates to `from_cli(...)` with all fields populated from the manager. + +```mermaid +sequenceDiagram + participant CLI as "CLI Core (ConfigManager)" + participant Config as "Config.from_config_manager" + participant FromCli as "Config.from_cli" + participant Validate as "validate_source_paths" + + CLI->>Config: from_config_manager(manager, repo_path, output_dir) + Config->>CLI: manager.get_config() + Config->>CLI: get_cluster_api_key / get_main_api_key / get_fallback_api_key + Config->>Config: check models and keys are present + Config->>FromCli: from_cli(repo_path, output_dir, models, keys, urls, tokens, temps) + FromCli->>FromCli: validation block keys urls types ranges enums + FromCli->>Validate: validate_source_paths() + Validate-->>FromCli: OK or raises ValueError or OSError + FromCli-->>Config: Config instance + Config-->>CLI: Config instance +``` + +## Integration with Other Modules + +- **[Backend Core](backend-core.md)** — the `AgentOrchestrator`, `DocumentationGenerator`, and dependency-analyzer components consume a fully constructed `Config` for repo paths, model selection, token limits, and prompt additions (`get_prompt_addition()`), and use `all_source_paths()`/`is_multi_path_mode()` to drive multi-path analysis. +- **[CLI Core](cli-core.md)** — `ConfigManager` and the `Configuration`/`AgentInstructions` models persist user settings to disk; `Config.from_config_manager` bridges that persisted state into the runtime `Config` used for a documentation run. +- **[Frontend Core](frontend-core.md)** — `BackgroundWorker` builds a `Config` per submitted job via `Config.from_web_job`, using job-specific `repo_path` and `docs_dir` values while relying on environment-configured models and keys. + +## Design Rationale + +- **Secrets never leak by default.** The `_RUNTIME_ONLY_SECRET_FIELDS` frozenset and the `include_secrets` flag on `to_dict()` ensure API keys are excluded from any dict representation used for caching, logging, or persistence, unless a caller explicitly opts in for same-process reconstruction. +- **Fail fast, fail clearly.** `from_cli` performs exhaustive validation (presence, type, range, enum) before constructing the object, and `validate_source_paths()` checks filesystem accessibility immediately after — so configuration errors surface with actionable messages before any expensive analysis or LLM calls begin. +- **One shape, many origins.** Regardless of whether a `Config` originates from CLI environment variables, a persisted `ConfigManager` configuration, or a web job submission, all paths converge on the same validated dataclass shape, so the rest of the pipeline ([Backend Core](backend-core.md)) never needs to know which caller produced it. diff --git a/docs/reference/architecture/frontend-core/frontend-core.md b/docs/reference/architecture/frontend-core/frontend-core.md new file mode 100644 index 00000000..3bb2ed7e --- /dev/null +++ b/docs/reference/architecture/frontend-core/frontend-core.md @@ -0,0 +1,257 @@ +# Frontend Core + +## Introduction + +The Frontend Core module implements CodeWiki's **web application layer** — a FastAPI-based service that lets users submit GitHub repository URLs, tracks documentation-generation jobs asynchronously, caches completed results, and serves the generated documentation back to the browser. + +It acts as the bridge between end users (submitting repositories through a web form) and the heavier documentation-generation machinery implemented in [Backend Core](backend-core.md) (specifically `DocumentationGenerator`) and the shared runtime settings in [Config Core](config-core.md) (`Config`). + +Unlike `cli-core` and `backend-core`, this module is not further decomposed into child sub-modules in the module tree — it is a compact, single-layer module. This document therefore covers all of its components directly, without separate sub-module pages. + +## Responsibilities + +- Validate and normalize submitted GitHub repository URLs +- Queue documentation-generation jobs and process them on a background thread +- Cache generated documentation by repository URL to avoid redundant regeneration +- Persist job status and cache metadata to disk so state survives restarts +- Render the web UI (submission form, job list, generated docs viewer) via Jinja2 templates +- Expose HTTP endpoints (via FastAPI route handlers) for submission, status polling, and documentation viewing + +## Architecture Overview + +The module is organized around a simple pipeline: a web request creates or looks up a `JobStatus`, which is queued to the `BackgroundWorker`. The worker clones the repository, invokes the documentation generator, and stores results through the `CacheManager`. All directories, timeouts, and queue sizing are centralized in `WebAppConfig`. + +```mermaid +flowchart TD + User["Browser / API Client"] -->|"submit repo_url"| Routes["WebRoutes"] + Routes -->|"validate URL"| GitProc["GitHubRepoProcessor"] + Routes -->|"check cache"| Cache["CacheManager"] + Routes -->|"enqueue job"| Worker["BackgroundWorker"] + Routes -->|"render HTML"| Templates["StringTemplateLoader / render_template"] + + Worker -->|"clone repository"| GitProc + Worker -->|"build Config.from_web_job"| ConfigCore["Config (config-core)"] + Worker -->|"generate docs"| DocGen["DocumentationGenerator (backend-core)"] + Worker -->|"store result path"| Cache + Worker -->|"persist status"| JobsFile[("jobs.json")] + + Cache -->|"persist index"| CacheFile[("cache_index.json")] + + subgraph models_group["Data Models"] + JobStatus["JobStatus"] + CacheEntry["CacheEntry"] + RepositorySubmission["RepositorySubmission"] + JobStatusResponse["JobStatusResponse"] + end + + Routes --> models_group + Worker --> models_group + Cache --> models_group +``` + +**Cross-module dependencies:** +- [Backend Core](backend-core.md) — `DocumentationGenerator` performs the actual dependency analysis and LLM-driven documentation generation invoked by `BackgroundWorker`. +- [Config Core](config-core.md) — `Config.from_web_job()` builds the runtime configuration (models, API keys, directories) used for each documentation job. + +## Core Components + +### WebAppConfig — Central Settings + +`WebAppConfig` (in `config.py`) is a plain class holding static configuration constants used across the whole module: + +- **Directories**: `CACHE_DIR`, `TEMP_DIR`, `OUTPUT_DIR` +- **Queue settings**: `QUEUE_SIZE` +- **Cache settings**: `CACHE_EXPIRY_DAYS` +- **Job cleanup**: `JOB_CLEANUP_HOURS`, `RETRY_COOLDOWN_MINUTES` +- **Server defaults**: `DEFAULT_HOST`, `DEFAULT_PORT` +- **Git clone settings**: `CLONE_TIMEOUT`, `CLONE_DEPTH` + +It also provides `ensure_directories()` (creates cache/temp/output folders) and `get_absolute_path()`. Every other component in this module reads its defaults from `WebAppConfig` unless overridden by an explicit constructor argument. + +### GitHubRepoProcessor — Repository Validation & Cloning + +`GitHubRepoProcessor` is a stateless utility class (all static methods) responsible for: + +- `is_valid_github_url(url)` — ensures the URL points to `github.com`/`www.github.com` with a valid `owner/repo` path +- `get_repo_info(url)` — extracts `owner`, `repo`, `full_name`, and a normalized `clone_url` +- `clone_repository(clone_url, target_dir, commit_id=None)` — clones via `git clone` (shallow, depth-limited by `WebAppConfig.CLONE_DEPTH`, unless a specific `commit_id` is requested, in which case a full clone + `git checkout` is performed); cleans up the target directory on failure + +This component has no dependency on any other module — it only shells out to `git` and reads settings from `WebAppConfig`. + +### CacheManager — Documentation Cache + +`CacheManager` maintains an on-disk index (`cache_index.json`) mapping a SHA-256 hash of the repository URL (`get_repo_hash`) to a `CacheEntry` describing where the generated docs live and when they were created/last accessed. + +Key behaviors: +- `get_cached_docs(repo_url)` — returns the cached docs path if the entry exists and has not expired (`CACHE_EXPIRY_DAYS`); expired entries are automatically removed +- `add_to_cache(repo_url, docs_path)` — creates/updates a `CacheEntry` and persists the index +- `remove_from_cache(repo_url)` / `cleanup_expired_cache()` — cache invalidation utilities +- Corrupted index files are detected and backed up rather than crashing the app + +### BackgroundWorker — Asynchronous Job Processing + +`BackgroundWorker` is the core orchestration engine of the module. It owns: + +- A bounded `Queue` (`processing_queue`, sized by `WebAppConfig.QUEUE_SIZE`) of job IDs waiting to be processed +- An in-memory `job_status: Dict[str, JobStatus]` map +- A `jobs.json` file for persisting completed job state across restarts + +Lifecycle: +1. `start()` launches a daemon thread running `_worker_loop()`, which polls the queue and dispatches jobs to `_process_job()`. +2. `add_job(job_id, job)` registers a new `JobStatus` and enqueues its ID. +3. `_process_job(job_id)`: + - Checks `CacheManager` first — if valid cached docs exist, marks the job `completed` immediately. + - Otherwise resolves repo info via `GitHubRepoProcessor.get_repo_info()`, clones the repository into a per-job temp directory (`GitHubRepoProcessor.clone_repository`, optionally checking out a specific `commit_id`). + - Builds a `Config` via `Config.from_web_job(repo_path, docs_dir)` (see [Config Core](config-core.md)). + - Instantiates `DocumentationGenerator(config, job.commit_id)` from [Backend Core](backend-core.md) and runs its async `run()` method in a dedicated event loop. + - On success, registers the output path with `CacheManager.add_to_cache()` and marks the job `completed`; on failure, marks it `failed` with an `error_message`. + - Always cleans up the temporary cloned repository directory. + +`load_job_statuses()` / `save_job_statuses()` persist only `completed` jobs to `jobs.json`. If no jobs file exists yet, `_reconstruct_jobs_from_cache()` rebuilds job entries directly from the `CacheManager`'s index for backward compatibility (older deployments that only had cache data). + +```mermaid +sequenceDiagram + participant Browser + participant Routes as "WebRoutes" + participant Worker as "BackgroundWorker" + participant Cache as "CacheManager" + participant Git as "GitHubRepoProcessor" + participant DocGen as "DocumentationGenerator" + + Browser->>Routes: POST / (repo_url, commit_id) + Routes->>Git: is_valid_github_url / get_repo_info + Routes->>Cache: get_cached_docs(repo_url) + alt Cache hit + Cache-->>Routes: docs_path + Routes-->>Browser: Render success message + else Cache miss + Routes->>Worker: add_job(job_id, JobStatus) + Worker-->>Routes: queued + Routes-->>Browser: Render "queued" message + Worker->>Git: clone_repository(clone_url, temp_dir, commit_id) + Worker->>DocGen: run() (async documentation generation) + DocGen-->>Worker: docs_dir populated + Worker->>Cache: add_to_cache(repo_url, docs_path) + Worker->>Worker: save_job_statuses() + end +``` + +### Data Models + +Defined in `models.py`, these dataclasses and Pydantic models flow between the components above: + +| Model | Kind | Purpose | +|---|---|---| +| `RepositorySubmission` | Pydantic `BaseModel` | Validates form input containing a `repo_url: HttpUrl` | +| `JobStatusResponse` | Pydantic `BaseModel` | Shape of the `/api/jobs/{job_id}` JSON response (status, timestamps, error, docs path, model used, commit) | +| `JobStatus` | `dataclass` | In-memory/on-disk representation of a job's lifecycle: `queued` → `processing` → `completed`/`failed` | +| `CacheEntry` | `dataclass` | Cache index entry: repo URL, its hash, docs path, creation and last-access timestamps | + +`JobStatus` and `CacheEntry` are pure data holders serialized manually (via `dataclasses.asdict`/manual dict construction) by `BackgroundWorker` and `CacheManager` respectively — they carry no business logic themselves. + +### WebRoutes — HTTP Route Handlers + +`WebRoutes` implements the FastAPI-facing handlers, wired to a `BackgroundWorker` and `CacheManager` instance: + +- `index_get(request)` — renders the main submission form plus the 100 most recent jobs +- `index_post(request, repo_url, commit_id)` — the primary submission flow: + 1. Cleans up expired jobs (`cleanup_old_jobs`) + 2. Validates the URL via `GitHubRepoProcessor` + 3. Normalizes the URL and derives a URL-safe `job_id` (`owner--repo`) + 4. Checks for an existing in-flight or recently-failed job (respecting `WebAppConfig.RETRY_COOLDOWN_MINUTES`) to prevent duplicate work + 5. Checks the cache; if a hit, synthesizes a `completed` `JobStatus` for immediate display + 6. Otherwise creates a new `queued` `JobStatus` and calls `BackgroundWorker.add_job()` +- `get_job_status(job_id)` — JSON API returning a `JobStatusResponse` +- `view_docs(job_id)` — redirects to the static documentation viewer for a completed job +- `serve_generated_docs(job_id, filename)` — resolves and renders a specific generated Markdown file (with path-traversal protection), falling back to reconstructing job state from the cache if no in-memory job exists; loads `module_tree.json`/`metadata.json` for navigation and converts Markdown to HTML for display +- Helper methods `_normalize_github_url`, `_repo_full_name_to_job_id`, `_job_id_to_repo_full_name`, `cleanup_old_jobs` support the above flows + +### Template Rendering + +`template_utils.py` provides a thin Jinja2 integration layer: + +- `StringTemplateLoader` — a custom `jinja2.BaseLoader` that serves a template directly from a Python string (no filesystem template directory needed), enabling templates to be defined inline as Python string constants elsewhere in the application +- `render_template(template, context)` — configures a `Jinja2` `Environment` (autoescaping HTML/XML, `trim_blocks`/`lstrip_blocks` enabled) and renders the given template string against a context dict +- `render_navigation(module_tree, current_page)` — renders sidebar navigation HTML from a documentation module tree structure +- `render_job_list(jobs)` — renders the recent-jobs list HTML fragment + +`WebRoutes` uses `render_template` to produce every `HTMLResponse` it returns. + +## Component Relationships + +```mermaid +classDiagram + class WebAppConfig { + +CACHE_DIR + +TEMP_DIR + +QUEUE_SIZE + +CACHE_EXPIRY_DAYS + +CLONE_TIMEOUT + +ensure_directories() + } + class GitHubRepoProcessor { + +is_valid_github_url(url) + +get_repo_info(url) + +clone_repository(clone_url, target_dir, commit_id) + } + class CacheManager { + +cache_index + +get_cached_docs(repo_url) + +add_to_cache(repo_url, docs_path) + +remove_from_cache(repo_url) + } + class BackgroundWorker { + +job_status + +processing_queue + +start() + +add_job(job_id, job) + +get_job_status(job_id) + } + class WebRoutes { + +index_get(request) + +index_post(request, repo_url, commit_id) + +get_job_status(job_id) + +serve_generated_docs(job_id, filename) + } + class JobStatus + class CacheEntry + class RepositorySubmission + class JobStatusResponse + class StringTemplateLoader + + WebRoutes --> BackgroundWorker + WebRoutes --> CacheManager + WebRoutes --> GitHubRepoProcessor + WebRoutes --> StringTemplateLoader + BackgroundWorker --> CacheManager + BackgroundWorker --> GitHubRepoProcessor + BackgroundWorker --> JobStatus + CacheManager --> CacheEntry + WebRoutes --> JobStatusResponse + GitHubRepoProcessor --> WebAppConfig + CacheManager --> WebAppConfig + BackgroundWorker --> WebAppConfig +``` + +## Job Lifecycle State Machine + +```mermaid +stateDiagram-v2 + [*] --> queued: "add_job()" + queued --> processing: "worker picks up job" + processing --> completed: "cache hit OR generation succeeds" + processing --> failed: "clone or generation error" + failed --> queued: "resubmission after cooldown" + completed --> [*] + failed --> [*] +``` + +## Integration with the Rest of CodeWiki + +- **Documentation generation**: `BackgroundWorker._process_job()` delegates the actual analysis and Markdown/diagram generation to `DocumentationGenerator` from [Backend Core](backend-core.md). Frontend Core does not implement any dependency analysis itself — it only manages the job lifecycle, caching, and presentation around that generator. +- **Configuration**: Every job builds a fresh runtime `Config` via `Config.from_web_job(repo_path, docs_dir)`, defined in [Config Core](config-core.md). This keeps LLM model selection, API keys, and output directories consistent with the rest of the pipeline while allowing the web app to supply job-specific paths. +- **Independent of CLI**: Unlike [CLI Core](cli-core.md), which drives documentation generation from the command line with its own `ConfigManager` and job models, Frontend Core is a self-contained HTTP-facing alternative entry point that shares the same downstream `DocumentationGenerator` and `Config` but has its own job-tracking (`JobStatus`) and caching (`CacheManager`) implementations tailored for a multi-user web environment (queueing, retry cooldowns, cache expiry). + +## Summary + +Frontend Core provides the web-facing shell around CodeWiki's documentation engine: validating and queueing repository submissions, running generation jobs on a background thread, caching results to avoid repeat work, and rendering both the submission UI and the generated documentation itself. Its five main building blocks — `WebAppConfig`, `GitHubRepoProcessor`, `CacheManager`, `BackgroundWorker`, and `WebRoutes` — form a straightforward pipeline, with `BackgroundWorker` acting as the connective tissue to the heavier [Backend Core](backend-core.md) documentation generator and [Config Core](config-core.md) configuration model. diff --git a/docs/reference/architecture/test-clustering/test-clustering.md b/docs/reference/architecture/test-clustering/test-clustering.md new file mode 100644 index 00000000..849c62b6 --- /dev/null +++ b/docs/reference/architecture/test-clustering/test-clustering.md @@ -0,0 +1,208 @@ +# Test Clustering + +## Purpose + +The Test Clustering module is a collection of standalone diagnostic and validation scripts used to exercise the LLM-driven **module clustering** functionality that lives inside the backend's documentation-generation pipeline (`cluster_modules`). Unlike a conventional `pytest` suite, these scripts are executable Python programs (`python3 script.py`) that print human-readable pass/fail reports to the console and exit with a non-zero status code on failure, making them suitable for quick manual runs, CI smoke checks, or debugging sessions when the clustering behavior of the underlying LLM changes. + +Clustering is the step in the CodeWiki pipeline where a flat list of code components (functions, classes, files) discovered by the [dependency analyzer](backend-core.md) is grouped by an LLM into a hierarchical module tree (e.g. "Auth Module", "API Module") that later becomes the basis for the generated documentation structure. Because this step depends on free-form LLM output, it is especially prone to format drift (e.g., the LLM returning quoted strings or class names instead of integer IDs). The scripts in this module were written to reproduce, isolate, and regression-test these failure modes. + +## Scope and Relationship to Other Modules + +This module does not define new production functionality; instead, it directly imports and drives components from other parts of the system: + +- **`cluster_modules`, `create_component_id_map`, `normalize_component_ids_by_lookup`** — the LLM clustering functions under test, part of the backend's documentation-generation pipeline (see [Backend Core](backend-core.md)). +- **`Node`** — the dependency-graph node model representing a single code component, documented as part of [Backend Core](backend-core/dependency-analyzer-models/dependency-analyzer-models.md). +- **`Config`** — the pipeline configuration object (model names, API keys, base URLs, token thresholds), documented in [Config Core](config-core.md). + +Because the clustering step is LLM-backed, most scripts in this module require valid API credentials (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, or the `MAIN_API_KEY` / `CLUSTER_API_KEY` / `FALLBACK_API_KEY` overrides) to be set in the environment or a local `.env.local` file before they can make live calls. The two validation-focused scripts (`test_clustering_validation.py` and `test_id_based_clustering.py`) are pure logic simulations and do **not** require any network access or API keys. + +## Architecture Overview + +```mermaid +flowchart TD + subgraph scripts["Test Clustering Scripts"] + Debug["test_clustering_debug.py"] + Forced["test_clustering_forced.py"] + Local["test_clustering_local.py"] + Integration["test_clustering_integration.py"] + Validation["test_clustering_validation.py"] + IdBased["test_id_based_clustering.py"] + end + + subgraph backend["Backend Clustering Pipeline"] + ClusterFn["cluster_modules()"] + IdMap["create_component_id_map()"] + Normalize["normalize_component_ids_by_lookup()"] + LLMClient["LLM Client"] + end + + ConfigMod["Config"] + NodeMod["Node"] + + Debug -->|"invokes"| ClusterFn + Forced -->|"invokes"| ClusterFn + Local -->|"invokes"| ClusterFn + ClusterFn -->|"calls"| LLMClient + + Integration -->|"invokes directly"| IdMap + Integration -->|"invokes directly"| Normalize + + Validation -->|"simulates validation logic of"| ClusterFn + IdBased -->|"simulates parsing/validation logic of"| ClusterFn + + Debug -->|"constructs"| NodeMod + Forced -->|"constructs"| NodeMod + Local -->|"constructs"| NodeMod + Debug -->|"constructs"| ConfigMod + Forced -->|"constructs"| ConfigMod + Local -->|"constructs"| ConfigMod +``` + +## The Common `TestResults` Pattern + +Nearly every script in this module defines its own local `TestResults` class rather than importing a shared one. This is a deliberate consequence of these being independent, copy-paste-friendly diagnostic scripts rather than a shared test library. Each implementation follows the same basic contract: + +1. Accumulate `(name, passed, details)` tuples via an `add_test(...)` method. +2. Print a formatted summary via `print_summary()`. +3. Return or expose an overall boolean/exit-code indicating whether all tests passed. + +| Script | `TestResults` Behavior | Exit Code Semantics | +|---|---|---| +| `test_clustering_debug.py` | Simple `add_test(name, passed, details)`; summary prints ✅/❌ per test | `sys.exit(0 if success else 1)` | +| `test_clustering_forced.py` | Same shape, message-only details | `sys.exit(0 if passed else 1)` | +| `test_clustering_local.py` | Same shape; returns `passed_count == total` | `sys.exit(0 if success else 1)` | +| `test_clustering_integration.py` | Tracks explicit `passed` / `failed` counters plus a list of dict-based test records | `main()` returns `0` or `1`, used as process exit code | +| `test_clustering_validation.py` | Tracks `passed` / `failed` counters and a `failures` list; exposes a `success` property | `exit(0 if success else 1)` | +| `test_id_based_clustering.py` | Minimal `(name, passed)` tuple list; `print_summary()` returns `all_passed` | `sys.exit(0)` / `sys.exit(1)` | + +```mermaid +classDiagram + class TestResultsBase { + +add_test(name, passed, details) + +print_summary() bool + } + class DebugResults + class ForcedResults + class LocalResults + class IntegrationResults { + +passed int + +failed int + } + class ValidationResults { + +success bool + } + class IdBasedResults + + TestResultsBase <|-- DebugResults + TestResultsBase <|-- ForcedResults + TestResultsBase <|-- LocalResults + TestResultsBase <|-- IntegrationResults + TestResultsBase <|-- ValidationResults + TestResultsBase <|-- IdBasedResults +``` + +Note: `TestResultsBase` above is a conceptual grouping for documentation purposes only — each script defines its own independent class with no shared base class or import relationship in the actual source code. + +## Script Reference + +### `test_clustering_debug.py` + +Runs a full, live invocation of `cluster_modules` against four hand-crafted `Node` components (`AuthController`, `AuthService`, `UserController`, `UserService`) and monkey-patches the LLM client factory (`create_llm_client`) so that the raw LLM response text can be captured and printed. This is the go-to script when the clustering output looks wrong and you need to see exactly what the LLM returned (including whether it emitted the expected `` tag). + +Key characteristics: +- Loads credentials from `.env.local` via `python-dotenv`. +- Builds a `Config` object with `repo_path` overridable via the `CODEWIKI_TEST_REPO` environment variable (defaults to a local `fixtures/sample_repo` directory). +- Prints up to 2000 characters of the captured LLM response for inspection. +- Reports failure with diagnostic detail when the module tree comes back empty (including whether the `` tag was present in the response). + +### `test_clustering_forced.py` + +Similar in structure to the debug script, but deliberately sets `max_token_per_module=50` on the `Config` to force the clustering logic to invoke the LLM (clustering is normally skipped when the component set is small enough to fit under the token threshold). It also lowers `cluster_max_tokens` to keep the forced call cheap. Ten synthetic `Component{i}` nodes are generated to guarantee the threshold is exceeded. + +This script is useful for confirming that: +- The token-based trigger for invoking the LLM clustering path actually fires. +- The LLM's response still respects the expected `` tag format under a forced, low-budget scenario. + +### `test_clustering_local.py` + +A more guarded, "safe to run locally" variant that requires an existing on-disk repository (`CODEWIKI_TEST_REPO`) and specific Java source files inside it (`AuthController.java`, `AuthService.java`, `UserController.java`, `UserService.java`, under an OpenFrame-style API project layout). If the repo or API keys are missing, the script exits early with a clear message rather than attempting an LLM call. It wraps the `cluster_modules` invocation in a `try`/`except` block and prints a full traceback on unexpected exceptions, in addition to the standard `TestResults` summary. + +### `test_clustering_integration.py` + +The most comprehensive script in this module. Rather than calling `cluster_modules` end-to-end, it exercises two lower-level helper functions directly: +- `create_component_id_map(components)` — builds the `id -> FQDN` lookup and human-readable ID descriptions used to keep LLM prompts compact. +- `normalize_component_ids_by_lookup(module_tree, id_to_fqdn)` — converts LLM-returned component IDs back into fully-qualified component names, filtering out anything invalid. + +It defines a lightweight `MockNode` class (with `fqdn`, `name`, `file_path` attributes) to avoid depending on the real `Node` model, and a `capture_log_warnings` decorator that redirects the `codewiki.src.be.cluster_modules` logger into an in-memory buffer so tests can assert on specific warning strings (e.g., `"Non-integer ID"`, `"Invalid ID 999"`, `"Valid range"`). + +Ten focused test functions cover the following normalization edge cases: + +| Test | Input | Expected Outcome | +|---|---|---| +| `test_component_id_map_creation` | 5 sample components | Sequential integer IDs `0..4` map 1:1 to FQDNs | +| `test_valid_integer_ids` | `[0, 1, 2]` | All normalize cleanly, no warnings | +| `test_invalid_quoted_integers` | `["0", "1", "2"]` | Accepted via `int()` conversion (resilient behavior), no warnings | +| `test_invalid_class_names` | `["AuthService", "CountedGenericQueryResult"]` | All rejected with `"Non-integer ID"` warnings | +| `test_mixed_invalid_ids` | `[0, "1", "AuthService", 999]` | Only `0` and `"1"` (→ 1) normalize; the rest are rejected | +| `test_out_of_range_ids` | `[0, 1, 999]` | `999` rejected with `"Invalid ID 999"` / `"Valid range"` warning | +| `test_json_loads_normalization` | `json.loads("[0, 1, 2]")` | Normalizes cleanly, confirming safe JSON parsing works | +| `test_empty_list` | `[]` | No components, no warnings | +| `test_duplicate_ids` | `[0, 1, 1, 2]` | Duplicates preserved (4 components, FQDN for `1` appears twice) | +| `test_negative_ids` | `[-1, 0, 1]` | `-1` rejected with `"Invalid ID -1"` warning | + +### `test_clustering_validation.py` + +A pure-simulation script that re-implements the validation logic from `cluster_modules.py` (referenced as lines 338–369 in the source comments) inside a local `simulate_validation(response_content, max_id)` function. It parses a raw JSON string with `json.loads` (explicitly avoiding `eval()` for safety) and checks that every component ID in every module is an integer within `[0, max_id]`. It is run against a table of ten hard-coded test cases covering valid bare integers, quoted integers, string class names, mixed types, out-of-range and negative IDs, malformed/trailing-comma JSON, empty component lists, and modules with no `components` key at all — asserting that the pass/fail outcome matches the expected `should_pass` flag for each case. + +### `test_id_based_clustering.py` + +The most isolated script — it has **no dependency on the `codewiki` package at all** and instead re-implements small snippets of the clustering logic inline to document and verify three specific historical bug fixes: + +1. **`json.loads()` instead of `eval()`** for parsing LLM responses (`test_json_parsing`) — demonstrates that quoted string IDs parse but would subsequently fail type validation. +2. **Correct 4-tuple unpacking** of a mocked `format_potential_core_components()` function (`test_return_types`) — guards against a regression where the last tuple element (a `Dict` of ID descriptions) was mistakenly treated as the string passed to a token counter. +3. **Integer ID range validation** (`test_id_validation`) and **ID-to-FQDN normalization** (`test_normalization`) — smaller, self-contained versions of the same checks performed by `test_clustering_integration.py`, useful as a minimal reproduction when debugging without the full backend installed. + +## Running the Scripts + +All scripts are plain Python entry points and can be run directly, for example: + +```bash +python3 test_clustering_debug.py +python3 test_clustering_forced.py +python3 test_clustering_local.py +python3 test_clustering_integration.py +python3 test_clustering_validation.py +python3 test_id_based_clustering.py +``` + +Scripts that perform live LLM calls (`test_clustering_debug.py`, `test_clustering_forced.py`, `test_clustering_local.py`) read credentials and model overrides from environment variables such as `MAIN_MODEL`, `CLUSTER_MODEL`, `FALLBACK_MODEL`, `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `MAIN_API_KEY`, `CLUSTER_API_KEY`, and `FALLBACK_API_KEY`, and optionally load a `.env.local` file via `python-dotenv`. The purely logical scripts (`test_clustering_validation.py`, `test_id_based_clustering.py`) run without any environment configuration. + +## Typical Debugging Flow + +```mermaid +sequenceDiagram + participant Dev as Developer + participant Debug as "test_clustering_debug.py" + participant Patch as "Monkey-patched LLM Client" + participant Cluster as "cluster_modules()" + + Dev->>Debug: python3 test_clustering_debug.py + Debug->>Patch: patch create_llm_client() + Debug->>Cluster: cluster_modules(leaf_nodes, components, config) + Cluster->>Patch: client.call(prompt) + Patch-->>Cluster: raw LLM response text + Patch-->>Debug: capture response (global variable) + Cluster-->>Debug: module_tree dict + Debug->>Dev: print captured response + module_tree summary + Debug->>Dev: TestResults.print_summary() +``` + +## Summary + +The Test Clustering module provides a layered set of confidence checks for the LLM-based clustering step of the CodeWiki pipeline: + +- **End-to-end scripts** (`test_clustering_debug.py`, `test_clustering_forced.py`, `test_clustering_local.py`) exercise the real `cluster_modules` function against live LLM calls, differing mainly in how they trigger the LLM path and how much diagnostic output they surface. +- **Targeted unit-style scripts** (`test_clustering_integration.py`) validate the ID-mapping and normalization helper functions in isolation, covering a wide range of malformed-input edge cases without requiring any LLM call. +- **Pure simulation scripts** (`test_clustering_validation.py`, `test_id_based_clustering.py`) re-implement small pieces of the validation logic to document and regression-test specific historical bugs (unsafe `eval()` usage, tuple-unpacking mistakes, ID range checks) without any dependency on the live backend or network access. + +For details on the underlying components these scripts depend on, see [Backend Core](backend-core.md) (clustering pipeline and dependency graph models) and [Config Core](config-core.md) (pipeline configuration). diff --git a/docs/reference/architecture/test-multi-path/sample_fixtures.md b/docs/reference/architecture/test-multi-path/sample_fixtures.md new file mode 100644 index 00000000..5f477219 --- /dev/null +++ b/docs/reference/architecture/test-multi-path/sample_fixtures.md @@ -0,0 +1,125 @@ +# Sample Fixtures + +The Sample Fixtures module provides a small, self-contained set of illustrative components used to exercise the multi-path analysis capabilities of the dependency analyzer. It models a minimal API-driven application composed of a request controller, a core service, and a pluggable data-processing interface. These fixtures are not part of the production CodeWiki system itself — they exist as test data to validate that the [Backend Core](backend-core.md) dependency analysis and call-graph tooling correctly resolve cross-file and cross-package relationships (controller → service → plugin) across multiple source roots. + +## Purpose and Scope + +Sample Fixtures serves as a synthetic fixture package with three cooperating pieces: + +- **`APIController`** — the entry point that receives API-style requests and routes them to the service layer. +- **`MainService`** — the core business logic component that validates and processes request data, and depends on helper utilities from an external `deps` package. +- **`PluginInterface` / `DataPlugin`** — an extensible plugin abstraction, defined in a separate `external` package, representing a third-party/pluggable extension point in the dependency graph. + +Because the components live in distinct sub-packages (`main.controller`, `main.service`, `external.plugin`), this module is specifically useful for verifying that multi-path/multi-root dependency resolution works correctly — hence its parent module name, [Test Multi Path](test-multi-path.md). + +## Module Position + +Sample Fixtures is a child module of [Test Multi Path](test-multi-path.md), alongside its sibling module [Test Suites](test_suites.md), which contains the test runners (`IntegrationTestRunner`, `TestResults`, `Colors`) that exercise these fixtures. + +```mermaid +graph TD + Parent["Test Multi Path"] --> SampleFixtures["Sample Fixtures"] + Parent --> TestSuites["Test Suites"] + SampleFixtures -.->|"exercised by"| TestSuites +``` + +## Component Architecture + +The three components form a simple layered call chain: a controller delegates to a service, and the service (conceptually) could delegate further to a plugin-based extension point. `DataPlugin` implements the `PluginInterface` contract, demonstrating polymorphic extension. + +```mermaid +classDiagram + class APIController { + +service: MainService + +request_count: int + +__init__() + +handle_request(endpoint, data) dict + } + class MainService { + +name: str + +active: bool + +__init__(name) + +start(config) bool + +process_request(data) dict + +stop() + } + class PluginInterface { + +initialize() bool + +execute(context) dict + } + class DataPlugin { + +name: str + +initialized: bool + +__init__(name) + +initialize() bool + +execute(context) dict + } + APIController --> MainService : "uses" + DataPlugin --|> PluginInterface : "implements" +``` + +### APIController + +`APIController` (defined in `test-multi-path/main/controller.py`) is the request-routing entry point of the fixture application. On construction it instantiates a `MainService` named `"api-service"` and tracks a running `request_count`. + +Its single public method, `handle_request(endpoint, data)`, performs simple string-based routing: + +- `"/process"` — delegates to `MainService.process_request(data)`. +- `"/health"` — returns a status payload including the current `request_count`. +- Any other endpoint — returns an `{"error": "Unknown endpoint"}` response. + +```mermaid +sequenceDiagram + participant Client + participant Controller as APIController + participant Service as MainService + + Client->>Controller: handle_request("/process", data) + Controller->>Controller: request_count += 1 + Controller->>Service: process_request(data) + Service->>Service: check active flag + Service-->>Controller: {"status": "success", "result": processed} + Controller-->>Client: response dict +``` + +### MainService + +`MainService` (defined in `test-multi-path/main/service.py`) encapsulates the core business logic of the fixture application. It is constructed with a `name` and starts in an inactive state (`active = False`). + +Key behaviors: + +- **`start(config)`** — validates the supplied configuration using `validate_input` (imported from an external `deps.helper` module, outside this module's scope) and, if valid, flips `active` to `True`. +- **`process_request(data)`** — raises a `RuntimeError` if the service has not been started; otherwise calls `process_data` (also from `deps.helper`) and wraps the result in a structured response dict containing `status`, `service`, and `result`. +- **`stop()`** — resets `active` to `False`. + +This component demonstrates a cross-package dependency: `MainService` lives in `main.service` but relies on helper functions declared in a separate `deps` package, which is a key scenario the dependency analyzer's multi-path resolution is designed to detect correctly. + +### PluginInterface and DataPlugin + +Defined in `test-multi-path/external/plugin.py`, these two classes model a pluggable extension mechanism located in a distinct `external` package: + +- **`PluginInterface`** is an abstract base defining two contract methods, `initialize()` and `execute(context)`, both of which raise `NotImplementedError` in the base class. +- **`DataPlugin`** is a concrete implementation. It tracks an `initialized` flag, requires `initialize()` to be called before `execute()` can run (otherwise raising `RuntimeError`), and its `execute(context)` method extracts a `"data"` key from the passed `context` dict and returns a formatted result string wrapped in a response dict. + +```mermaid +stateDiagram-v2 + [*] --> Uninitialized: "DataPlugin(name)" + Uninitialized --> Initialized: "initialize()" + Initialized --> Initialized: "execute(context)" + Uninitialized --> Error: "execute() before initialize()" +``` + +## Cross-Module Dependency Flow + +The fixture's package layout — `main.controller`, `main.service`, and `external.plugin` — intentionally spans multiple directories/roots so that the dependency analysis pipeline described in [Backend Core](backend-core.md) can be validated against realistic multi-path import resolution scenarios (e.g., resolving `from service import MainService` and `from deps.helper import process_data, validate_input` across different source roots). + +```mermaid +graph LR + Controller["main.controller.APIController"] -->|"imports"| Service["main.service.MainService"] + Service -->|"imports"| Deps["deps.helper (external)"] + Plugin["external.plugin.DataPlugin"] -->|"implements"| Iface["external.plugin.PluginInterface"] +``` + +## Usage in Testing + +Sample Fixtures components are consumed by the sibling [Test Suites](test_suites.md) module, whose `IntegrationTestRunner` exercises the controller-to-service call chain and validates that analysis output (call graphs, dependency edges) matches expectations for this intentionally cross-package layout. diff --git a/docs/reference/architecture/test-multi-path/test-multi-path.md b/docs/reference/architecture/test-multi-path/test-multi-path.md new file mode 100644 index 00000000..72143a16 --- /dev/null +++ b/docs/reference/architecture/test-multi-path/test-multi-path.md @@ -0,0 +1,87 @@ +# Test Multi Path + +## Purpose + +The Test Multi Path module is a self-contained **test fixture and validation suite** used to verify the multi-path source analysis capability of CodeWiki's dependency analysis pipeline. It does not implement production application logic; instead, it provides: + +1. **Sample application code** (a small "service + controller + plugin" codebase spread across multiple simulated source roots: `main/`, `deps/`, `external/`, and `vendor/`) that mimics a real repository with cross-directory imports. +2. **Test runner scripts** that configure the [Config](config-core.md) object with `additional_source_paths`, invoke `DependencyGraphBuilder` from the [Backend Core](backend-core.md) module, and assert that components are correctly discovered, namespaced, and (where applicable) linked across path boundaries. + +This module exists to answer a specific engineering question for CodeWiki itself: *"When a repository's source code is split across multiple root directories (e.g., a monorepo with `main`, `deps`, and `vendor` folders), can the dependency analyzer correctly parse, namespace, and graph components from all of them without ID collisions or missed dependencies?"* + +## Architecture Overview + +The module has two cooperating halves: **fixtures** (sample code to be analyzed) and **suites** (scripts that drive the analysis and assert on results). + +```mermaid +flowchart TD + subgraph fixtures["Sample Application Fixtures"] + Service["MainService"] + Controller["APIController"] + Plugin["DataPlugin / PluginInterface"] + end + + subgraph suites["Test Suites and Runners"] + MultiPathSuite["test_multi_path.py suite"] + IntegrationRunner["IntegrationTestRunner"] + end + + subgraph external_deps["External Dependencies"] + ConfigCls["Config"] + Builder["DependencyGraphBuilder"] + end + + MultiPathSuite -->|"constructs"| ConfigCls + IntegrationRunner -->|"constructs"| ConfigCls + ConfigCls -->|"declares additional_source_paths"| Builder + Builder -->|"parses"| Service + Builder -->|"parses"| Controller + Builder -->|"parses"| Plugin + MultiPathSuite -->|"asserts on"| Builder + IntegrationRunner -->|"asserts on"| Builder +``` + +- **Config** and **DependencyGraphBuilder** are core components of the [Config Core](config-core.md) and [Backend Core](backend-core.md) modules respectively; Test Multi Path exercises them but does not own them. +- The fixtures (`main/service.py`, `main/controller.py`, `external/plugin.py`) are ordinary Python source files that stand in for a target repository being documented. +- The suites (`test_multi_path.py`, `integration_test.py`) are executable scripts (not pytest-based) that print colorized pass/fail output and return process exit codes. + +## Sub-modules + +### [Sample Fixtures](sample_fixtures.md) + +Contains the sample application code used as analysis input: `MainService`, `APIController`, `DataPlugin`, and `PluginInterface`. These classes simulate a layered application (controller → service → helper) spread across the `main/` and `external/` source roots so that the analyzer has realistic cross-file and cross-path relationships to discover. + +### [Test Suites](test_suites.md) + +Contains the executable validation logic: the `test_multi_path.py` suite (`Colors`, `TestResults`) which runs eight discrete scenario checks (single path, multiple paths, namespacing, cross-path dependencies, invalid paths, empty paths, relative vs. absolute paths), and `integration_test.py` (`IntegrationTestRunner`, `TestResults`) which runs a single comprehensive end-to-end scenario against a freshly generated temporary repository with `main/`, `deps/`, and `vendor/` roots. + +## How the Suites Use the Analysis Pipeline + +Both test scripts follow the same general pattern when invoking the dependency analysis pipeline documented in [Backend Core](backend-core.md): + +```mermaid +sequenceDiagram + participant Suite as "Test Script" + participant Cfg as "Config" + participant Builder as "DependencyGraphBuilder" + + Suite->>Cfg: Create Config with repo_path and additional_source_paths + Suite->>Cfg: validate_source_paths() (optional, for invalid-path test) + Suite->>Builder: DependencyGraphBuilder(config) + Suite->>Builder: build_dependency_graph() + Builder-->>Suite: Returns (components, leaf_nodes) + Suite->>Suite: Assert namespaces, counts, and cross-path edges + Suite->>Suite: Print summary and exit code +``` + +Key points validated by these scripts: + +- **Namespacing:** component IDs are prefixed by their top-level source root (for example `main.service.MainService`) so that same-named classes in different roots (e.g., `main/service.py` vs. `deps/helper.py`) never collide. +- **Multi-path detection:** `Config.is_multi_path_mode()` correctly reports whether more than one source root was configured. +- **Path validation:** `Config.validate_source_paths()` raises an error when an additional path does not exist on disk. +- **Cross-path dependency behavior:** at the time these tests were written, import statements crossing source-root boundaries are parsed but not yet resolved into graph edges — the suites document this as expected/known behavior rather than a failure. + +## Relationship to Other Modules + +- **[Config Core](config-core.md):** Test Multi Path constructs `Config` instances (including the `additional_source_paths` and `validate_source_paths` behavior) to drive the scenarios under test. +- **[Backend Core](backend-core.md):** Test Multi Path exercises `DependencyGraphBuilder`, which in turn relies on the dependency-analyzer pipeline (AST parsing, node/graph models) documented in that module. diff --git a/docs/reference/architecture/test-multi-path/test_suites.md b/docs/reference/architecture/test-multi-path/test_suites.md new file mode 100644 index 00000000..a2bcae90 --- /dev/null +++ b/docs/reference/architecture/test-multi-path/test_suites.md @@ -0,0 +1,261 @@ +# Test Suites + +## Introduction + +The Test Suites module contains the executable test scripts that validate CodeWiki's **multi-path dependency analysis** feature — the ability to analyze source code spread across several independent root directories (e.g. `main/`, `deps/`, `vendor/`) as a single logical repository while keeping component identifiers correctly namespaced and dependency edges correctly resolved. + +This module provides two complementary test scripts: + +- **`test_multi_path.py`** — a scenario-based smoke-test suite that exercises the `DependencyGraphBuilder` (see [Backend Core](backend-core.md)) with a series of independent, self-contained checks (single path, multiple paths, namespacing, cross-path dependencies, invalid paths, empty paths, relative vs. absolute paths). +- **`integration_test.py`** — a full end-to-end integration test that programmatically builds a temporary multi-directory sample repository, runs the complete analysis pipeline against it, and asserts on namespace counts, cross-namespace dependency detection, and overall component totals. + +Both scripts are executable as standalone CLI programs (`python test_multi_path.py`, `python integration_test.py`) and return a process exit code suitable for use in CI pipelines. + +This module is a child of the [Test Multi Path](test-multi-path.md) module and is a sibling of the [Sample Fixtures](sample_fixtures.md) module, which supplies some of the on-disk fixture files (`main/controller.py`, `main/service.py`, `external/plugin.py`) referenced by `test_multi_path.py`. + +## Module Purpose in the System + +Test Suites does not implement product functionality — it is a verification harness for the dependency analysis subsystem documented in [Backend Core](backend-core.md), specifically its multi-path graph construction logic. It exercises: + +- **`Config`** (from the [Config Core](config-core.md) module) — used to declare a primary `repo_path` plus a list of `additional_source_paths`. +- **`DependencyGraphBuilder`** (part of the dependency analyzer pipeline in [Backend Core](backend-core.md)) — the component under test, responsible for parsing all configured paths and producing a namespaced component graph. + +Because these scripts drive real analysis runs, they act as a living specification for how multi-path namespacing and cross-path dependency resolution are expected to behave, including documenting known limitations (e.g. cross-namespace dependency resolution is not yet implemented at the AST-parsing level). + +## Core Components + +| Component | File | Responsibility | +|---|---|---| +| `Colors` | `test_multi_path.py` | ANSI color code constants used for terminal output formatting | +| `TestResults` (scenario suite) | `test_multi_path.py` | Accumulates pass/fail results for the seven scenario tests and prints a summary | +| `IntegrationTestRunner` | `integration_test.py` | Orchestrates the full end-to-end integration test: environment setup, config creation, pipeline execution, and multi-stage validation | +| `TestResults` (integration suite) | `integration_test.py` | Accumulates detailed assertion results (with optional `details` text) for the integration run and prints a validation summary | + +> **Note:** Both scripts define a class named `TestResults`, but they are distinct, independent implementations local to each file — there is no shared base class or import relationship between them. + +### Colors + +A simple constants class holding ANSI escape codes (`GREEN`, `RED`, `YELLOW`, `BLUE`, `BOLD`, `END`) used by the module-level print helpers (`print_header`, `print_success`, `print_error`, `print_warning`, `print_info`) in `test_multi_path.py` to produce readable, color-coded console output. + +### TestResults (test_multi_path.py) + +```text +class TestResults: + def __init__(self): + self.results = {} # test_name -> bool + + def add_test(self, name, passed) -> None + def print_summary(self) -> int # returns 0 (all passed) or 1 (failures) +``` + +Used by `run_all_tests()` as a simple dictionary-backed accumulator. Each of the seven scenario test functions returns a boolean, which is recorded under a descriptive test name. `print_summary()` renders a pass/fail report and computes the suite's overall exit code. + +### IntegrationTestRunner + +`IntegrationTestRunner` is a stateful orchestrator class that owns the entire lifecycle of the integration test: + +```text +class IntegrationTestRunner: + def __init__(self): + self.test_dir # temp directory root + self.main_path # main/ source path + self.deps_path # deps/ source path + self.vendor_path # vendor/ source path + self.config # Config instance + self.builder # DependencyGraphBuilder instance + self.components # Dict[str, Node] result + self.leaf_nodes # leaf node result + self.results # TestResults instance +``` + +Its `run()` method executes the following ordered steps, each implemented as a dedicated method: + +1. `setup_test_environment()` — creates a temporary directory with three sub-paths (`main/`, `deps/`, `vendor/`) and populates each with hand-written Python fixture files (`service.py`, `api.py`, `models.py`, `controller.py`, `utils.py` in `main/`; `helper.py`, `validator.py`, `cache.py` in `deps/`; `logger.py`, `metrics.py` in `vendor/`) that intentionally cross-import between paths. +2. `create_config()` — builds a `Config` (see [Config Core](config-core.md)) with `repo_path` set to `main/` and `additional_source_paths` set to `[deps/, vendor/]`. +3. `validate_paths()` — confirms the root and all additional paths exist on disk. +4. `execute_dependency_parser()` — instantiates `DependencyGraphBuilder` (see [Backend Core](backend-core.md)), verifies `config.is_multi_path_mode()` returns `True`, and calls `build_dependency_graph()`. +5. `verify_namespaces()` — groups discovered component IDs by their top-level namespace (`main`, `deps`, `vendor`) and checks expected per-namespace component counts. +6. `verify_cross_path_dependencies()` — inspects component `dependencies` for edges that cross namespace boundaries. +7. `verify_no_warnings()` — checks the builder for any recorded warnings. +8. `verify_file_counts()` — checks the total component count against the expected total (14). +9. `print_detailed_output()` — dumps all component IDs, cross-namespace dependencies, and per-namespace counts for manual inspection. +10. `cleanup()` — removes the temporary test directory (always executed via `finally`). + +### TestResults (integration_test.py) + +A richer accumulator than the scenario-suite version — each recorded test entry carries an optional `details` string, allowing informational warnings to be attached to otherwise-passing assertions: + +```text +class TestResults: + def __init__(self): + self.tests = [] # list of {"name", "passed", "details"} + + def add_test(self, name, passed, details="") -> None + def print_summary(self) -> None # prints PASS/FAIL + warnings block + def all_passed(self) -> bool +``` + +`IntegrationTestRunner.run()` uses `all_passed()` to determine the final process exit code (`0` on success, `1` on any failed assertion, `2` if the test crashes with an unhandled exception). + +## Architecture + +### Class Relationships + +```mermaid +classDiagram + class Colors { + +GREEN + +RED + +YELLOW + +BLUE + +BOLD + +END + } + class ScenarioTestResults { + -results dict + +add_test(name, passed) + +print_summary() int + } + class IntegrationTestRunner { + -test_dir + -main_path + -deps_path + -vendor_path + -config + -builder + -components + -leaf_nodes + -results + +setup_test_environment() + +create_config() + +validate_paths() + +execute_dependency_parser() + +verify_namespaces() + +verify_cross_path_dependencies() + +verify_no_warnings() + +verify_file_counts() + +print_detailed_output() + +cleanup() + +run() int + } + class IntegrationTestResults { + -tests list + +add_test(name, passed, details) + +print_summary() + +all_passed() bool + } + IntegrationTestRunner --> IntegrationTestResults : records into + IntegrationTestRunner --> Config : creates + IntegrationTestRunner --> DependencyGraphBuilder : invokes + ScenarioTestResults ..> Colors : uses for output +``` + +### Scenario Test Suite Flow (`test_multi_path.py`) + +```mermaid +flowchart TD + Start["run_all_tests()"] --> T1["test_single_path()"] + Start --> T2["test_multiple_paths()"] + Start --> T3["test_component_namespacing()"] + Start --> T4["test_cross_path_dependencies()"] + Start --> T5["test_invalid_path_handling()"] + Start --> T6["test_empty_additional_paths()"] + Start --> T7["test_relative_vs_absolute_paths()"] + T1 --> Builder["DependencyGraphBuilder.build_dependency_graph()"] + T2 --> Builder + T3 --> Builder + T4 --> Builder + T6 --> Builder + T7 --> Builder + T5 --> Validate["Config.validate_source_paths()"] + Builder --> Results["TestResults.add_test()"] + Validate --> Results + Results --> Summary["TestResults.print_summary()"] + Summary --> ExitCode["Process exit code (0 or 1)"] +``` + +Each scenario test constructs its own `Config` (see [Config Core](config-core.md)) via the local `create_test_config()` helper, pointing `repo_path` at the `main/` fixture directory and, where relevant, supplying `additional_source_paths` pointing at `deps/` and/or `external/` fixture directories (see [Sample Fixtures](sample_fixtures.md)). + +### Integration Test Pipeline (`integration_test.py`) + +```mermaid +sequenceDiagram + participant Main as "main()" + participant Runner as "IntegrationTestRunner" + participant Cfg as "Config" + participant Builder as "DependencyGraphBuilder" + participant Results as "IntegrationTestResults" + + Main->>Runner: run() + Runner->>Runner: setup_test_environment() + Note over Runner: Creates main/, deps/, vendor/ with cross-importing fixture files + Runner->>Cfg: create_config() + Cfg-->>Runner: Config(repo_path, additional_source_paths) + Runner->>Runner: validate_paths() + Runner->>Builder: execute_dependency_parser() + Builder->>Builder: build_dependency_graph() + Builder-->>Runner: components, leaf_nodes + Runner->>Results: verify_namespaces() + Runner->>Results: verify_cross_path_dependencies() + Runner->>Results: verify_no_warnings() + Runner->>Results: verify_file_counts() + Runner->>Runner: print_detailed_output() + Runner->>Results: print_summary() + Results-->>Runner: all_passed() + Runner->>Runner: cleanup() + Runner-->>Main: exit code 0, 1, or 2 +``` + +### Dependency on Backend Core + +```mermaid +flowchart LR + subgraph TS["Test Suites"] + SM["Scenario Test Suite (test_multi_path.py)"] + IT["Integration Test Runner (integration_test.py)"] + end + subgraph CC["Config Core"] + Config["Config"] + end + subgraph DA["Dependency Analyzer Core"] + DGB["DependencyGraphBuilder"] + end + SM --> Config + IT --> Config + SM --> DGB + IT --> DGB +``` + +Both scripts treat `DependencyGraphBuilder` (see [Backend Core](backend-core.md)) as the system under test and `Config` (see [Config Core](config-core.md)) as the primary input mechanism for expressing multi-path analysis intent via `additional_source_paths` and `is_multi_path_mode()`. + +## Test Coverage Summary + +### Scenario Test Suite (`test_multi_path.py`) + +| # | Test | Validates | +|---|---|---| +| 1 | Single Path (Backward Compatibility) | Analyzing one `repo_path` with no additional paths still discovers expected components | +| 2 | Multiple Paths With Unique Components | Components from `main/`, `deps/`, and `external/` are all discovered when passed as `additional_source_paths` | +| 3 | Component ID Namespacing | No duplicate component IDs occur across paths; IDs reflect their source path | +| 4 | Cross-Path Dependencies | Dependency edges are built between components across different source paths | +| 5 | Invalid Path Handling | `Config.validate_source_paths()` raises `ValueError`/`OSError` for a nonexistent additional path | +| 6 | Empty Additional Paths | An empty `additional_source_paths` list behaves identically to single-path mode | +| 7 | Relative vs. Absolute Paths | Absolute additional paths resolve correctly (relative-path resolution is noted as needing further logic) | + +### Integration Test (`integration_test.py`) + +| Step | Validates | +|---|---| +| Path validation | Root and all additional paths exist on disk before analysis begins | +| Multi-path mode detection | `config.is_multi_path_mode()` returns `True` when additional paths are configured | +| Namespace presence | All expected top-level namespaces (`main`, `deps`, `vendor`) appear in the component graph | +| Namespace component counts | Each namespace yields the expected number of class/function-level components (`main`: 7, `deps`: 5, `vendor`: 2 — 14 total) | +| Cross-namespace dependencies | Dependency edges spanning namespaces are detected where present; documents current limitation that import-level cross-path resolution is not yet fully implemented | +| Warning tracking | No unexpected warnings are recorded by the builder during analysis | + +## Related Modules + +- [Test Multi Path](test-multi-path.md) — parent module; defines the overall multi-path test harness this module belongs to. +- [Sample Fixtures](sample_fixtures.md) — sibling module providing on-disk sample components (`MainService`, `APIController`, `DataPlugin`, `PluginInterface`) referenced by the scenario test suite's `main/` and `external/` fixture directories. +- [Config Core](config-core.md) — supplies the `Config` class used to declare `repo_path` and `additional_source_paths` for every test scenario. +- [Backend Core](backend-core.md) — houses the dependency analyzer pipeline, including the `DependencyGraphBuilder` under test.