Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,18 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
directly instead of going through `python -m clgraph.mcp`.
- `clgraph.toml` / `[tool.clgraph]` project configuration, with keys
`sql_dir` and `dialect`.
- `clgraph detect` - proposes which directory holds your SQL and which
dialect it is, with the evidence and a confidence rating attached.
Evidence is ranked: a dbt profile's adapter `type:` (high confidence),
then dialect-specific syntax markers, then parse scoring. It writes
nothing and decides nothing. `--json` for machine-readable output.
- `clgraph init` - records the answer in `clgraph.toml`, or in
`[tool.clgraph]` with `--into-pyproject`, and adds `.clgraph/` to
`.gitignore`. Prompts when run bare, using detection to pre-fill; with
`--yes` it never prompts and fails instead, which is how agents and CI
invoke it. A missing dialect is an error, never a guess.
- `clgraph.detect` and `clgraph.config_writer` public API: `detect()`,
`Detection`, `Confidence`, `write_config()`, `known_dialects()`.

### Changed

Expand All @@ -36,6 +48,16 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
in clients as a failed server with no explanation.
- `--pipeline` is now optional for `clgraph mcp` and `clgraph-mcp`.

### Fixed

- Dialect parse scoring counted only raised `ParseError`s, but sqlglot
more often degrades unsupported syntax into an `exp.Command` node
without erroring. Both now count as a failed parse, so scoring can
actually distinguish dialects.
- sqlglot warnings emitted while probing candidate dialects no longer
reach stdout, where they corrupted `clgraph detect --json` for any
caller parsing it.

## [0.0.8] - 2026-08-11

### Added
Expand Down
41 changes: 38 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -775,8 +775,43 @@ pip install 'clgraph[mcp]'

#### Project Configuration

Run `clgraph init` and it will find your SQL, propose a dialect with its
reasoning, and write the config:

```bash
clgraph init
```

To see what it would propose without writing anything:

```bash
clgraph detect
```

```
SQL directories:
models (34 files, dbt)

Dialect: snowflake (confidence: high)
- profiles.yml declares adapter type 'snowflake'
```

Confidence is the part to read. `high` comes from a dbt profile's own
adapter type and can be accepted as-is; anything lower is a real question,
because **the dialect is never guessed** — sqlglot parses most of a corpus
under the wrong grammar without complaining, and the resulting lineage
graph is plausible and wrong.

For scripted or CI use, pass everything explicitly and it will never
prompt:

```bash
clgraph init --sql-dir models/ --dialect snowflake --yes
```

The server works out what to index from the project itself, so hosts can
launch it with no arguments. Put this in `clgraph.toml` at your project root:
launch it with no arguments. `clgraph init` writes this for you, or write
it by hand at your project root:

<!-- skip-test -->
```toml
Expand All @@ -799,8 +834,8 @@ what to do instead of showing a dead connection.

**The dialect is never guessed.** Pointing clgraph at SQL files without
saying which dialect they are is an error, not a silent fall back to
BigQuery — sqlglot parses most of a corpus under the wrong grammar without
complaining, and the resulting lineage graph is plausible and wrong.
BigQuery. Run `clgraph init` to be asked once and have the answer
recorded.

#### Claude Desktop Configuration

Expand Down
172 changes: 172 additions & 0 deletions src/clgraph/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -255,6 +255,178 @@ def _print_diff_summary(diff_result):
console.print(f" [yellow]~ {cd.full_name} ({cd.field_name})[/yellow]")


@app.command()
def detect(
path: Annotated[
Path,
typer.Argument(help="Project directory to inspect"),
] = Path("."),
as_json: Annotated[
bool,
typer.Option("--json", help="Emit machine-readable JSON"),
] = False,
):
"""Propose which SQL to index and which dialect it is.

Reports evidence and a confidence level; writes nothing and decides
nothing. `clgraph init` records the answer.

Confidence is what matters: "high" comes from a dbt profile's own
adapter type and can be accepted as-is. Anything lower is a genuine
question, because the wrong dialect produces a lineage graph that
looks right and is not.
"""
from clgraph.detect import detect as run_detection

if not path.exists():
typer.echo(f"Error: path does not exist: {path}", err=True)
raise typer.Exit(code=1)

result = run_detection(path)

if as_json:
typer.echo(json.dumps(result.to_dict(), indent=2))
return

_print_detection(result)


def _print_detection(result) -> None:
"""Render a Detection for a human reader."""
if not result.sql_dirs:
typer.echo("No SQL files found.")
else:
typer.echo("SQL directories:")
for candidate in result.sql_dirs:
relative = candidate.to_dict(result.root)["path"]
typer.echo(f" {relative} ({candidate.file_count} files, {candidate.kind})")

guess = result.dialect
typer.echo("")
if guess.best is None:
typer.echo("Dialect: no evidence found.")
else:
typer.echo(f"Dialect: {guess.best} (confidence: {guess.confidence})")
for reason in guess.evidence:
typer.echo(f" - {reason}")

if guess.confidence != "high":
typer.echo("")
typer.echo("Confirm the dialect before indexing; clgraph will not guess it.")


@app.command()
def init(
sql_dir: Annotated[
Optional[str],
typer.Option("--sql-dir", help="Directory of SQL files, or a JSON pipeline"),
] = None,
dialect: Annotated[
Optional[str],
typer.Option(help="SQL dialect. Required for SQL files; never inferred."),
] = None,
project_dir: Annotated[
Path,
typer.Option("--project-dir", help="Project root to configure"),
] = Path("."),
into_pyproject: Annotated[
bool,
typer.Option("--into-pyproject", help="Write [tool.clgraph] instead of clgraph.toml"),
] = False,
force: Annotated[
bool,
typer.Option("--force", help="Overwrite existing clgraph configuration"),
] = False,
yes: Annotated[
bool,
typer.Option("--yes", "-y", help="Never prompt; fail instead of asking"),
] = False,
):
"""Configure this project for clgraph.

Run bare, it detects candidates and asks. With --yes it never prompts,
which is how agents and CI invoke it: everything must be passed
explicitly, and a missing dialect is an error rather than a guess.
"""
from clgraph.config_writer import write_config
from clgraph.detect import detect as run_detection

detection = run_detection(project_dir)

resolved_sql_dir = sql_dir or _choose_sql_dir(detection, yes)
resolved_dialect = dialect or _choose_dialect(detection, resolved_sql_dir, yes)

try:
result = write_config(
project_dir,
sql_dir=resolved_sql_dir,
dialect=resolved_dialect,
into_pyproject=into_pyproject,
force=force,
)
except (ValueError, FileExistsError) as err:
typer.echo(f"Error: {err}", err=True)
raise typer.Exit(code=1) from err

typer.echo(f"Wrote {result.describe()}")
if result.gitignore_updated:
typer.echo("Added .clgraph/ to .gitignore")
typer.echo("Next: clgraph index")


def _choose_sql_dir(detection, yes: bool) -> str:
"""Pick the SQL directory, prompting only when allowed."""
candidates = detection.sql_dirs

if yes:
if not candidates:
typer.echo(
"Error: no SQL directory found and --sql-dir was not given.",
err=True,
)
raise typer.Exit(code=1)
return candidates[0].to_dict(detection.root)["path"]

if not candidates:
return typer.prompt("Where are your SQL files?")

default = candidates[0].to_dict(detection.root)["path"]
typer.echo("SQL directories found:")
for candidate in candidates:
relative = candidate.to_dict(detection.root)["path"]
typer.echo(f" {relative} ({candidate.file_count} files, {candidate.kind})")
return typer.prompt("Which directory holds your SQL?", default=default)


def _choose_dialect(detection, sql_dir: str, yes: bool) -> Optional[str]:
"""Pick the dialect, prompting only when allowed.

A JSON pipeline carries its own dialect, so it is never asked about.
"""
if sql_dir.endswith(".json"):
return None

guess = detection.dialect

if yes:
typer.echo(
"Error: --dialect is required with --yes. clgraph does not guess: "
"the wrong dialect yields a plausible but incorrect lineage graph.",
err=True,
)
if guess.best:
typer.echo(f"Detection suggests: {guess.best} ({guess.confidence})", err=True)
raise typer.Exit(code=1)

if guess.best:
typer.echo(f"Detected dialect: {guess.best} (confidence: {guess.confidence})")
for reason in guess.evidence:
typer.echo(f" - {reason}")
return typer.prompt("Which SQL dialect?", default=guess.best)

return typer.prompt("Which SQL dialect?")


@app.command()
def mcp(
pipeline: Annotated[
Expand Down
Loading