Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

AzLemma

A lightweight Azerbaijani lemmatizer backed by a compact SQLite database.

AzLemma performs deterministic dictionary-based lemmatization without loading a neural model or requiring external Python packages. It supports strict lookup, optional context-aware POS disambiguation, text reconstruction, JSON output, file processing, and interactive use.

Features

  • Fast SQLite-based word lookup
  • Azerbaijani-aware Unicode normalization and lowercasing
  • Lemmatization of nouns, proper nouns, adjectives, verbs, participles, verbal nouns, infinitives, and converbs
  • Strict ambiguity preservation by default
  • Optional contextual POS disambiguation
  • Batch lookup for separate words
  • Text tokenization with punctuation preservation
  • Lemma-based text reconstruction
  • JSON output
  • UTF-8 file processing
  • Interactive mode
  • No third-party Python dependencies

How it works

The runtime database contains mappings from an observed word form to one or more lemma and part-of-speech candidates:

surface_norm -> lemma, POS, source

For example:

yaratdığı -> yaratmaq, V, moraz
dəyişir   -> dəyişmək, V, moraz
sürətlə   -> sürət, N, moraz
sürətlə   -> sürətlə, ADV, identity

When all candidates point to the same lemma, AzLemma resolves the word directly. When several different lemmas are possible, strict mode reports the word as ambiguous.

Contextual mode uses generic coarse-POS transition scores and a Viterbi-style sequence algorithm. It does not contain word-specific exceptions. A candidate is selected only when its score exceeds the second-best candidate by the configured confidence margin.

Requirements

  • Python 3.10 or newer
  • The azlemma.sqlite runtime database

No pip install step is required.

Installation

Clone the repository:

git clone https://github.com/vrashad/azlemma.git
cd azlemma

Download the database

The SQLite database is too large for normal Git storage, so it is distributed as a GitHub Release asset.

Release page:

https://github.com/vrashad/azlemma/releases/tag/v1.0.0

Create the local data directory:

New-Item -ItemType Directory -Force .\data

Download the database using GitHub CLI:

gh release download v1.0.0 `
  --repo vrashad/azlemma `
  --pattern "azlemma.sqlite" `
  --dir .\data

Or download it directly with PowerShell:

Invoke-WebRequest `
  -Uri "https://github.com/vrashad/azlemma/releases/download/v1.0.0/azlemma.sqlite" `
  -OutFile ".\data\azlemma.sqlite"

You can also download azlemma.sqlite manually from the release page and place it inside the data directory.

Verify that the file exists:

Get-Item .\data\azlemma.sqlite

Project structure

After downloading the database, the project should look like this:

azlemma/
├── azlemma.py
├── README.md
├── .gitignore
└── data/
    └── azlemma.sqlite

Quick start

python azlemma.py text "Allahın yaratdığı dünya sürətlə dəyişir."

Example output:

Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
dünya -> dünya [N; moraz]
sürətlə -> AMBIGUOUS: sürət/N, sürətlə/ADV
dəyişir -> dəyişmək [V; moraz]

Enable contextual disambiguation:

python azlemma.py text `
  --contextual `
  "Allahın yaratdığı dünya sürətlə dəyişir."

Example output:

Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
dünya -> dünya [N; moraz]
sürətlə -> sürətlə [ADV; identity; CONTEXTUAL]
dəyişir -> dəyişmək [V; moraz]

Database location

By default, AzLemma looks for:

data/azlemma.sqlite

You can provide another path with the global --db option:

python azlemma.py --db C:\data\azlemma.sqlite words yaratdığı

The global option must appear before the subcommand.

Usage

Look up separate words

python azlemma.py words Allahın yaratdığı yaradaraq dəyişir

Example output:

Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
yaradaraq -> yaratmaq [V; moraz]
dəyişir -> dəyişmək [V; moraz]

Separate-word lookup never uses contextual disambiguation:

python azlemma.py words sürətlə
sürətlə -> AMBIGUOUS: sürət/N, sürətlə/ADV

Analyze text in strict mode

python azlemma.py text "Allahın yaratdığı dünya sürətlə dəyişir."
Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
dünya -> dünya [N; moraz]
sürətlə -> AMBIGUOUS: sürət/N, sürətlə/ADV
dəyişir -> dəyişmək [V; moraz]

Strict mode never forces a choice between different lemmas.

Analyze text with contextual disambiguation

python azlemma.py text `
  --contextual `
  "Allahın yaratdığı dünya sürətlə dəyişir."
Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
dünya -> dünya [N; moraz]
sürətlə -> sürətlə [ADV; identity; CONTEXTUAL]
dəyişir -> dəyişmək [V; moraz]

The contextual algorithm prefers the ADV candidate because the coarse POS sequence around the word supports an adverb before a verb.

Change the contextual confidence threshold

python azlemma.py text `
  --contextual `
  --context-margin 2.5 `
  "Allahın yaratdığı dünya sürətlə dəyişir."

The default margin is 2.0.

A higher value makes the algorithm more conservative. When the score gap is too small, the result remains ambiguous.

Reconstruct text with lemmas

Strict reconstruction:

python azlemma.py text `
  --reconstruct `
  "Allahın yaratdığı dünya sürətlə dəyişir."

Contextual reconstruction:

python azlemma.py text `
  --contextual `
  --reconstruct `
  "Allahın yaratdığı dünya sürətlə dəyişir."

Output:

Allah yaratmaq dünya sürətlə dəyişmək.

Reconstruction is intended for normalization, search, indexing, and NLP pipelines. It does not produce a grammatically inflected sentence.

JSON output

python azlemma.py text `
  --contextual `
  --json `
  "Allahın yaratdığı dünya sürətlə dəyişir."

Include punctuation tokens:

python azlemma.py text `
  --contextual `
  --json `
  --punctuation `
  "Allahın yaratdığı dünya sürətlə dəyişir."

Process a UTF-8 file

python azlemma.py file input.txt

With contextual disambiguation:

python azlemma.py file `
  --contextual `
  input.txt

Reconstruct a file as lemma-normalized text:

python azlemma.py file `
  --contextual `
  --reconstruct `
  input.txt

Interactive mode

python azlemma.py interactive

Contextual interactive mode:

python azlemma.py interactive --contextual

Exit with Ctrl+C or Ctrl+Z.

Runtime database schema

The main lookup table has the following logical structure:

CREATE TABLE forms (
    surface_norm TEXT NOT NULL,
    surface      TEXT NOT NULL,
    lemma        TEXT NOT NULL,
    pos          TEXT NOT NULL,
    entry_source TEXT NOT NULL,
    PRIMARY KEY (surface_norm, lemma, pos)
) WITHOUT ROWID;

A single surface form may have several rows when it has several possible lemmas or parts of speech.

Command reference

python azlemma.py [--db PATH] [--mutable] COMMAND ...

Available commands:

words        Look up separate words without context
text         Analyze a text string
file         Analyze a UTF-8 file
interactive  Analyze one input line at a time

Useful text and file options:

--contextual          Enable contextual POS disambiguation
--context-margin N    Set the minimum contextual score gap
--reconstruct         Replace resolved words with lemmas
--json                Produce JSON output
--punctuation         Include punctuation in JSON output
--analyses            Include full analyses when the database supports them

Display complete CLI help:

python azlemma.py --help
python azlemma.py text --help

Contextual disambiguation

Contextual mode uses general coarse-POS sequence preferences such as:

DET  -> ADJ
DET  -> N
ADJ  -> N
NUM  -> N
ADV  -> V
PRON -> V
N    -> POSTP

The implementation does not include rules such as:

if word == "sürətlə":
    ...

The same scoring mechanism is applied to every ambiguous token.

Limitations

  • Contextual disambiguation is heuristic, not a full syntactic parser.
  • The quality of the result depends on database coverage.
  • Unknown words are returned unchanged.
  • Low-confidence ambiguity is intentionally preserved.
  • Text reconstruction produces dictionary lemmas, not a natural sentence.
  • The compact database does not include detailed morphological analyses unless they are explicitly retained during database creation.

Data generation and distribution

The full morphological database is generated separately and then compacted for runtime use. The current database includes forms produced from Azerbaijani morphological paradigms, including finite verbs and non-finite verb forms such as participles, infinitives, verbal nouns, and converbs.

The generated SQLite database is not committed to the Git repository. It is distributed as the azlemma.sqlite asset attached to:

https://github.com/vrashad/azlemma/releases/tag/v1.0.0

After downloading it, place the file at:

data/azlemma.sqlite

Contributing

Issues and pull requests are welcome. When reporting a lemmatization problem, include:

  • the original word or sentence;
  • the actual output;
  • the expected lemma and POS;
  • whether strict or contextual mode was used;
  • the database version or generation date.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages