A lightweight Azerbaijani lemmatizer backed by a compact SQLite database.
AzLemma performs deterministic dictionary-based lemmatization without loading a neural model or requiring external Python packages. It supports strict lookup, optional context-aware POS disambiguation, text reconstruction, JSON output, file processing, and interactive use.
- Fast SQLite-based word lookup
- Azerbaijani-aware Unicode normalization and lowercasing
- Lemmatization of nouns, proper nouns, adjectives, verbs, participles, verbal nouns, infinitives, and converbs
- Strict ambiguity preservation by default
- Optional contextual POS disambiguation
- Batch lookup for separate words
- Text tokenization with punctuation preservation
- Lemma-based text reconstruction
- JSON output
- UTF-8 file processing
- Interactive mode
- No third-party Python dependencies
The runtime database contains mappings from an observed word form to one or more lemma and part-of-speech candidates:
surface_norm -> lemma, POS, source
For example:
yaratdığı -> yaratmaq, V, moraz
dəyişir -> dəyişmək, V, moraz
sürətlə -> sürət, N, moraz
sürətlə -> sürətlə, ADV, identity
When all candidates point to the same lemma, AzLemma resolves the word directly. When several different lemmas are possible, strict mode reports the word as ambiguous.
Contextual mode uses generic coarse-POS transition scores and a Viterbi-style sequence algorithm. It does not contain word-specific exceptions. A candidate is selected only when its score exceeds the second-best candidate by the configured confidence margin.
- Python 3.10 or newer
- The
azlemma.sqliteruntime database
No pip install step is required.
Clone the repository:
git clone https://github.com/vrashad/azlemma.git
cd azlemmaThe SQLite database is too large for normal Git storage, so it is distributed as a GitHub Release asset.
Release page:
https://github.com/vrashad/azlemma/releases/tag/v1.0.0
Create the local data directory:
New-Item -ItemType Directory -Force .\dataDownload the database using GitHub CLI:
gh release download v1.0.0 `
--repo vrashad/azlemma `
--pattern "azlemma.sqlite" `
--dir .\dataOr download it directly with PowerShell:
Invoke-WebRequest `
-Uri "https://github.com/vrashad/azlemma/releases/download/v1.0.0/azlemma.sqlite" `
-OutFile ".\data\azlemma.sqlite"You can also download azlemma.sqlite manually from the release page and
place it inside the data directory.
Verify that the file exists:
Get-Item .\data\azlemma.sqliteAfter downloading the database, the project should look like this:
azlemma/
├── azlemma.py
├── README.md
├── .gitignore
└── data/
└── azlemma.sqlite
python azlemma.py text "Allahın yaratdığı dünya sürətlə dəyişir."Example output:
Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
dünya -> dünya [N; moraz]
sürətlə -> AMBIGUOUS: sürət/N, sürətlə/ADV
dəyişir -> dəyişmək [V; moraz]
Enable contextual disambiguation:
python azlemma.py text `
--contextual `
"Allahın yaratdığı dünya sürətlə dəyişir."Example output:
Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
dünya -> dünya [N; moraz]
sürətlə -> sürətlə [ADV; identity; CONTEXTUAL]
dəyişir -> dəyişmək [V; moraz]
By default, AzLemma looks for:
data/azlemma.sqlite
You can provide another path with the global --db option:
python azlemma.py --db C:\data\azlemma.sqlite words yaratdığıThe global option must appear before the subcommand.
python azlemma.py words Allahın yaratdığı yaradaraq dəyişirExample output:
Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
yaradaraq -> yaratmaq [V; moraz]
dəyişir -> dəyişmək [V; moraz]
Separate-word lookup never uses contextual disambiguation:
python azlemma.py words sürətləsürətlə -> AMBIGUOUS: sürət/N, sürətlə/ADV
python azlemma.py text "Allahın yaratdığı dünya sürətlə dəyişir."Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
dünya -> dünya [N; moraz]
sürətlə -> AMBIGUOUS: sürət/N, sürətlə/ADV
dəyişir -> dəyişmək [V; moraz]
Strict mode never forces a choice between different lemmas.
python azlemma.py text `
--contextual `
"Allahın yaratdığı dünya sürətlə dəyişir."Allahın -> Allah [PROPN; fallback_propn]
yaratdığı -> yaratmaq [V; moraz]
dünya -> dünya [N; moraz]
sürətlə -> sürətlə [ADV; identity; CONTEXTUAL]
dəyişir -> dəyişmək [V; moraz]
The contextual algorithm prefers the ADV candidate because the coarse POS
sequence around the word supports an adverb before a verb.
python azlemma.py text `
--contextual `
--context-margin 2.5 `
"Allahın yaratdığı dünya sürətlə dəyişir."The default margin is 2.0.
A higher value makes the algorithm more conservative. When the score gap is too small, the result remains ambiguous.
Strict reconstruction:
python azlemma.py text `
--reconstruct `
"Allahın yaratdığı dünya sürətlə dəyişir."Contextual reconstruction:
python azlemma.py text `
--contextual `
--reconstruct `
"Allahın yaratdığı dünya sürətlə dəyişir."Output:
Allah yaratmaq dünya sürətlə dəyişmək.
Reconstruction is intended for normalization, search, indexing, and NLP pipelines. It does not produce a grammatically inflected sentence.
python azlemma.py text `
--contextual `
--json `
"Allahın yaratdığı dünya sürətlə dəyişir."Include punctuation tokens:
python azlemma.py text `
--contextual `
--json `
--punctuation `
"Allahın yaratdığı dünya sürətlə dəyişir."python azlemma.py file input.txtWith contextual disambiguation:
python azlemma.py file `
--contextual `
input.txtReconstruct a file as lemma-normalized text:
python azlemma.py file `
--contextual `
--reconstruct `
input.txtpython azlemma.py interactiveContextual interactive mode:
python azlemma.py interactive --contextualExit with Ctrl+C or Ctrl+Z.
The main lookup table has the following logical structure:
CREATE TABLE forms (
surface_norm TEXT NOT NULL,
surface TEXT NOT NULL,
lemma TEXT NOT NULL,
pos TEXT NOT NULL,
entry_source TEXT NOT NULL,
PRIMARY KEY (surface_norm, lemma, pos)
) WITHOUT ROWID;A single surface form may have several rows when it has several possible lemmas or parts of speech.
python azlemma.py [--db PATH] [--mutable] COMMAND ...
Available commands:
words Look up separate words without context
text Analyze a text string
file Analyze a UTF-8 file
interactive Analyze one input line at a time
Useful text and file options:
--contextual Enable contextual POS disambiguation
--context-margin N Set the minimum contextual score gap
--reconstruct Replace resolved words with lemmas
--json Produce JSON output
--punctuation Include punctuation in JSON output
--analyses Include full analyses when the database supports them
Display complete CLI help:
python azlemma.py --help
python azlemma.py text --helpContextual mode uses general coarse-POS sequence preferences such as:
DET -> ADJ
DET -> N
ADJ -> N
NUM -> N
ADV -> V
PRON -> V
N -> POSTP
The implementation does not include rules such as:
if word == "sürətlə":
...The same scoring mechanism is applied to every ambiguous token.
- Contextual disambiguation is heuristic, not a full syntactic parser.
- The quality of the result depends on database coverage.
- Unknown words are returned unchanged.
- Low-confidence ambiguity is intentionally preserved.
- Text reconstruction produces dictionary lemmas, not a natural sentence.
- The compact database does not include detailed morphological analyses unless they are explicitly retained during database creation.
The full morphological database is generated separately and then compacted for runtime use. The current database includes forms produced from Azerbaijani morphological paradigms, including finite verbs and non-finite verb forms such as participles, infinitives, verbal nouns, and converbs.
The generated SQLite database is not committed to the Git repository. It is
distributed as the azlemma.sqlite asset attached to:
https://github.com/vrashad/azlemma/releases/tag/v1.0.0
After downloading it, place the file at:
data/azlemma.sqlite
Issues and pull requests are welcome. When reporting a lemmatization problem, include:
- the original word or sentence;
- the actual output;
- the expected lemma and POS;
- whether strict or contextual mode was used;
- the database version or generation date.