Skip to content

Repository files navigation

clScan

Tests

A local job-search workbench that collects Craigslist listings, prioritizes them with configurable rules and a local language model, and turns the results into a shortlist you can read, listen to, and sort yourself.

Collection and assessment run concurrently: eligible listings enter a persistent queue as they arrive. Software-related keywords can move jobs to the front of that queue, so useful results appear before the full search finishes.

What it does

  • Searches across markets using multiple profiles, rotating queries and configurable budgets.
  • Filters duplicates, stale ads and poor matches before spending model time.
  • Scores listings with llama.cpp or Ollama and preserves prompts, responses and timing for inspection.
  • Provides a live website with run history, shortlist filters, skill keywords and optional local Piper audio.
  • Saves Bad result, Maybe not, Maybe yes, and Hell yes judgments in SQLite. Sorted listings leave the shortlist so other matches can appear.
  • Resumes queued work after restart and can reassess stored listings without collecting again.
  • Retrieves reply addresses through the normal browser interface when enabled. It does not send applications or messages.

This is a single-user, local application. Model scores are triage suggestions, not verification of an employer or offer. Human ratings are recorded for later analysis; they do not currently train the model or alter its scores.

Quick start

Requires Linux, Bash, Python 3.11+, and a local llama.cpp or Ollama server. Linux is the supported platform because the pipeline uses POSIX file locking and the reply helper supports Linux clipboard tools. Chromium is installed by the setup script.

git clone https://github.com/OperationAzura/clScan.git
cd clScan
./setup.sh

The script creates .venv, installs the application and Chromium, and copies config.example.toml to config.toml if needed. It preserves existing configuration and data. Model weights, voice files and system services are configured separately; see setup instructions.

  1. Edit config.toml: set your location, profiles, search terms, ranking preferences and model endpoint. The supplied North Carolina/Virginia geography is an example to customize.

  2. Start your local model server, then check it:

    .venv/bin/clscan doctor
  3. Start the website:

    .venv/bin/clscan-web
  4. Open http://127.0.0.1:8765, choose New run, review the configuration and queue a search.

The example allows 120 queries, up to 600 new downloads and 1,600 listing/profile assessments, with a software-heavy shortlist. These are ceilings, not guaranteed coverage. Start with smaller budgets while tuning a new profile. Piper is optional; without a configured voice the text workflow works, but audio requests cannot play.

Sort the results

Button Meaning
Bad result This should not have been selected; retain it to investigate filtering mistakes.
Maybe not Reasonable selection, but less interesting; retain it for future ranking analysis.
Maybe yes Good listing to revisit later.
Hell yes Excellent match; retain it as a positive ranking example.

Every choice saves the ad and assessment context, removes that content from the shortlist across profiles and runs, and allows qualifying replacements to appear. Use Show to revisit a category or Undo sorting to restore normal shortlist eligibility. Saved snapshots remain available when the original posting changes or expires.

Documentation

Development

./setup.sh --dev --skip-browser
.venv/bin/python -m unittest discover -s tests -v
.venv/bin/python -m build
# After installing Chromium:
.venv/bin/python tests/browser_feedback.py

Unit/API tests mock external services and use temporary databases. The feedback browser test also uses an isolated database. The separate audio browser check requires a configured test instance; see the contributing guide.

Personal configuration, SQLite records, collected ads, contact details, models, audio and browser state are excluded from Git. Source data is not bundled. examples/listings.jsonl contains synthetic examples only. Review the source site's access terms before collecting; access blocks and verification challenges stop collection rather than being bypassed.

About

Local job-search workbench with streaming LLM triage, accessible audio review, and persistent human feedback

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages