CommonsDownloadTool is a Python script used to download and package thousands of files from Wikimedia Commons (made for Lingua Libre recordings) into .zip archives, plus helpers/ (titles_fetcher.py, titles_filterer.py) to list and filter titles.
- Operations/create_datasets.sh - LinguaLibre.org server script which calls and runs CommonsDownloadTool.
- lingualibre.org/datasets/ - LinguaLibre.org page to access the generated archives.
# Download from list of titles in ../titles-filtered.txt
python3 commons_download_tool.py --threads 1 --titles ../titles-filtered.txt --keep --fileformat wav --directory out --nozip
# Download from a Wikimedia Commons category
python3 commons_download_tool.py --category "Lingua_Libre_pronunciations-oci"
# Download from a category with multiple threads and keep files unzipped
python3 commons_download_tool.py --category "Lingua_Libre_pronunciations-fr" --threads 8 --keep --directory out
# Download from a category, force WAV format, and output to a specific zip file
python3 commons_download_tool.py --category "Lingua_Libre_pronunciations-en" --fileformat wav --output communications-en.zippython3 commons_download_tool.py --titles list.txt --user-agent 'MyBot/1.0 (contact)' --comment 'Archive of {count} files' --redirects ./redirects.txt # identify the bot, zip comment, record redirect titles (skipped)Build the titles list with the helpers, then download them:
python3 helpers/titles_fetcher.py -p 'LL-Q' -o ./tmp/titles.txt # list file titles (namespace 6) starting with LL-Q from the monthly dump
python3 helpers/titles_fetcher.py -n 6 -x 'LL-Q' -d ./dumps -o ./tmp/titles.txt # same, and keep the dump in ./dumps for later runs
python3 helpers/titles_fetcher.py -n 6 -x 'LL-Q' -p '^LL-Q\d{3,9}[_ ].+\.(?:wav|ogg)$' -o ./tmp/titles.txt # same, keep only wav/ogg names
python3 helpers/titles_filterer.py -i ./tmp/titles.txt -p 'Q1307730' -o ./tmp/filtered.txt # narrow the list, e.g. one language QID
python3 commons_download_tool.py --titles ./tmp/filtered.txt --threads 4 --output out.zip # download and zip the listRedirect pages are skipped and, with --redirects, recorded one per row.
.zip archives generated by this tool ship a comment detailing how many recordings they contain.
You can read this comment with the following commands:
- On Linux:
zipinfo -z archive.zip
- helpers/README.md - fetching and filtering titles (replica, dumps, Toolforge setup).
- Phabricator: Lingua-libre > Datasets and mass downlaods column — tickets manager
- Github: Lingua-libre/CommonsDownloadTool — code (Python)