Hi everyone! π
This is a Multi-Threaded Web Crawler written in Go! π€π
This program crawls the web like a spider πΈοΈ:
- Starts with a website you give it π₯οΈ.
- Finds all the links on the webpage π.
- Recursively visits those links π (but only to a certain depth so it doesn't crawl forever! π ).
- Shows you all the crawled URLs in the terminal! π»
When you start the program, it might look like this in your terminal:
Starting web crawl at https://google.com with a depth of 2
Crawling: https://google.com (Depth: 2)
Crawling: https://accounts.google.com/ServiceLogin?hl=en&passive=true&continue=https://www.google.com/&ec=GAZAAQ (Depth: 1)
Crawling: https://www.youtube.com/?tab=w1 (Depth: 1)
Crawling: https://drive.google.com/?tab=wo (Depth: 1)
Crawling: https://www.google.com/imghp?hl=en&tab=wi (Depth: 1)
Crawling: https://news.google.com/?tab=wn (Depth: 1)
Crawling: https://maps.google.com/maps?hl=en&tab=wl (Depth: 1)
Crawling: https://play.google.com/?hl=en&tab=w8 (Depth: 1)
Crawling: https://www.google.com/intl/en/about/products?tab=wh (Depth: 1)
Crawling: https://mail.google.com/mail/?tab=wm (Depth: 1)
Crawling: http://www.google.com/history/optout?hl=en (Depth: 1)
Crawling completed.
Pretty cool, huh? π
-
Install Go: First, make sure you have Go installed π οΈ.
-
Clone the repository:
git clone https://github.com/Antot-12/Web-Crawler.git
-
Navigate to the project folder:
cd Web-Crawler -
Run it: Use this command:
go run crawler.go
-
Customize it: Change the
startURLandmaxDepthvalues in the code to crawl other websites or limit how deep it goes π§.