lorien/awesome-web-scraping

List of libraries, tools and APIs for web scraping and data processing.

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 27 minutes ago
Added to GitGenius on September 8th, 2026
Created on August 12th, 2015
Open Issues & Pull Requests: 2 (+0)
GitHub issues: Enabled
Number of forks: 940
Total Stargazers: 8,146 (+0)
Total Subscribers: 227 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 11.2 hours
Mean response time: 6.4 days
90th percentile: 17.3 days
Tracked items: 12

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 0
New in 7 days: 0
Closed in 7 days: 0
Avg open age: N/A days
Stale 30+ days: 0
Stale 90+ days: 0

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Awesome Web Scraping is a curated list of libraries, tools, and resources for web scraping and data processing across multiple programming languages.

The project addresses the challenge of discovering suitable scraping tools by organizing them into language-specific categories covering Python, PHP, Ruby, JavaScript, and Go, alongside command-line tools and educational manuals. It also includes references to headless browsers, captcha solving services, proxy marketplaces, and community discussion groups, providing a centralized index for developers seeking both technical solutions and practical guidance.

Developers should use this list when evaluating scraping technology stacks or learning web scraping fundamentals. It suits teams building crawlers or data extraction pipelines who need to compare available options across their preferred programming language. The resource is particularly valuable for those new to scraping who benefit from curated recommendations and educational materials rather than searching through fragmented documentation.

The project maintains an organized structure with separate documentation files for each language and topic area, making it straightforward to navigate to relevant tools. Community contributions are welcomed through a documented contribution process, enabling the list to expand as new libraries and services emerge. Discussion groups in multiple languages provide spaces for practitioners to share knowledge and troubleshoot implementation challenges.