apache/nutch

Apache Nutch is an extensible and scalable web crawler

View on GitHub ↗Jump to charts ↓

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 40 minutes ago
Added to GitGenius on September 21st, 2026
Created on May 21st, 2009
Open Issues & Pull Requests: 7 (+0)
GitHub issues: Disabled - open counts may still include pull requests.
Number of forks: 1,280
Total Stargazers: 3,298 (+0)
Total Subscribers: 224 (+0)

Repository Insights (GitGenius)

Most active contributors

Sign in to see contributor activity.
Sign in

Related repositories by overlapping contributors

No overlapping-contributor repos identified yet.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

GitHub issues are disabled for this repository, so issue analytics and the issue explorer are not available.

Detailed Description

Apache Nutch is an extensible and scalable web crawler built on Java that automates the discovery and indexing of web content at large scale.

Nutch addresses the challenge of crawling the web efficiently by distributing crawl tasks across multiple machines using Hadoop. The tool handles the core responsibilities of a web crawler: fetching pages, parsing content, extracting links, and managing the crawl frontier. Its architecture separates concerns into distinct plugins and components, allowing operators to customize behavior for different crawling scenarios without modifying core code. The crawler manages politeness through configurable delays and respects robots.txt directives, while supporting both focused crawls of specific domains and broad web-scale operations.

Nutch suits organizations that need to build search indexes, monitor web content at scale, or integrate web crawling into data pipelines. It is particularly valuable for projects requiring distributed crawling across commodity hardware rather than single-machine solutions. The tool works well when you need fine-grained control over crawl behavior and can invest in operational complexity. Teams should choose Nutch when extensibility matters more than simplicity, as the framework requires understanding its plugin architecture and Hadoop integration to be effective.

The project maintains steady development activity with regular commits addressing bug fixes, dependency updates, and feature enhancements. Pull requests receive review and discussion from maintainers before merging. The codebase shows active maintenance of its core crawling logic and integration with evolving Hadoop ecosystems. Issue tracking reflects ongoing user engagement with both feature requests and problem reports being addressed over time.