tesseract-ocr/tesseract

Tesseract Open Source OCR Engine (main repository)

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 59 minutes ago
Added to GitGenius on August 28th, 2026
Created on August 12th, 2014
Open Issues & Pull Requests: 489 (+0)
GitHub issues: Enabled
Number of forks: 10,779
Total Stargazers: 76,266 (+0)
Total Subscribers: 1,710 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.3 hours
Mean response time: 24.9 days
90th percentile: 26.9 days
Tracked items: 215

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 83% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "feature request" is answered fastest, typically in about an hour, while "layout analysis" waits about 2 days. Only 5% of issues opened in the past year have been closed. Three people close 83% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 102
New in 7 days: 3
Closed in 7 days: 2
Avg open age: 1,310 days
Stale 30+ days: 97
Stale 90+ days: 81

Recent activity

Opened in 7 days: 3
Closed in 7 days: 2
Comments in 7 days: 2
Events in 7 days: 5

Top labels

  • question (30)
  • bug (27)
  • feature request (19)
  • accuracy (13)
  • layout analysis (10)
  • traineddata (10)
  • build process (8)
  • unexpected termination (8)

Detailed Description

Tesseract is an OCR engine that converts images of text into machine-readable text through optical character recognition.

The tool solves the problem of extracting text from images by combining two recognition approaches. Tesseract 4 introduced a neural network-based LSTM engine focused on line recognition, while maintaining backward compatibility with the legacy character-pattern-based engine from Tesseract 3 through a legacy OCR mode. The engine requires trained data files to function and benefits from image quality preprocessing to achieve better results.

Tesseract suits projects that need to extract text from images without a graphical interface. The tool recognizes more than 100 languages out of of the box with UTF-8 support, processes multiple image formats including PNG, JPEG, and TIFF, and outputs results in various formats such as plain text, hOCR, PDF, TSV, ALTO, and PAGE. It can be trained to recognize additional languages. Projects requiring a GUI should look to third-party applications built on top of Tesseract rather than the engine itself.

Development is led by Stefan Weil with Zdenko Podobny as maintainer. The project maintains an active issue tracker and planning documentation. The codebase is written in C++ and receives updates across minor versions and bugfix releases on the main branch.