datalab-to/surya

OCR, layout analysis, reading order, table recognition in 90+ languages

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 46 minutes ago
Added to GitGenius on September 2nd, 2026
Created on January 10th, 2024
Open Issues & Pull Requests: 196 (+0)
GitHub issues: Enabled
Number of forks: 1,537
Total Stargazers: 21,359 (+0)
Total Subscribers: 128 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.0 days
Mean response time: 21.6 days
90th percentile: 67.9 days
Tracked items: 149

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 60% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. Only 4% of issues opened in the past year have been closed. Three people close 67% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 118
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 472 days
Stale 30+ days: 109
Stale 90+ days: 100

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 1
Events in 7 days: 1

Top labels

  • bug: breaking (18)
  • bug: output (10)
  • enhancement (6)

Most active issues this week

Detailed Description

Surya is an OCR and document intelligence tool that performs optical character recognition, layout analysis, reading order detection, and table recognition across 90+ languages.

The tool addresses the challenge of extracting and understanding text from scanned documents and images. It uses a 650-million-parameter neural network model trained to recognize text with high accuracy while simultaneously analyzing document structure. The model identifies layout elements such as tables, images, and headers, determines the logical reading order of content, and recognizes table structure including rows and columns. It handles multilingual documents, supporting over 90 languages in a single model.

Surya suits projects that need to process diverse document types at scale, particularly those requiring both text extraction and structural understanding. The tool achieves 83.3% accuracy on a standard OCR benchmark while maintaining fast throughput of 5 pages per second on high-end hardware. It also includes smaller companion models for specialized tasks like line-level text detection and OCR error detection. For teams preferring managed infrastructure rather than self-hosting, the project's creators offer a cloud platform with free trial credits and a public playground for testing.

The project maintains active development with regular model improvements and expanded language support. The codebase receives consistent updates addressing performance optimization and feature additions. Community engagement is supported through an official Discord channel for user discussion and feedback.