google/langextract

A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 39 minutes ago
Added to GitGenius on December 22nd, 2025
Created on July 8th, 2025
Open Issues & Pull Requests: 120 (+0)
GitHub issues: Enabled
Number of forks: 2,720
Total Stargazers: 38,931 (+0)
Total Subscribers: 172 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 2.3 days
Mean response time: 24.6 days
90th percentile: 56.8 days
Tracked items: 180

Maintainer activity

1 person did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

Roughly one issue in four opened in the past year never receives a reply. 97% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "discussion" is answered fastest, typically in about 11 hours, while "alternative-llm" waits about 3 days. Only 56% of issues opened in the past year have been closed. Three people close 85% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 71
New in 7 days: 4
Closed in 7 days: 11
Avg open age: 166 days
Stale 30+ days: 66
Stale 90+ days: 55

Recent activity

Opened in 7 days: 4
Closed in 7 days: 11
Comments in 7 days: 7
Events in 7 days: 16

Top labels

  • discussion (13)
  • alternative-llm (11)
  • enhancement (11)
  • bug (8)
  • plugin (5)
  • documentation (2)
  • more details required (2)
  • question (2)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

LangExtract is a Python library for extracting structured information from unstructured text using large language models with precise source grounding and interactive visualization.

The library addresses the challenge of converting unstructured documents like clinical notes or reports into organized, structured data while maintaining verifiable connections to the source material. It works by chunking input text, processing it in parallel through an LLM based on user-defined extraction instructions, and then grounding every extracted value back to its exact character span in the original document. This approach ensures that extracted information can be traced and verified against the source.

Developers should choose this tool when they need reliable structured data extraction with accountability—particularly in domains like healthcare, legal review, or any context where tracing extracted values to their source is critical. The library supports multiple LLM providers including Google's Gemini models, OpenAI, and local models via Ollama, as well as custom model providers. It includes interactive visualization capabilities and has been applied to use cases ranging from medication extraction to radiology report structuring.

The project maintains a substantial base of real-world adopters, with almost all open issues raised by outside users rather than the core team. Maintainers typically respond to new issues and pull requests within a few days. Work in the issue tracker is dominated by discussions around alternative LLM providers and enhancement requests, reflecting active community engagement with the tool's extensibility and model support.