timescale/pgvectorscale

Postgres extension for vector search (DiskANN), complements pgvector for performance and scale. Postgres OSS licensed.

View on GitHub ↗Jump to charts ↓

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 2 hours ago
Added to GitGenius on August 18th, 2025
Created on July 1st, 2023
Open Issues & Pull Requests: 20 (+0)
GitHub issues: Enabled
Number of forks: 157
Total Stargazers: 3,137 (+0)
Total Subscribers: 25 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 4.0 days
Mean response time: 31.3 days
90th percentile: 97.8 days
Tracked items: 89

Maintainer activity

1 person did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 10
New in 7 days: 2
Closed in 7 days: 0
Avg open age: 339 days
Stale 30+ days: 7
Stale 90+ days: 5

Recent activity

Opened in 7 days: 2
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • community (53)
  • pgvectorscale (51)
  • bug (41)
  • enhancement (10)
  • question (8)
  • good first issue (4)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

pgvectorScale is a project by Timescale aimed at dramatically accelerating vector similarity search within PostgreSQL using pgvector, TimescaleDB, and specialized hardware like GPUs. It’s essentially a drop-in replacement for pgvector’s indexing functionality, offering significantly improved query performance, particularly for large datasets, without requiring changes to application code. The core idea is to leverage TimescaleDB’s columnar storage and parallel processing capabilities alongside GPU acceleration to perform approximate nearest neighbor (ANN) searches much faster than traditional methods.

At its heart, pgvectorScale utilizes a technique called HNSW (Hierarchical Navigable Small World) graph indexing, similar to pgvector, but with key optimizations. Instead of relying solely on CPU processing, pgvectorScale offloads the computationally intensive indexing and search operations to GPUs. This is achieved through a custom CUDA kernel implementation designed to efficiently traverse the HNSW graph. The project focuses on maximizing GPU utilization and minimizing data transfer between CPU and GPU, which are common bottlenecks in GPU-accelerated database systems. It supports various distance metrics including L2 (Euclidean), cosine similarity, and inner product.

The architecture involves a PostgreSQL extension that intercepts pgvector index creation and search requests. When pgvectorScale is enabled, these requests are redirected to the GPU for processing. The extension manages the transfer of vector data to the GPU, executes the ANN search using the CUDA kernel, and then returns the results back to PostgreSQL. Crucially, pgvectorScale is designed to be compatible with existing pgvector workflows. Applications using pgvector can continue to use the same API calls (e.g., `CREATE INDEX USING pgvector`, `SELECT ... ORDER BY vector_column <-> query_vector`) without modification. This ease of integration is a major advantage.

Currently, pgvectorScale is focused on NVIDIA GPUs and requires a compatible CUDA installation. The project provides Docker images for simplified deployment and testing. Performance gains are substantial, with benchmarks demonstrating speedups of several orders of magnitude compared to CPU-based pgvector indexing, especially as the dataset size increases. The speedup is dependent on factors like GPU model, vector dimensionality, dataset size, and query complexity. TimescaleDB’s hyperfunctions framework is used to define and execute the GPU kernels within PostgreSQL.

The project is still under active development, but it represents a significant step forward in making vector similarity search practical for large-scale applications. Future development plans include support for additional distance metrics, improved indexing algorithms, and potentially support for other GPU vendors. pgvectorScale is particularly well-suited for applications like recommendation systems, image/video search, natural language processing, and fraud detection, where efficient similarity search is critical. It effectively bridges the gap between the flexibility of PostgreSQL and the performance of specialized hardware for vector database workloads.