Databricks Adds Built-In Search to Lakebase Postgres on AWS and Azure
The release puts meaning-based and keyword search beside application data. Databricks claims a cost and speed advantage, but enabling the feature restarts project computes and cannot be undone.
Databricks is making Lakebase Search generally available on AWS and Azure, bringing vector retrieval, BM25 keyword ranking, and operational-table joins into Postgres. The commercial case rests on company-run tests: 97% true-neighbor retrieval and 71 ms p99 latency on 100 million vectors, plus a claimed fourfold cost advantage versus a pgvector cloud-Postgres setup; results depend on workload and test configuration. Teams evaluating it should also account for an irreversible project-level enablement that restarts每?
01
lakebase_vector accepts pgvector types, operators, and query syntax; lakebase_text ranks keyword matches with BM25, and hybrid results are merged from separate searches.
02
After scale-to-zero, Databricks measured a 1.13-second p90 first-query response in its 100-million-vector test, distinct from active-search latency.
03
Conexiom says it searches more than 100 million rows with half the compute footprint of its former pgvector setup; this is a customer comparison.
Applications using Lakebase Postgres can now search for both a phrase’s meaning and its exact words without sending data to a separate search system. Databricks has made Lakebase Searchgenerally available on AWS and Azure, pitching one database for search and everyday application records.
Two ways to find a match
The release adds two Postgres extensions. A vector is a numerical representation of content; vector search can find a relevant row even if it uses different words from the query. Keyword search instead helps when the wording itself matters, such as a name or code. Lakebase can combine the two methods and join the results with live operational tables in a Postgres query.
lakebase_vector finds approximate nearest matches among vectors. Its index accepts pgvector’s vector types, distance operators and query syntax, according to Databricks’ documentation.
lakebase_text uses BM25 to rank keyword matches. That method weighs terms using information from the wider collection, so uncommon words can carry more weight than common ones.
Hybrid search is not just two indexes on the same table. Databricks’ example runs vector and keyword searches separately, then combines their ranked results. Teams can tune how those rankings are merged; the example is an implementation, not a published measure of hybrid-search quality.
Why the index works differently
Databricks argues that pgvector’s graph-based search becomes costly at large scale because fast queries depend on keeping the index in memory. When it spills to disk, following the graph requires many scattered reads. Lakebase takes another route: it groups vectors into blocks, checks which groups look promising and reads those blocks rather than traversing the whole index.
It also stores compact, roughly one-bit-per-dimension versions of vectors to narrow the candidates before checking a shortlist at full precision. Lakebase separates durable storage from computing power, so compute can stop when idle and resume for a new query. Databricks measured a 1.13-second 90th-percentile response for the first query after scale-to-zero in a 100-million-vector test. That cold-start figure is different from its latency result for active search.
A benchmark and a customer example
Those are company-run results, not a guarantee for every search workload. Databricks says its pgvector and DiskANN performance tests used one large instance apiece, a qualification worth keeping in view when comparing throughput across systems. The cost figure compares Lakebase with a cloud Postgres vendor using pgvector; it does not by itself price every possible way to build a search service.
Databricks also cites Conexiom, which says it runs hybrid search over more than 100 million rows with half the compute footprint of its previous pgvector setup. That is a customer comparison, not the same test as the 100-million-vector benchmark. For a prospective buyer, it points to a concrete question: whether the new index cuts resources on their own data while preserving the matches they need.
The decision before installation
There is an operational step before anyone installs an extension. Lakebase Search must be enabled in project settings, and Databricks’ documentation says that action restarts every compute in the project, drops active connections and cannot be reversed. The feature requires Postgres 16 or later. Teams planning to try it on a running application must account for that restart before testing the promised search gains.
Reader comments
Newest comments first. Replies stay oldest first.