pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Jonathan S. Katz <jkatz@postgresql.org>
To: Giuseppe Broccolo <g.broccolo.7@gmail.com>
To: Nathan Bossart <nathandbossart@gmail.com>
Cc: pgsql-hackers@postgresql.org
Cc: mail@joeconway.com
Subject: Re: vector search support
Date: Fri, 26 May 2023 10:37:57 -0400
Message-ID: <49c7ba52-818a-6d0b-b8fd-eadef8e195a1@postgresql.org> (raw)
In-Reply-To: <CAFtuf8CR6LKu0sVfOBgEKjPtRf6=n=QZSWyD_+yWkSnKYMWD-A@mail.gmail.com>
References: <20230422000723.GB1527017@nathanxps13>
	<CAFtuf8CR6LKu0sVfOBgEKjPtRf6=n=QZSWyD_+yWkSnKYMWD-A@mail.gmail.com>

On 4/26/23 9:31 AM, Giuseppe Broccolo wrote:
> Hi Nathan,
> 
> I find the patches really interesting. Personally, as Data/MLOps 
> Engineer, I'm involved in a project where we use embedding techniques to 
> generate vectors from documents, and use clustering and kNN searches to 
> find similar documents basing on spatial neighbourhood of generated 
> vectors.

Thanks! This seems to be a pretty common use-case these days.

> We finally opted for ElasticSearch as search engine, considering that it 
> was providing what we needed:
> 
> * support to store dense vectors
> * support for kNN searches (last version of ElasticSearch allows this)

I do want to note that we can implement indexing techniques with GiST 
that perform K-NN searches with the "distance" support function[1], so 
adding the fundamental functions to help with this around known vector 
search techniques could add this functionality. We already have this 
today with "cube", but as Nathan mentioned, it's limited to 100 dims.

> An internal benchmark showed us that we were able to achieve the 
> expected performance, although we are still lacking some points:
> 
> * clustering of vectors (this has to be done outside the search engine, 
> using DBScan for our use case)

 From your experience, have you found any particular clustering 
algorithms better at driving a good performance/recall tradeoff?

> * concurrency in updating the ElasticSearch indexes storing the dense 
> vectors

I do think concurrent updates of vector-based indexes is one area 
PostgreSQL can ultimately be pretty good at, whether in core or in an 
extension.

> I found these patches really interesting, considering that they would 
> solve some of open issues when storing dense vectors. Index support 
> would help a lot with searches though.

Great -- thanks for the feedback,

Jonathan

[1] https://www.postgresql.org/docs/devel/gist-extensibility.html


Attachments:

  [application/pgp-signature] OpenPGP_signature (839B, ../49c7ba52-818a-6d0b-b8fd-eadef8e195a1@postgresql.org/2-OpenPGP_signature)
  download

view thread (8+ messages)  latest in thread

Message-ID: <49c7ba52-818a-6d0b-b8fd-eadef8e195a1@postgresql.org>
Permalink:  ../49c7ba52-818a-6d0b-b8fd-eadef8e195a1@postgresql.org/
Also on:    postgresql.org/message-id/49c7ba52-818a-6d0b-b8fd-eadef8e195a1@postgresql.org

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: jkatz@postgresql.org, g.broccolo.7@gmail.com, nathandbossart@gmail.com, mail@joeconway.com
  Subject: Re: vector search support
  In-Reply-To: <49c7ba52-818a-6d0b-b8fd-eadef8e195a1@postgresql.org>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox