a.k.a. Prompted subspaces
Reference: https://jasperlu.com/blog/prompted-similarity-search/
Problem statement: We want to do similarity search but the similarity will be conditioned with respect to a given concept. E.g. compare sentences by similarity conditioned on topic, or on language style, etc etc.
Background: I had tried doing something similar once, trying to define similarity 'along a dimension' of 'complexity'. I had tried using points of low complexity and high complexity and tried finding an 'average dimension vector' which points from low to high. Not super rigorous, I know. Not even sure if it worked properly. That project was mostly abandoned so I didn't get to explore this more.
What I understood: The technique in the blog post is doing something more mathematically rigorous... it's using the examples of the concept and applying the "first part of [[Principle Component Analysis|PCA]]", i.e. finding the subspace with the highest variance in the given concept. And then projecting down the rest of the data into that subspace and using the familiary concepts of cosine similarity.
Let's say we were to compare the similarity of things based on a concept of 'colour' we could take a set of colour embedding vectors (red, green, blue, yellow, so on...) and then use PCA to find a subspace where these vectors have maximum variance. Then we'd project down any other embedding down to this subspace and we'd get a 'color embedding' and a residual embedding. The residual will tell us how much this thing is _not related to color_ while the color embedding would be mostly aligned with the colour of the thing.
-------------
What I still didn't understand:
So let's say if we were to give it vectors of 'low complexity' and 'high complexity' it would likely find a single vector which has the highest variance between these two (doesn't it run into the same trouble? if you have n=1 example of each? I also seem to recall reading about 'contrastive PCA' back then). But the original two vectors need not lie on the straight line. (What? I think PCA would mostly give THE straight line between the vectors, no?)
I think I need to experiment with this and study this in more detail with text embeddings to get a better idea of how this works.
-----------------
Idea / Experiment:
Let's embed `/usr/share/dict/words` using a small embedding model and try this technique on some concepts.