---
I'm curious if anyone knows whether it is better to pass structured data or unstructured data to embedding api's? If I ask ChatGPT, it says it is better to send unstructured data. (looking at the authors github, it looks like he generated embeddings from json strings)
My use case is for jsonresume, I am creating embeddings by sending full json versions as strings, but I've been experimenting with using models to translate resume.json's into full text versions first before creating embeddings. The results seem to be better but I haven't seen any concrete opinions on this.
My understanding is that unstructured data is better because it contains textual/semantic meaning because of natural lanaguage aka
skills: ['Javascript', 'Python']
is worse than; Thomas excels at Javascript and Python
Another question: What if the search was also a json embedding? JSON <> JSON embeddings could also be great?