Full threadnostrebored·Avoiding the server call seems to only be time-efficient if calls to embed are time-efficient. The nice thing about server side inference is that you can use embedding models that can handle web-scale data independent of hardwareView on HN