HNHacker News
TopNewBestAskShowJobs

juliensalinas

41 karma · joined January 2, 2018

submissionscomments
juliensalinas··on Cursor IDE support hallucinates lockout policy, causes user cancellations
Relying on GenAI for support like that without a human in the loop is a huge mistake...
juliensalinas··on Comparing GenAI Inference Engines: TensorRT-LLM, VLLM, HF TGI, and LMDeploy
You can read the full comparison here: https://nlpcloud.com/genai-inference-engines-tensorrt-llm-vs...
juliensalinas··on Ask HN: What are you working on? (March 2025)
Sounds very cool. I'm curious how you manage to monitor Linkedin though. The only tool that seems capable of monitoring Linkedin is https://kwatch.io , so if you manage to achieve that too it's impressive.
juliensalinas··on Ask HN: Founders, what was the major sourcing channel for your first 100 users?
Social listening on HN, Reddit, X... I used https://kwatch.io and jumped into the relevant conversations to mention my product.
juliensalinas··on Ask HN: What is used instead of mention.com nowadays?
I use KWatch.io (https://kwatch.io) for social listening and it works very well for HN monitoring in my case. They also support other platforms (Reddit, Linkedin, Twitter..). But they don't propose advanced features like dashboards, analytics...
juliensalinas··on [dead]
Many are trying to install and deploy their own LLaMA 3 model, so here is a tutorial I just made showing how to deploy LLaMA 3 on an AWS EC2 instance: https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-...

Deploying LLaMA 3 8B is fairly easy but LLaMA 3 70B is another beast. Given the amount of VRAM needed you might want to provision more than one GPU and use a dedicated inference server like vLLM in order to split your model on several GPUs.

LLaMA 3 8B requires around 16GB of disk space and 20GB of VRAM (GPU memory) in FP16. As for LLaMA 3 70B, it requires around 140GB of disk space and 160GB of VRAM in FP16.

I hope it is useful, and if you have questions please don't hesitate to ask!

Julien

juliensalinas··on Show HN: Crowdlens – AI-powered social listening
How does this solution compare to platforms like https://kwatch.io or brand24 for hacker news monitoring?

Does it monitor hacker news in real time?

juliensalinas··on Who uses Google TPUs for inference in production?
We tried hard to move some of our inference workloads to TPUs at NLP Cloud, but finally gave up (at least for the moment) basically for the reasons you mention. We now only perform our fine-tunings on TPUs using JAX (see https://nlpcloud.com/how-to-fine-tune-llama-openllama-xgen-w...) and we are happy like that.

It seems to me that Google does not really want to sell TPUs but only showcase their AI work and maybe get some early adopters feedback. It must be quite a challenge for them to create a dynamic community around JAX and TPUs if TPUs stay a vendor locked-in product...

juliensalinas··on Ask HN: Best Alternatives to OpenAI ChatGPT?
Claude (Anthropic) might be the closest direct alternative to ChatGPT (but it's not available in alls countries). You might also want to try ChatDolphin by NLP Cloud (a company I created 3 years ago as an OpenAI alternative): https://chat.nlpcloud.com Open-source is also catching up very quickly. The best models you might want to try today are LLaMA 2 70B, Yi 34B, or Mistral 7B.
juliensalinas··on Mistral 7B
For those who want to try Mistral 7b, here is a video that shows how to do it on an A10 GPU on AWS: https://www.youtube.com/watch?v=88ByWjM-KGM
juliensalinas··on GPT-4 API General Availability
Thank you.
juliensalinas··on GPT-4 API General Availability
Thank you for the update! Do you happen to know if there are quality comparisons somewhere, between llama.cpp and exllama? Also, in terms of VRAM consumption, are they equivalent?
juliensalinas··on GPT-4 API General Availability
Oh it seems you're right, I had missed that.

As far as I can see llama.cpp with CUDA is still a bit slower than ExLLaMA but I never had the chance to do the comparison by myself, and maybe it will change soon as these projects are evolving very quickly. Also I am not exactly sure whether the quality of the output is the same with these 2 implementations.

juliensalinas··on GPT-4 details leaked?
LLaMA 30B or 60B can be very impressive when correctly prompted. Deploying the 60B version is a challenge though and you might need to apply 4-bit quantization with something like https://github.com/PanQiWei/AutoGPTQ or https://github.com/qwopqwop200/GPTQ-for-LLaMa . Then you can improve the inference speed by using https://github.com/turboderp/exllama .

If you prefer to use an "instruct" model à la ChatGPT (i.e. that does not need few-shot learning to output good results) you can use something like this: https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored... The interesting thing with these Uncensored models is that they don't constantly answer that they cannot help you (which is what ChatGPT and GPT-4 are doing more and more).

juliensalinas··on GPT-4 API General Availability
llama.cpp focuses on optimizing inference on a CPU, while exllama is for inference on a GPU.
juliensalinas··on ChatGPT loses users for first time, shaking faith in AI revolution
Totally agree. Actually a couple of months ago Sam Altman even admitted that they had a very hard time doing proper "engineering" (meaning that they had the right team to create a very good LLM but not the right team to productionize and their models and APIs). Many people are actually finding the OpenAI API very unstable and do not plan to rely on OpenAI for their production workloads.
juliensalinas··on ChatGPT loses users for first time, shaking faith in AI revolution
NLP Cloud (especially the Dolphin and Fine-tuned GPT-NeoX 20B models)
juliensalinas··on Ask HN: Best UNCENSORED language model comparable to ChatGPT?
You might want to try our ChatDolphin model on NLP Cloud that is very similar to Vicuna and uncensored: https://nlpcloud.com/home/playground/text-generation (select the ChatDolphin model at the top right).

I hope it will be useful.

juliensalinas··on Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
As the founder of NLP Cloud (https://nlpcloud.com) I can only guess how costly it must be for OpenAI to maintain several versions of GPT-4 in parallel. I think that the main reason why they don't provide you with a way to pin a specific model version is because of the huge GPU costs involved. There might also be this "alignment" thing that makes them delete a model because they realize that it has specific capacities that they don't want people to use anymore.

On NLP Cloud we're doing our best to make sure that once a model is released it is "pinned" so our users can be sure that they won't face any regression in the future. But again it costs money so profitability can be challenging.

juliensalinas··on Ask HN: Anyone know any OpenAI API hosted alternatives?
You might want to try NLP Cloud: https://nlpcloud.com
juliensalinas··on Ask HN: Open-source ChatGPT alternatives?
Some alternatives are explained in this article: https://nlpcloud.com/chatgpt-open-source-alternatives.html But it will require some prompt engineering in order to get the same level of instruction as ChatGPT.
juliensalinas··on Ask HN: How do you plan on making money with Open AI API?
Smaller models like Curie can work quite well too. But they are less "instruct-like" models so you will need to properly use few-shot learning (aka "prompt engineering") in order to get good results: https://nlpcloud.com/effectively-using-gpt-j-gpt-neo-gpt-3-a...

I takes a bit more work though, and it makes your requests bigger so more expensive.

juliensalinas··on Ask HN: Self-hosted/open-source ChatGPT alternative? Like Stable Diffusion
You might want to have a look at this article that mentions a couple of open-source alternatives: https://nlpcloud.com/chatgpt-open-source-alternatives.html None of them are easy to self-host though...
juliensalinas··on Ask HN: Self-hosted/open-source ChatGPT alternative? Like Stable Diffusion
The best open-source alternatives you can find today are GPT-NeoX 20B, GPT-J, Bloom, and OPT. But these are all generative models à la GPT-3. In order to turn them into a chatbot you will need to use few-shot learning: https://nlpcloud.com/effectively-using-gpt-j-gpt-neo-gpt-3-a...

These models can be self-hosted but they will require advanced hardware and some specific skills related to AI model deployment.

juliensalinas··on GPT-3 can create both sides of an Interactive Fiction transcript
You could follow EleutherAI's official guide: https://github.com/EleutherAI/gpt-neox You could also use a hosted service that proposes GPT-NeoX like https://nlpcloud.com or https://goose.ai
juliensalinas··on 1 week of Stable Diffusion
I worked on the Stable Diffusion and GPT-J integrations on NLP Cloud (https://nlpcloud.com/). Both can be used in FP16 without any noticeable quality drop (in my opinion). Stable diffusion requires 7GB of VRAM on a Tesla T4 GPU. GPT-J requires 12GB of VRAM (but if you really try to use the 2048 tokens context, the VRAM will go up and reach something like 20GB of VRAM).
juliensalinas··on OpenAI API pricing update FAQ
I've been testing BLOOM for a while and it seems it is working very well with good few-shot learning. See this post about prompt examples: https://nlpcloud.com/effectively-using-gpt-j-gpt-neo-gpt-3-a... All of these examples work great with BLOOM.

But for the moment you can't just use pure natural language instructions with BLOOM indeed. Maybe it will come!

juliensalinas··on Alpa.ai: Free, Unlimited Opt-175B Text Generation
For the moment I can't use it. I'm getting the following error:

    <html>
<head><title>403 Forbidden</title></head>

<body>

<center><h1>403 Forbidden</h1></center>

<hr><center>Microsoft-Azure-Application-Gateway/v2</center>

</body>

</html>

juliensalinas··on Alpa.ai: Free, Unlimited Opt-175B Text Generation
Great to see that someone proposes a way to try OPT-175B, at last. Great job! Can you say more about the hardware you're using behind this service?
juliensalinas··on Textsynth: Bellard's free GPT-NeoX-20B, GPT-J playground and paid API
Totally, and that's because GPT models don't really support multilingual content. It works, but very poorly. It's the case for GPT-J, GPT-NeoX 20B, and even GPT-3.

I recently integrated GPT-NeoX 20B on NLP Cloud: https://nlpcloud.io . I had hopes that non-English languages would be better supported than with GPT-J since the model was trained on 20B parameters instead of 6B parameters, but quality still leaves to be desired. In my opinion, the best way to handle text generation in non-English languages for the moment is to couple it with a good translation module. I actually wrote an article about that: https://nlpcloud.io/multilingual-nlp-how-to-perform-nlp-in-n... .

But there is hope! Bigscience is about to release a huge NLP model that should theoretically work very well in almost 50 languages: https://bigscience.huggingface.co/ . We'll soon see if it's true!

Page 1 of 2Next →