HNHacker News
TopNewBestAskShowJobs

pilooch

777 karma · joined January 7, 2011

https://www.jolibrain.com/
submissionscomments
pilooch··on Introducing Gemma 3n
Fix: it's the E2B
pilooch··on Introducing Gemma 3n
This model is fully compatible with anything previously done with gemma3. Just passed it to one of my vlm fine-tuning scripts and it started without issues (hf transformer code). On a single GPU with Lora the E4B model takes 18Gb of VRAM in batch size 1 where gemma-4B was 21Gb. Nice one from deepmind, the gemma3 family tops the open weights VLLMs.
pilooch··on AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms
Yes the 2023 reference on island based evolution with LLMs (nature article) https://www.nature.com/articles/s41586-023-06924-6 has more details.

Agreed the dimensions/features are key. These white papers are an insult to science...

pilooch··on Gonzalo Guerrero
An inspiration to Avatar maybe!
pilooch··on Show HN: VectorVFS, your filesystem as a vector database
Using it for a RAG is smart indeed, especially with a multimodal encoder (vision-rag), as the implementation would be straightforward from what you already have.
pilooch··on MCP server for Ghidra
Or just we could forget about code and have model act directly :) That's my bet.
pilooch··on Gemma 3 Technical Report [pdf]
Good, FYI the number one usage is vision RAGs (RAGs that deal with documents as images instead of text).
pilooch··on Gemma 3 Technical Report [pdf]
Someone knows whether there is support for multiple images as input ? I don't see it from the docs yet.
pilooch··on Mistral OCR
Fun, but LLMs would follow them post OCR anyways ;)

I see OCR much like phonemes in speech, once you have end to end systems, they become latent constructs from the past.

And that is actually good, more code going into models instead.

pilooch··on Mistral OCR
But what's the need exactly for OCR when you have multimodal LLMs that can read the same info and directly answer any questions about it ?

For a VLLM, my understanding is that OCR corresponds to a sub-field of questions, of the type 'read exactly what's written in this document'.

pilooch··on Zelensky leaves White House after angry meeting
Because you don't overrun a nuclear state with weapons, but with influence and the true promise of scaling up.
pilooch··on Benchmarking vision-language models on OCR in dynamic video environments
The question is what is OCR for ? If it's to answer questions and work with a document, then VLMs do actually contain self correcting mechanisms. That is, the end to end image + text input to text output is statistically grounded, by training. So the question to ask is what do you need OCR for ? Fedding an LLM? Then feed it to the VLM instead. Some other usage ? Well, to be decided. But near now, CTX and lstms are done with, because VLMs do everything: finding the area to read, reading, embedding, and answering. OCR was a mid-step, it's going away.
pilooch··on Scaling up test-time compute with latent reasoning: A recurrent depth approach
It could be argued that "thinking" / CoT in latent space abstracts away the language issue, and that in fact language in reasoning steps doesn't matter. Latent tokens could actually be decoded afterwards to any target language. Much more powerful IMO.

On a side note, there's decent research on how well bilingual humans do actually think in both language, and are actually better at decisive thinking outside of their mother tongue.

pilooch··on Why LLMs still have problems with OCR
It's good and useful to see empirical analyses like this. I use open & custom VLMs a lot. The point of VLMs is that OCR is not needed anymore: it's intrinsic to the model. For instance at work we've developed a family vision-based RAG, and it's performance is twice that of a text-based one. The point I'd like to make here is that OCR is an intermediate step that is not explicitly needed anymore, un many cases. My hunch is that pure OCR will go away.
pilooch··on OpenAI Furious DeepSeek Might Have Stolen All the Data OpenAI Stole from Us
Any ML based service with an API is basically a dataset builder for more ML. This has been known forever and is actually a useful "law" of ML-based systems.
pilooch··on DeepSeek releases Janus Pro, a text-to-image generator [pdf]
Sure but it's good to recognize Meta never stopped publishing even after Openai and deepmind most notably stopped sharing the good sauce. From clip to dinov2 and llama series, it's a serious track to be remembered.
pilooch··on Don't use cosine similarity carelessly
Statistically you want the retriever to be trained for cosine similarity. Vision LLM retriever such as DSE do this correctly. No need for reranker once done.
pilooch··on All You Need Is 4x 4090 GPUs to Train Your Own Model
I do this for many application. 2 to 4 RTXA5000 do the job (Lora finetune). As for dataset, depending on your task, you need image / text pairs.
pilooch··on Microsoft acquires twice as many Nvidia AI chips as tech rivals
Google has custom made TPUs.
pilooch··on Parkinson's Law: It’s real, so use it
Opportunity to say that other Parkinson's papers, and especially shorts and opinions that can be found in a book are both scientifically interesting and hilarious. A marvelous one is about importance of people and their physical trajectories in cocktail parties. A wonderful mix of social and economical sciences.
pilooch··on PaliGemma 2: Powerful Vision-Language Models, Simple Fine-Tuning
Paligemma proves easy to train and useful in fine-tuning. It's main drawback was not being able to handle multiple images without being partly retrained. This new version dies not seem to support multiple images as input at once. Qwen2vl does. This is useful for vision rag typically.
pilooch··on QwQ: Alibaba's O1 Like Reasoning LLM
I don't see deeper technical details nor how to control the sampling depth. Has anyone found more ?
pilooch··on Ask HN: What Are You Working On? (October 2024)
A custom email sorter / spam filter that uses a fineruned multimodal LLM: my emails are turned into images (turning them / extracting html then rendered with selenium) and passes to the vision LLM. I went from ~200 to ~15 useful emails a day.

Raw code is here: https://gitHub.com/beniz/llmbox

All runs locally, the finetune is a plaigemma-3b (from mix-448). Acc/F1/prec/recall are all within 99.99%, including on llm-generated spam.

pilooch··on I Am Tired of AI
By AI here, it is meant generative systems relying on neural networks and semi/self supervised training algorhms.

It's a reduction of what AI is as a computer science field and even of what the subfield of generative AI is.

On a positive note, generative AI is a malleable statiscally-geounded technology with a large applicative scope. At the moment the generalistic commercial and open models are "consumed" by users, developers etc. But there's a trive of forthcoming, personalized use cases and ideas to come.

It's just we are still more in a contemplating phase than a true building phase. As a machine learnist myself, I recently replaced my spam filter with a custom fineruned multimodal LLM that reads my emails a pure images. And this is the early early beginning, imagination and local personalization will emerge.

So I'd say, being tired of it now is missing much later. Keep the good spirit on and think outside the box, relax too :)

pilooch··on I Am Tired of AI
It's intended as a joke and a demonstration no ? Like this is exactly the type of text and words that a commercial grade LLM would never let you generate :) At least that's how I got that comment...
pilooch··on Forget ChatGPT: why researchers now run small AIs on their laptops
I run a fineruned mulmodal LLM as a spam filter (reads emails as images). Game changer. Removes all the stuff I wouldn't read anyways, not only spam.
pilooch··on Visualizing Weather Forecasts Through Landscape Imagery
Can't wait for the stable diffusion version :)
pilooch··on Hezbollah pager explosions kill several people in Lebanon
They could have seen it while fixing a pager... Or those things work so well, never break...
pilooch··on Show HN: Tune LLaMa3.1 on Google Cloud TPUs
Indeed, a lora finetune of llama 3.1 8B works on a single 24GB GPU and takes from a few hours to a few days depending on the dataset size.
pilooch··on RLHF is just barely RL
Yes, same for maths. As long as a true reward 'surface' can be optimized. Approximate rewards are similar to approximate and non admissible heuristics,search eventually misses true optimal states and favors wrong ones, with side effects in very large state spaces.
← PreviousPage 2 of 11Next →