HNHacker News
TopNewBestAskShowJobs

rom1504

68 karma · joined May 30, 2015

[ my public key: https://keybase.io/rom1504; my proof: https://keybase.io/rom1504/sigs/C3OLSKGL50IJtOmw2w7TmG8p3GPTI61C44e2qX8KwWk ]
submissionscomments
rom1504··on Why Software Factories Fail (or: harness engineering is not enough)
Why is a very good question indeed! My guess is we have not figured out how to include long term enough tasks into the RL training.

As to whether they can today, it's pretty easy to figure out that they can't by observing the difference between asking for very high level things like "please make me a website that will be very successful" vs sending it a detailed plans on the goals and many of the specifics and then holding it accountable towards the planning.

Doing appropriate project structure and trying to optimize code quality are definitely good methods to try and increase the chance that the project will succeed long term but I would say they are only methods towards the long term goal and the LLM need to be thinking of the long term goal and how to plan towards that to achieve a much wider and subtile range of methods to achieve long term goals.

rom1504··on Why Software Factories Fail (or: harness engineering is not enough)
I think the simple answer is LLM cannot do long term planning.

Maintainability means thinking what will happen to this code in the next year and possibly hundreds of changes and based on that selecting the right abstractions that will work in the long term.

LLMs currently cannot do long term planning and specifically cannot pick abstractions that will work in the long term.

I think that's the important problem to solve. Teaching LLMs how to pick the right abstractions.

rom1504··on Ask HN: What happens to James Halliday ( Substack)?
yeah it's sad
rom1504··on Same.energy: Image Search by Similarity
Hehe, well you know, PR welcome, the front end is 500 lines https://github.com/rom1504/clip-retrieval/blob/main/front/sr...

Other people have done a few alternate front ends already

This one is meant to be functional, but could sure be made prettier

rom1504··on Laion-5B: A new era of open large-scale multi-modal datasets
yes indeed. Video is the clear next step.
rom1504··on Laion-5B: A new era of open large-scale multi-modal datasets
Looks like you missed the whole point of this dataset.

The idea that we proved is you can get a dataset with decent caption and images (that do match yes, you can see for yourself at https://rom1504.github.io/clip-retrieval/ ) that can be used to trained well performing models (eg openclip and stable diffusion) while using only automated filtering of a noisy source (common crawl)

We further proved that idea by using aesthetic prediction, nsfw and watermark tags to select the best pictures.

Is it possible to write caption manually? sure, but that doesn't scale much and won't make it possible to train general models.

rom1504··on Exploring 12M of the 2.3B images used to train Stable Diffusion
Done https://github.com/rom1504/clip-retrieval/commit/53e3383f58b...

Using clip for searching is better than direct text indexing for a variety of reasons but here for example because it matches better what stable diffusion sees

Still interesting to have a different view over the dataset!

If you want to scale this out, you could use elastic search

rom1504··on Exploring 12M of the 2.3B images used to train Stable Diffusion
This is due to the aesthetic scoring in the UI. Simply disable it if you want precise results rather than aesthetic ones.

It works for your example

I guess I'll disable it by default since it seems to confuse people

rom1504··on Exploring 12M of the 2.3B images used to train Stable Diffusion
Hi, laion5b author here,

Nice tool!

You can also explore the dataset there https://rom1504.github.io/clip-retrieval/

Thanks to approximate knn, it's possible to query and explore that 5B datasets with only 2TB of local storage, anyone can download the knn index and metadata to run that locally too.

Regarding duplicates, indeed it's an interesting topic!

Laion5b deduplicated samples by url+text, but not by image.

To deduplicate by image you need to have an efficient way to compute whether image a and b are the same.

An idea to do that is to compute an hash based on clip embeddings. A further idea would be to train a network actually good at dedup and not only similarity by training on positive and negative pairs, eg with triple loss.

Here's my plan on the topic https://docs.google.com/document/d/1AryWpV0dD_r9x82I_quUzBuR...

If anyone is interested to participate, I'd be happy to guide them to do that. This is an open effort, just join laion discord server and let's talk.

rom1504··on Large-Scale Artificial Intelligence Open Network
Yeah we have regular contact with many EAI members. EleutherAI is pretty great at open research.
rom1504··on Large-Scale Artificial Intelligence Open Network
Let me start by saying that laion is a non profit, open to anyone that want to contribute.

Agreed about the website css. Do you want to contribute?

What's the problem with the dataset name exactly? Seems to work pretty well.

Yes the dataset is an extract of common crawl, this is an accessible to all method to produce valuable dataset. This is unlike supervised dataset which are reserved to organization with millions of dollars to spend on annotation and do not scale.

Non annotated datasets are the base of self supervised learning, which is the future of machine learning. Image/text with no human label is a feature, not a bug. We provide safety tags for safety concerns and watermark tags to improve generations.

It also so happens that this dataset collection method has been proven by using laion400m to reproduce clip model. (And by a bunch of other models trained on it)

rom1504··on Laion-400M: open-source dataset of 400M image-text pairs
a safe mode is now in place
rom1504··on Laion-400M: open-source dataset of 400M image-text pairs
That's interesting indeed! Note that the description search as well as the image search are both using a knn index on embeddings and not exact search. That helps for finding semantically close by items but indeed for exact reference match it might not be the best solution. Re-indexing the dataset with something like elastic search would give the reference search results you expect.
rom1504··on Overview of methods and softwares for computer vision
I wrote an overview of what methods and softwares are available for computer vision. Computer vision is really powerful these days !
rom1504··on Enable Node.js to Run with Microsoft's ChakraCore Engine
For example the chakra engine has a better coverage of es6 than v8. see edge score there https://kangax.github.io/compat-table/es6/
rom1504··on 2^74207281-1 is Prime
So I tried it and it's 72135 bytes using tar czf (created by `fs.writeFileSync('/tmp/file',new Buffer(74207281).fill("1"),'utf8')`)

Not that it's the most efficient representation at all, 74207281 fits in 4 bytes.