HNHacker News
TopNewBestAskShowJobs

baptiste1

16 karma · joined August 31, 2023

submissionscomments
baptiste1··on Is supervised learning dead for computer vision?
Thanks, James for your insights.

Your library looks nice.

You are right some computer vision systems do have real-time requirements and do need to be run on the edge. It is in the current roadmap of Datasaurus. I would like to capture logs of API calls and foundational model responses so that they can be later used as training data for smaller models.

baptiste1··on Is supervised learning dead for computer vision?
I used LLaVA. Unfortunately, I signed a NDA :( so I cannot share the code and the data is private. We fine-tuned it with example images, labels, and text prompts. We also tried in-context learning. Indeed, the prompt was static but we could do data augmentation and provide a series of equivalent prompts. We just used the prompt that gave us the best performance during initial model testing with in-context learning. I am unsure if the existence of equivalent prompts creates instability because a sentence with the same meaning should be quite close in the latent space of the foundation model so it understands them in a similar manner.
baptiste1··on Is supervised learning dead for computer vision?
Thanks for your comment. I agree, I think both method are quite complementary
baptiste1··on Is supervised learning dead for computer vision?
Thanks for your question. You are right, current vision-language foundation models are quite heavy. However, for example in NLP there are some works on smaller foundation models. In addition, you could also use a foundational model to help train your smaller model or label more data.
baptiste1··on Is supervised learning dead for computer vision?
Yes, you can. The model that I was talking about LLaVA only output text but other models such as SEEM (https://github.com/UX-Decoder/Segment-Everything-Everywhere-...) outputs a segmentation map. You could prompt the model "Where is the pickleball in the image?" and get a segmentation map that you could then use to compute its center. Please let me know if you would be interested to have SEEM available in Datasaurus
baptiste1··on Is supervised learning dead for computer vision?
Yes, true the fine-tuning is not new and indeed I also view it as "starting with an incredibly well-initialized network"

However, the promotable aspects of those vision models are completely new. You can define your tasks at runtime and steer the model behavior. I think this makes it easier and faster to insights from your images. Lastly, those models are trained on a lot of different tasks compared to previous models that were general classifiers and that could then be trained on a specific domain. This allows them for example to be reused in an organisation and prevents you from creating multiple task-specific models

baptiste1··on Is supervised learning dead for computer vision?
Yes, your understanding is correct. However, instead of adding a head on top of the network, most fine-tuning is currently done with LoRA (https://github.com/microsoft/LoRA). This introduces low-rank matrices between different layers of your models, those are then trained using your training data while the rest of the models' weights are frozen.
baptiste1··on Is supervised learning dead for computer vision?
Yes, true. Indeed, your phrasing is better! I am not sure however if I can change the title now. I will keep your comment in mind for the future.
baptiste1··on Is supervised learning dead for computer vision?
Thanks for your comment.

I did not know about "Betteridge's law of headlines", quite interesting. Thanks for sharing :)

You raise some interesting points.

1) Safety: It is true that LVMs and LLMs have unknown biases and could potentially create unsafe content. However, this is not necessarily unique to them, for example, Google had the same problem with their supervised learning model https://www.theverge.com/2018/1/12/16882408/google-racist-go.... It all depends on the original data. I believe we need systems on top of our models to ensure safety. It is also possible to restrict the output domain of our models (https://github.com/guidance-ai/guidance). Instead of allowing our LVMs to output any words, we could restrict it to only being able to answer "red, green, blue..." when giving the color of a car.

2) Cost: You are right right now LVMs are quite expensive to run. As you said are a great way to go to market faster but they cannot run on low-cost hardware for the moment. However, they could help with training those smaller models. Indeed, with see in the NLP domain that a lot of smaller models are trained on data created with GPT models. You can still distill the knowledge of your LVMs into a custom smaller model that can run on embedded devices. The advantage is that you can use your LVMs to generate data when it is scarce and use it as a fallback when your smaller device is uncertain of the answer.

3) Labelling data: I don't think labeling data is necessarily cheap. First, you have to collect the data, depending on the frequency of your events could take months of monitoring if you want to build a large-scale dataset. Lastly, not all labeling is necessarily cheap. I worked at a semiconductor company and labeled data was scarce as it required expert knowledge and could only be done by experienced employees. Indeed not all labelling can be done externally.

However, both approaches are indeed complementary and I think systems that will work the best will rely on both.

Thanks again for the thought-provoking discussion. I hope this answer some of the concerns you raised

baptiste1··on Is supervised learning dead for computer vision?
Thank you for your comment. Indeed, you are right not every company has terabytes of data to train their model. I like your example "Another company needs to detect loose screw heads in engine blocks”.

I actually got the idea for Datasaurus because of a similar problem. My brother wanted to check if sheets of metal were bent and needed to be rejected in a production line setting. However, he did not have any data and could maybe annotate a couple of images manually but not create a full dataset. We tested the fine-tuning approach and he was able to have good results in a couple of minutes.

That’s why I think this could be quite valuable and I decided to package it into an open-source application.

baptiste1··on Is supervised learning dead for computer vision?
Foundational models are generally trained on internet scale level of data. They have seen billions of images, so they would have seen some medical images. For example, extracted from public datasets or textbooks. However, indeed, they may not be specialized to your use case. You could still fine-tune the model with a couple of examples to be more tailored to what you desire. Having a foundation model does not exclude training and your data could still be valuable. Indeed, you could achieve better performance by fine-tuning the larger model than just using your training data alone to train a model from scratch.

Also for the medical domain, I think vision-text segmentation models like SEEM (https://github.com/UX-Decoder/Segment-Everything-Everywhere-...) are really cool. You could for example ask “Where is the tumor located on that image?” and then the tumor is highlighted in the picture.