Show HN: Simply Reading Analog Gauges – GPT4, CogVLM Can't
huggingface.co
huggingface.co
Synthetic data showing relative gauge positions and a wide variety of contexts anchored to a particular category, like "dial gauge showing 31%" is the solution, just like watch images showing a full range of perspectives and times is the solution to the 10:10 problem.
I think the trickiest problem in transformer models is identifying situations like this where there's a mode collapse, but it's one or more degrees of separation from being apparent. This can lead to generation being technically and structurally correct (like a beautifully rendered wristwatch,) but the range of available answers will be limited to whatever the model has collapsed on (i.e. all watches and clocks only show one time, even if you ask it to display 4:59 PM and construct a perfect prompt about something right before work is done for the day.)
These are subtle distortions in the world model, but things like the 10:10 problem likely also affect other concepts that overlap with clocks and watches, like analog gauges. Maybe there are patterns in the model that can be associated with these types of issues, and using noise and synthetic data can target and unstick a model?
Nobody wants to upgrade when the current system still works and can be supported (however painfully).
Put this on a 'spot' robot from Boston dynamics and you have an operator that won't be hurt when there is a steam leak, or whatever.
There are definitely use cases.
The question is how much you trust this vision system, not whether or not you still need a guy to do rounds.
I am suspect that the training data is good. Many tasks highly depend on labeled datasets and unfortunately these need to be very large (given current ML methods). The major problem is that these labels are generated from cheap or free labor where labelers are also expected to know the nuance and label appropriately. Nearly (if not every) dataset you've seen has major errors in them that matter. Honestly, one of the only ways to combat this is be insistent on viewing benchmarks as guides. There's been countless times I've demonstrated to others how I can perform worse the training dataset but absolutely demolish results when using customer datasets or real world data (i.e. generalization). Remember, it is absolutely possible to overfit your model even if your validation accuracy does not decrease. There are many ways to overfit and evaluation is unfortunately substantially harder than looking at results on a test set. There's a lot of complexity to this and if you want an effective product you should probably have at least one person on your team that cares about statistics and will nitpick stuff (though you can also over-saturate your team with these people too. I say that as one of these people).
Edit: Actually I believe every example shows an incorrect reading.
[0] https://synanthropic-reading-analog-gauge.hf.space/--replica...
> What do you think is the most important requirement?
Data quality and augmentation. I looked at the dataset[0] and the quality is very poor. There is very low diversity and it appears that it is composed of a few images and increased through augmentation. Augmentation should not be handled through dataset generation but dataset processing. Torch will even move bounding box keypoints when cropping or resizing images.
Data quality is one of the most important factors and I know it is common to just get labels to Turk or cheap services but you absolutely get what you pay for.
[0] Please don't harvest emails and usernames to require downloading. I created a dummy account for this because that's absolutely not acceptable. DO NOT SEND PEOPLE SPAM. DO NOT HARVEST EMAILS AND SELL THOSE EMAILS. I happily flag emails that come to my account through means like this until Google identifies all emails from that sender as spam. DO NOT DO THIS. It will just make people hate you and build a devoted base against you. There is no justification for this. You should absolutely rescind the gate and delete all emails you harvested. Your licensing is already an agreement. Don't be shady.
EDIT: ah, it's for the other types of gauges. Errors out on mine :)
I wanted to record speed, rpm and acceleration to graph that and show at the repair station.
At one visit in a authorised repair station the mechanic couldn't diagnose it (unrelated heavy snowing conditions).
This sounds problematic and like it explains why the first example image is improperly labeled. There is a general form to dial gauges and the results make me suspect of overfitting. Data diversity is essential. Quality of data matters more than quantity (though you still need quantity).
I am not sure if it even has the right foundation data set to even understand the myriad of gauges that exist in the world, so I doubt fine tuning will fix anything.
LLMs trained in Internet data are Internet simulators.
How often have you asked the Internet to read an analog gauge for you? Probably never.
This could also be a way to train a bot to do grandpa's job.