Large-scale vehicle classification
blog.aqnichol.com
blog.aqnichol.com
This seems to be because you scraped new car prices instead of used. The used market has a wild variance in price data, even if viewing only CPO vehicles directly from dealerships. That in itself could be a whole other fascinating week spent training it on that data. It could also be that Kelley Blue Book's data is wrong as always. KBB pumps up the prices on used cars well beyond their actual market value, and I've yet to figure out what benefits they receive from doing so. It might be used dealerships gaming things. A Chevrolet Cobalt regardless of year or location sells for about $3,500 from private sellers, but KBB says they go for $6,600, or almost twice that.
I'd also be interested in seeing market location weighting for the data. I expect vehicles to be more expensive in places like California and New York compared to Florida or Oklahoma for example.
>After some Googling, I found out that this was a limitation of EXT4 when creating directories with millions of files in them.
I've experienced this myself with faulty Flatpaks spamming my drives with log files. And also Blender, interestingly. If you split a frame into six layers (AO, diffuse, glossy, alpha, shadows, Z-depth) and the animation is about sixty six thousand frames, that's three hundred and ninety six thousand images for one forty five minute animation. Often times you split the scene into three pieces; foreground, focus, and background, which triples that output to a million a hundred and eighty eight thousand images. When put into separate folders to separate each scene or shot, and with multiple renders to test or be sent for approval, it's very easy to run into the EXT4 hash table limitation for just three or four animations over twenty minutes.
For a starting point I'd have it look at vehicles from 2012 or newer. This should prevent "classic" or newly collectible vehicles from spiking the data. Secondly, be aware that Craigslist prices on the front of the ad are often listed as fake numbers such as "0$" or "12345$" with the actual soft or hard price contained within the ad's body of text. Facebook Marketplace can be much worse about this, and much more manipulative with the number of fake listings, so check for duplicates or suspiciously low prices to prevent spiking the data. eBay Motors has a quirk shared with eBay, where they have both an auction price and a Buy It Now price. The only reliable way to gauge prices there would be using Buy It Now listings.
If op is here, I would love to hear your thoughts - and which was more accurate.
(Also I could be wrong but the Audi is an A5/S5 and the model should be able to predict it’s not a A4/S4 quite easily from the rear view?)
I'll know more in a week or two once all the models have converged, and can update here!
A possible commercial application is: evaluating shopping mall customers demographic by parking lot. Real estate market shifts (possible hypothesis is vehicles values go up before commercial or residential property values). At which point, you would need to include used cars in the database.
Second, I realize make and model are the best indicators, but as interested in vehicle design and history, is there anything about the innate design or shape that looks “expensive” or looks “modern”/current. Are there patterns here that without knowing the badge? Related, do the new BMW’s look any more expensive than a Genesis if you didn’t know it was a BMW? If there was a way to remove model/make from the model, what would it say as brands that look more expensive than they are and vice versa? Which brands’ design language looks old?
I would hypothesize this is an artifact of the training data being photos taken for the purpose of selling the cars.
As a seller, it doesn’t make a lot of sense to take a photo of the middle of the car, unless a) you’re trying to hide something or b) you’re being lazy (implying the car isn’t worth much of your time). I would imagine that better photography in general - composition, framing, lighting, setting - would be correlated with higher car prices.
1. Ford F-150
2. Chevrolet Silverado
3. Ram 1500
It is a bit mindblowing to me that these are so popular. Smart of their manufacturers to evolve them into luxury vehicles with thick profit margins.
One explanation is that price labels are super noisy. If there is enough noise in the primary labels, you could imagine that adding in the more predictable target variables could help reduce gradient noise and speed up training. That's my current hypothesis, but I'm very open to others. If I had more time I'd try to do more experiments on this.