Show HN: Cherry2.0 – text classification without ML knowledge needed
github.com
github.com
> Gamble / Porn / Political / Nomal (model='harmful')
Not sure what Nomal is, but 'Political' is considered harmful? I suppose they mean politics has the potential to be harmful. Would be interesting to see what material it was trained on (a cross section of pro and anti-CCP? pro-CCP would still be political, but then why would that need to be classified under a harmful model?)
> Lottery ticket / Finance / Estate / Home / Tech / Society / Sport / Game / Entertainment (model='news')
Interesting bit here is that Lottery Tickets are included in news and yet gambling is included in harmful. Is the lottery not considered gambling in China? Or gambling just has the potential to be harmful, despite also being news?
It's hard to explain why 'Political' is considered harmful in China. Most of the Chinese companies try to avoid the political text classify problem by using a hashtable to store 'sensitive words'. One example would be we can't search '64g RAM' in Xianyu, which is one of the biggest c2c second-hand online markets developed from Alibaba because 64 indicate June 4th which is a special day in China.
> Would be interesting to see what material it was trained on (a cross section of pro and anti-CCP? pro-CCP would still be political, but then why would that need to be classified under a harmful model?)
Good question. It's quite subjective and no one knows the answer. Companies have to do lots of self-censorships. However, unlike porn or gamble. We don't have a standard here.
> Is the lottery not considered gambling in China?
Lottery ticket is gambling, but in China, we call it 福利彩票 which indicate the money will be used in charity. It's quite different from gambling on football or basketball and Lottery ticket
Yes, that's one way to put it.
Like with AWS Rekognition where you’re billed on raw usage, which makes no sense when what matters is going to be the precision & recall & false detection rate on your data distribution (not whatever Amazon’s team trains it on). How many true positives / false positives / etc. will you get on your data?
ML is uniquely poorly suited to be treated as just some API or just some black box library. I really wish people would stop popularizing footgun approaches to this!