And unlike other nations that might build these tools for internal use only, it's basically certain that this tech will be sold to e.g. the US.
See: NSO Group, Cellebrite, Paragon, Oosto/Anyvision, Corsight. They've all had or have contracts with US domestic agencies that needed surveillance tech.
Chinese (esp. youth) are locked up, there are a lot of stories about migration being denied on a large scale for Han Chinese, including within China itself (not allowed to move cities, like you saw long ago in the Soviet Union)
And yet we hear nothing. Meanwhile there's "social credit score", mass unemployment in the capital city, and expert LLM teams (that have put the best US and EU teams to shame 2 or 3 times now) ...
I think we'll suddenly find out China far ahead of absolutely anyone else in this field.
And it's things like this that are ripe to come to other countries. An interesting book that goes into this concept is "The Palestine Laboratory." Sad times indeed.
This isn't a tool to better understand arabic in some innocent sense, it's a compensation for israelis refusing to learn arabic and having trouble gathering intelligence from linguistic sources. When you're actively involved in the genocide of a people you'll also have problems finding informants among them, it's much easier under a less intense apartheid regime.
So not some journalistic investigative scoop. Its PR !
> building the LLM “focused only on the dialects that hate us”.
Well duh..
> “We had no clue how to train a foundation model,
Immediately followed by typical MIC over spending instead of fine-tuning
> training data eventually consisted of approximately 100bn words
So closer to gpt3 than any recent sota models
> existing .. Arabic models using standard written Arabic.. rather than spoken Arabic
I think this is the only ah-ha item in the paper.