Tokens, n-grams, and bag-of-words models (2023)
zilliz.com
zilliz.com
It's making me wonder, why are models usually in Python? Could these models be implemented in say, Scala, Kotlin, or NodeJS, and have there been attempts to do so?
Python is widely adapted not really because the language itself is any good or particularly performant (it's not at all), but because it presents easy to use wrapper APIs to developers who may have a poor background in Computer Science, but are rather stronger in statistics, general data analysis, or applied fields like economics.
MATLAB is a proprietary computing platform with its own language that was very widely used (and probably still the standard in some fields of engineering). The fact that it is proprietary and the language is not great as a general programming language were significant drawbacks to wide adoption.
As far as ML is concerned, the deep learning revolution happened when Python was the dominant language for scientific computing (mostly due to NumPy and SciPy), so naturally a lot of the ecosystem was built to have Python as the main (scripting) language. The rest is history.
As far as "attempts to recreate similar ecosystem":
PyTorch (currently the most popular deep learning framework), was originally Torch (initial release: 2002 - long before "Deep Learning" was a thing) with Lua as the scripting language. Python's momentum in 2010's meant that eventually it was rewritten in Python thus becoming PyTorch.
Julia language is a famous somewhat recent example (first release 2012, stable 1.0 in 2018) of a language that was built partially to address some of Python's shortcomings and to "replace" it as the default for scientific computing. It didn't succeed - it's hard to move people away from an ecosystem with so much head start and momentum as Python had in the 2010s.
I don't know if that is a handicap to the idea you mention, but it may be...perhaps dropping into C for hardware performance and then back up to the JVM is just too annoying since things representable in hardware had no non-super-tedious representation back in the JVM (Scala or Java in my case).
This is incredibly important in an environment where you are exploring datasets and generating new hypotheses.
It's used in the context of multiple "words" as in this article, but also tuples of characters. The confusion stems from the fact that "token" may refer to words as is used in this article (this is common in modern NLP) but older sources often use the word token to refer to graphemes ("letters") and various other breakdowns. This wikipedia article[1] is a good example of such usage.
Afaik this is pretty canonical.
Of course. Python itself could be reimplemented in Scala or Kotlin.
I'm not trying to be snarky, but my teenage self didn't realize all languages could do all things when I was young, so I'm speaking to anyone else out there who might not have learned this.
It's just a matter of ease. Python is easy for short programs (although maintainability can suffer on large programs), and, for better or worse, Python has become the most popular language in the world, especially for numerical computing.
Going forward, I assume we'll see more Python interfacing with more Rust as well. (Polars is one prominent example.) Python is just a fantastic UI over “close to the metal” languages.