135 karma · joined August 1, 2016
> Once we have a source-coding scheme, we can "invert" it to get conditional probabilities; we could even sample from it to get a generator. (We'd need a little footwork to deal with some technicalities, but not a heck of a lot.) So something I'd really love to see done, by someone with the resources, is the following experiment:
> - Code up an implementation of Lempel-Ziv without the limitations built in to (e.g.) gzip; give it as much internal memory to build its dictionary as a large language model gets to store its parameter matrix. Call this "LLZ", for "large Lempel-Ziv".
> - Feed LLZ the same corpus of texts used to fit your favorite large language model. Let it build its dictionary from that. (This needs one pass through the corpus...)
> - Build the generator from the trained LLZ.
> - Swap in this generator for the neural network in a chatbot or similar. Call this horrible thing GLLZ.
> In terms of perplexity, GLLZ will be comparable to the neural network, because Lempel-Ziv does, in fact, do universal source coding.
Maybe someone on HN will have resources for such an experiment?
> However, the regulator hit back, saying: "It is the CMA's job to do what is best for the people, businesses and economy of the UK, not merging firms with commercial interests."
> But the Competition and Markets Authority (CMA) said its job was not to serve the interests of merging firms.
The CIFAR experiments I mentioned were https://arxiv.org/pdf/1806.00451.pdf. It doesn't contain this argument (unfortunate wording) but appears to support it well.
(Sorry, can’t resist.)
For more complex equations people use line breaks and indentations, and/or macros.
The point of my original post is that the asymptotics break down here, and this phenomenon is not poorly understood, at least in some other communities. It is not meant to provide an alternative that is always well-defined and useful, although as I said in the grandparent comment, there is the useful implication that you can stay safe by sticking to the asymptotic regime.
> On analysis of the file, I found that the vast majority of the records are actually related to sex, porn, and other smartphone brands. There are mentions of Tibet, Hong Kong, and other religious groups, however, mentions of the CCP and “China” are also included, too
> I think it’s pretty clear that the filter is specifically used for filtering advertisements
Personally I find it more comfortable to use TeXmacs for quick/throwaway notes, and macros in longer documents for better readability.
It definitely helps that he has a lifetime of research experience (in another field).
http://dustintran.com/blog/a-research-to-engineering-workflo...
This is nuanced: in some cases it’s omission for brevity [1] and in other cases it’s how the phonemes sound 100 years ago [2]. And if you pronounce them fast enough there isn’t much difference.
The omitted diaeresis is certainly for brevity, although using v is more common to me.
[1]: https://www.zhihu.com/question/26010099 [2]: https://www.zhihu.com/question/313646560
> Vastaamo board fires CEO, says he kept [a previous] data breach secret for year and a half
> [The CEO's family] owned the company until it was bought by Helsinki-based private equity firm Intera Partners in May 2019, not long after a second breach of the psychotherapy provider’s data security systems.