HNHacker News
TopNewBestAskShowJobs

iflp

135 karma · joined August 1, 2016

submissionscomments
iflp··on Scalable watermarking for identifying large language model outputs
The seed is computed using the last 4 tokens, but to get the predictive distribution they should still need the full history?
iflp··on First order transition in [LK-99] containing Cu2S
> The superconducting-like behavior in LK-99 most likely originates from a magnitude reduction in resistivity caused by the first-order structural phase transition of Cu2S. [...] It is important to note that this first-order structural transition differs significantly from the second-order superconducting transition.
iflp··on Plain old gzip+kNN outperforms BERT and other DNNs
This seems to be about easier classification tasks with not too many samples, for which TF-IDF also works well (Table 3). But more generally gzip for text modeling might make sense. Quoting http://bactra.org/notebooks/nn-attention-and-transformers.ht... :

> Once we have a source-coding scheme, we can "invert" it to get conditional probabilities; we could even sample from it to get a generator. (We'd need a little footwork to deal with some technicalities, but not a heck of a lot.) So something I'd really love to see done, by someone with the resources, is the following experiment:

> - Code up an implementation of Lempel-Ziv without the limitations built in to (e.g.) gzip; give it as much internal memory to build its dictionary as a large language model gets to store its parameter matrix. Call this "LLZ", for "large Lempel-Ziv".

> - Feed LLZ the same corpus of texts used to fit your favorite large language model. Let it build its dictionary from that. (This needs one pass through the corpus...)

> - Build the generator from the trained LLZ.

> - Swap in this generator for the neural network in a chatbot or similar. Call this horrible thing GLLZ.

> In terms of perplexity, GLLZ will be comparable to the neural network, because Lempel-Ziv does, in fact, do universal source coding.

Maybe someone on HN will have resources for such an experiment?

iflp··on Microsoft in furious attack on UK as major gaming deal blocked
Content aside, this is terrible editing on BBC’s part:

> However, the regulator hit back, saying: "It is the CMA's job to do what is best for the people, businesses and economy of the UK, not merging firms with commercial interests."

> But the Competition and Markets Authority (CMA) said its job was not to serve the interests of merging firms.

iflp··on GPT-4
These are all good reasons, but it’s really a new level of openness from them.
iflp··on A stack of feed-forward layers does surprisingly well on ImageNet
No. I was referring to the "standard concentration bound" in that paper, which applies when you have separate validation and test sets. I think the argument can usually be improved by applying small-variance inequalities such as Bernstein's, to excess risk-like quantities such as l(f_hat(x), y) - l(f_ref(x), y), to show that accuracy difference / relative rank enjoys better guarantees. For ImageNet we can use the 01 loss and set f_{ref} to a SoTA classifier which, while having its loss bounded away from 0, is "mostly similar" to most f_hat's, and thus leads to a small excess risk.

The CIFAR experiments I mentioned were https://arxiv.org/pdf/1806.00451.pdf. It doesn't contain this argument (unfortunate wording) but appears to support it well.

iflp··on A stack of feed-forward layers does surprisingly well on ImageNet
If you only care about identically distributed test data, test set overfitting doesn't happen that fast: if you evaluate M models on N test samples, the overfitting error is on the order of sqrt(log M / N). And even as this error becomes more noticeable, the relative ranks among the models are even more stable, as you can apply the small variance bounds. This is actually verified on models proposed for CIFAR-10.
iflp··on Microsoft's ex-UI chief is shocked about the Windows 11 start menu
StartAllBack
iflp··on I love my GPD Micro PC
surface go?
iflp··on Are language models deprived of electric sleep?
The wake-sleep algorithm is designed to provide generative models with necessary sleep, at least in their developmental stages.

(Sorry, can’t resist.)

iflp··on I hate LaTeX, I love LaTeX
The equation example is artificial. In practice there will be no curly brackets around these single-token sub/superscripts, nor should the \left / \right present in this example. With properly added whitespaces this equation becomes quite legible.

For more complex equations people use line breaks and indentations, and/or macros.

iflp··on An Account of the Shanghai Lockdown
The lockdown policy was successful in these aspects before, but it’s very unclear if it can still work with omicron.
iflp··on Entropy isn't sufficient to measure password strength
It is a random variable in this setting, as it is a function of the randomly generated password. Given a deterministic sequence, you find the definition of its Kolmogorov complexity in textbooks/Wikipedia/etc. By saying the Kolmogorov complexity will disagree with Shannon entropy, I meant the former, which is a random variable here, does not converge to the latter, contrary to the standard asymptotic setting which probably gives people the idea of using entropy to characterize password (I don't know, don't work in security).

The point of my original post is that the asymptotics break down here, and this phenomenon is not poorly understood, at least in some other communities. It is not meant to provide an alternative that is always well-defined and useful, although as I said in the grandparent comment, there is the useful implication that you can stay safe by sticking to the asymptotic regime.

iflp··on Entropy isn't sufficient to measure password strength
Kolmogorov complexity is only unambiguously defined asymptotically, and "asymptotics is merely a heuristic". It is also uncomputable. So, to use entropy arguments for passwords, the only correct way I could think of is to generate long and (elementwise) random passwords.
iflp··on Entropy isn't sufficient to measure password strength
Kolmogorov complexity/entropy is more suitable for this purpose, under the implicit assumption that password crackers don't have tailored prior knowledge and are just enumerating "simple" sequences. It only agrees with Shannon entropy on long ergodic sequences. The author basically constructed an example where the two notions don't agree.
iflp··on A Xiaomi phone might’ve shipped with a censorship list in Europe
https://www.xda-developers.com/xiaomi-secret-blacklist-expla...

> On analysis of the file, I found that the vast majority of the records are actually related to sex, porn, and other smartphone brands. There are mentions of Tibet, Hong Kong, and other religious groups, however, mentions of the CCP and “China” are also included, too

> I think it’s pretty clear that the filter is specifically used for filtering advertisements

iflp··on LaTeX Input for Impatient Scholars
This is interesting and follows the line of [1].

Personally I find it more comfortable to use TeXmacs for quick/throwaway notes, and macros in longer documents for better readability.

[1]: https://castel.dev/post/lecture-notes-1/

iflp··on LaTeX Input for Impatient Scholars
Overleaf and VSCode (with LaTeX workshop) support this.
iflp··on Physics Student Earns PhD at Age 89
> Steiner is not prepared to rest on his laurels. He is currently reworking part of his dissertation for publication and plans to continue his theoretical physics work.

It definitely helps that he has a lifetime of research experience (in another field).

iflp··on Cognition Without Computation
I was about to post the same thing. Self-organisation maps seem a classical computational model to me. If the author’s point was that computational models should be biologically plausible, there are many other examples as well. I’ve never really understood what neuroscientists are talking about…
iflp··on Ask HN: Do you also alternate between super productive and slow days?
The nice thing about the amortisation strategy is that you can e.g. exercise on the slow days, which improves your stamina and thus efficiency in the long run.
iflp··on Jam 80 Cores, 768GB of RAM into E-ATX Case with This Tiny Board
This is 7.2$/mo, at which rate you can rent xeon cores as well.
iflp··on A visual introduction to Gaussian Belief Propagation
Try Bayesian Reasoning and Machine Learning specifically, as these are all about Bayesian reasoning.
iflp··on Ask HN: Whats your ideal PhD workflow
Obligatory link: a research to engineering workflow

http://dustintran.com/blog/a-research-to-engineering-workflo...

iflp··on Show HN: Neural-hash-collider – Find target hash collisions for NeuralHash
You don’t need to have the weights. “Transfer attack” is a thing.
iflp··on Pronouncing non-English names for English speakers
> liu" should be pronounced "liou" and "shui" should be pronounced "shuei"

This is nuanced: in some cases it’s omission for brevity [1] and in other cases it’s how the phonemes sound 100 years ago [2]. And if you pronounce them fast enough there isn’t much difference.

The omitted diaeresis is certainly for brevity, although using v is more common to me.

[1]: https://www.zhihu.com/question/26010099 [2]: https://www.zhihu.com/question/313646560

iflp··on Positive Energy Warp Drive from Hidden Geometric Structures
If you are willing to assume CTC: https://www.scottaaronson.com/democritus/lec19.html
iflp··on Show HN: I built an online interactive course that helps you learn Vim faster
Maybe the problem is that you are not supposed to rely solely on hjkl for navigation? Not sure how others use it, but for me most of the navigation are performed with <C-d> / <C-u> / search (long range) or f, w, b etc (shorter range, usually with modifiers). hjkl are almost always used in combination with these to adjust the initial / final location of the cursor.
iflp··on The Next Decade Could Be Even Worse
Has anyone carefully checked his work? Skimming through his first few papers it seems he just built datasets covering the past several millennia and ran PCA (or models with similar complexity). Fine for explanatory purposes, but not so great if you want to make predictions when most complexity of the society arguably comes from the past few centuries.
iflp··on Therapy patients blackmailed for cash after clinic data breach
From the first two links:

> Vastaamo board fires CEO, says he kept [a previous] data breach secret for year and a half

> [The CEO's family] owned the company until it was bought by Helsinki-based private equity firm Intera Partners in May 2019, not long after a second breach of the psychotherapy provider’s data security systems.

Page 1 of 2Next →