HNHacker News
TopNewBestAskShowJobs

mcyc

355 karma · joined November 11, 2020

@me if you wanna talk about Korean + computers.

Also interested in high school cs education

Bsky: https://bsky.app/profile/mcognetta.bsky.social

Mastodon: https://sigmoid.social/@mc

Twitter: @marco_computers

Site: theoreticallygoodwithcomputers.com

meet.hn/city/35.6768601,139.7638947/Tokyo

submissionscomments
mcyc··on Doing a Machine Learning PhD While Working in Japan
Ah, too bad we never crossed paths!

Thanks for your comments here. I agree that a lot of labs were pretty unserious. One other thing that I should have mentioned in the article is that labs in Japan are like little contracting businesses, and the professors have to find funding in a way that seems much harder than the US. This leads to a lot of variability in the financial situations of labs, so yeah my concerns about funding/etc were probably understated (and my lab was likely one of the best funded in the nation).

As for domestic conferences, typically my lab used them as staging for their papers that they would send to top international conferences, so the quality was still pretty good. I do get the sense that the quality outside of the T1 labs drops off pretty quickly.

> ... but I never collaborated with them (in fact, our lab has hardly ever published a paper with >2 authors).

This was something I had in the original draft of the article, but it was removed for being a bit niche. TiTech (and perhaps other schools, but not universal) had a restriction in place where only one student could claim primary authorship on a paper for the purposes of their graduation requirements (typically 1 conference + 1 journal paper). Even if two students contributed equally, they had to sign a form designating one of them the "true" primary author. This massively discourages collaboration. I absolutely hated this policy, but there was basically nothing I could do at an institutional level to change it. I ended up being able to collaborate pretty freely, since I reached my publication requirements quickly, but this policy really hurts students who do their PhD in Japan just because of the numbers game. If two candidates are up for a position, and one has N papers (all first author, but 2-3 authors max) and the other has N + 10 papers (with <N first author, but all with huge groups), the one with N papers is at a disadvantage at first glance. And it just removes one of the primary skills you need to cultivate during a PhD: collaborating with other researchers.

mcyc··on Doing a Machine Learning PhD While Working in Japan
Hmm, the tuition is extremely low relative to the US. On the order or $4-5k USD/year. I am not sure how many additional fees (like facilities, etc.) there are, but I guess not many. Being self funded was common for Japanese students. AFAIK no international students that I knew were self funded, up to the tail of their graduate program when their MEXT ran out.

Now that I think about it, one thing I am not sure about is if self-funded students' travel/conference fees were covered. I would guess they are covered through a grant and it is just the tuition that is self funded, but I do not know for certain. Going to >1 international conference a year would be far more than the tuition.

So, I would guess it would cost ~10k/year to be fully self funded. You can recoup some of that cost working (on a student visa this is limited though), and your lab may also have an RA stipend (but this is both very low [<$20/hr] and counts towards your 28 hour/week limit).

My advice though is to never do a self-funded PhD. Being a PhD "student" is a misnomer. You are a research employee and should be compensated as such.

mcyc··on Doing a Machine Learning PhD While Working in Japan
Yeah, it is definitely a harder limit for international students for the reason you stated (but this article is geared towards foreigners after all).

What field were you in? I am surprised to hear that you think Japan's lab culture is much better than the US. I do think it varies a lot by lab, but overall the internal collaboration and guidance (if not from the PI, from senior members) was much better in American labs than in Japanese labs in my experience.

mcyc··on Doing a Machine Learning PhD While Working in Japan
Thanks!

My decision to start my PhD in Korea was under very different circumstances. I was studying a totally different field (automata theory, so closer to mathematics than ML) where longer PhDs were normal, I was used to the US system, I wasn't working or married, etc. I expected it to take ~6 years when I started.

I basically had to choose between living in the Canadian countryside, Helsinki, or Seoul in order to be in a lab studying what I wanted. I was already pretty into Korean culture (I minored in Korean in college, studied abroad there, and was a serious StarCraft player) so that choice was easy!

When I started looking again, I knew I wanted a shorter PhD (especially since I had a masters and industry + industry research experience), so Japan was pretty attractive.

I will admit that I think 3 years is a little short to make substantive progress outside of machine learning, so unless you come in with a lot of research experience and a good idea of what you want to work on, I think a longer PhD would be beneficial, but the opportunity cost is so high that it really doesn't seem like a great option on paper.

mcyc··on Doing a Machine Learning PhD While Working in Japan
Thank you! And, re: 3 years, that's good to know, I should have been a bit more precise.

What field were you supervising?

3 years in CS is, imo, a little tight but not bad. But I always wondered how people in other fields (stem and otherwise) did it. CS simply doesn't have to deal with the extraneous factors that come with, e.g., biology or climatology.

mcyc··on Doing a Machine Learning PhD While Working in Japan
I am the author of this. Happy to discuss!

One thing to note is that, after I submitted the final draft of this article, the Japanese government ended the MEXT university track or is at least planning to. There was a memo, but the status is not clear.

Anyway, I had a great time doing my PhD this way, and generally I recommend that people work before a PhD and try to work during it (if you are in CS). It gives you a lot of freedom and, at least for me, it is very helpful to have more than one pipeline of work at a time so that when one slows down (as it inevitably does in research), you can draw on the other.

mcyc··on Prove you're human by winning a claw machine
Lichess has a checkmate captcha that I think is cute.

It requires you to solve a mate-in-one puzzle to, e.g., post on the forums.

(Sorry, don't have a better link, there wasn't any non-technical I could find about it).

https://www.reddit.com/r/chess/comments/q19wgq/til_lichess_d...

mcyc··on Rio 3.5 Open 397B – from Rio de Janeiro's city government
Yeah, it is interesting to me that it is coming from the _city_'s government. I've seen sovereign AI things at the country level, but this is the first municipal one I have seen.
mcyc··on AMÁLIA and the future of European Portuguese LLMs
You are right about most tokenizers being heavily biased towards English, but the situation is not so bad for Portuguese. Here are some results on the Goldfish corpus [1] with a few different tokenizers. This measures #characters in corpus / #subwords in tokenized corpus.

```

Llama3

english, 0.216

portuguese, 0.285

italian, 0.287

greek, 0.592

```

```

Gemma4

english, 0.219

portuguese, 0.246

italian, 0.249

greek, 0.537

```

```

Kimi2.6

english, 0.214

portuguese, 0.310

italian, 0.308

greek, 0.716

```

Portuguese is worse than English certainly, but it is on par with Italian (which I think has more overlap with English) and much better than Greek (since it doesn't use the Latin script and is definitely not prioritized in the tokenizer construction).

On your second point, tokenizer transfer allows for extending/modifying a tokenizer without retraining the model from scratch. The simplest version of this is tokenizer extension + continual pretraining, where you just add a bunch more tokens to the vocab for the language/domain that you want to improve and train a little more. It's been done for Japanese [2] and Indic languages, but afaik not Portuguese.

So I think that continual pretraining for a large base model would have probably been fine for this case with huge cost savings. But it is good to have the ability to train your own base models, so I don't think this is such a bad idea.

-----------------------

[1]: https://huggingface.co/datasets/goldfish-models/fish-food

[2]: https://arxiv.org/abs/2404.17790

mcyc··on LLM Structured Outputs Handbook
This is a fantastic guide! I did a lot of work on structured generation for my PhD. Here are a few other pointers for people who might be interested:

Some libraries:

- Outlines, a nice library for structured generation

  - https://github.com/dottxt-ai/outlines
- Guidance (already covered by FlyingLawnmower in this thread), another nice library

  - https://github.com/guidance-ai/guidance
- XGrammar, a less-featureful but really well optimized constrained generation library

  - https://github.com/mlc-ai/xgrammar

  - This one has a lot of cool technical aspects that make it an interesting project
Some papers:

- Efficient Guided Generation for Large Language Models

  - By the outlines authors, probably the first real LLM constrained generation paper

  - https://arxiv.org/abs/2307.09702
- Automata-based constraints for language model decoding

  - A much more technical paper about constrained generation and implementation

  - https://arxiv.org/abs/2407.08103
- Pitfalls, Subtleties, and Techniques in Automata-Based Subword-Level Constrained Generation

  - A bit of self-promotion. We show where constrained generation can go wrong and discuss some techniques for the practitioner

  - https://openreview.net/pdf?id=DFybOGeGDS
Some blog posts:

- Fast, High-Fidelity LLM Decoding with Regex Constraints

  - Discusses adhering to the canonical tokenization (i.e., not just the constraint, but also what would be produced by the tokenizer)

  - https://vivien000.github.io/blog/journal/llm-decoding-with-regex-constraints.html
- Coalescence: making LLM inference 5x faster

  - Also from the outlines team

  - This is about skipping inference during constrained generation if you know there is only one valid token (common in the canonical tokenization setting)

  - https://blog.dottxt.ai/coalescence.html
mcyc··on Biconnected components
This is a nice attitude. I think HN is overall pretty nice for geeking out and also hearing other people geek out, but there is still a strain of elitism (not like StackExchange thankfully) and so I'm happy to see comments like this.
mcyc··on Tokenization for language modeling: BPE vs. Unigram Language Modeling (2020)
You can cross whitespace boundaries by setting flag `--split-on-whitespace` to false (it's true by default).

https://github.com/google/sentencepiece/blob/master/doc/opti...

mcyc··on Tokenization for language modeling: BPE vs. Unigram Language Modeling (2020)
Just a minor nit: SentencePiece is a library, not a tokenization algorithm. It implements two tokenization algorithms, Unigram and BPE.

BPE builds vocabularies from the base up so I assume you are talking about Unigram which starts with a big vocabulary and trims it.

The details of UnigramLM are here https://arxiv.org/pdf/1804.10959, and the part about vocabulary seeding is Section 3.2.

Basically, it just selects all substrings that appear in the corpus up to a certain length (and then maybe trims it a little by discarding rare substrings or something to reduce the initial size a bit and make things faster).

mcyc··on Show HN: Transductive regular expressions for text editing
People may also be interested in Pynini [1], a python wrapper (+ a lot of additional ease-of-use functionality) of OpenFst [2] (a really great library for transducers).

There are some good tutorials in the form of homework assignments (from like Johns Hopkins and some others) that go through Pynini use cases.

[1] https://www.openfst.org/twiki/bin/view/GRM/Pynini

[2] https://www.openfst.org/

mcyc··on Ask HN: Have you ever taken a career break or gap year to hack?
You might be interested in the Recurse Center (https://www.recurse.com/) and the experiences of people who have gone through it (they heavily encourage blogging about your time there so there is lots to read).

Note: I am not affiliated with the Recurse Center, just a big fan.

mcyc··on Tokenisation Is NP-Complete
Thanks!

Our paper [1] is kind of a goofy adversarial thing where we thought "here's this cool metric, how can we break it?". The tokenizers we propose are definitely not tokenizers you should use in practice.

The original paper that proposes the metric is, imo, much more interesting theoretically [2].

[1]: https://aclanthology.org/2024.lrec-main.1469/

[2]: https://aclanthology.org/2023.acl-long.284/

mcyc··on Tokenisation Is NP-Complete
It's two problems:

1) the sequence length increases too much. Idk what the average token length is for Llama, but imagine it's like 5+ bytes. Using individual bytes as tokens immediately makes the context 5x longer which is super bad for inference speed and memory requirements (since attention inference is quadratic in the length of the sequence).

2) individual bytes have essentially no meaning, so byte embeddings are harder to learn. Subword tokens aren't a perfect solution, but they definitely often have some standalone meaning where embeddings make sense.

I'll give another example from a recent paper that tries to eliminate tokenizers (this is a popular research direction) [1].

Figure 4 is a really good example of why byte-level models are wasting computation. Once part of a word is generated, most of the remaining bytes are assigned basically probability 1. But a byte-level model would still have to spend time decoding them. With a subword-level model most of these easy-to-decode bytes would be packed together in a single token so you don't have to decode them individually.

When model APIs bill by the token, this is an important consideration.

[1]: https://arxiv.org/abs/2412.09871

mcyc··on Tokenisation Is NP-Complete
NB: Can't edit my original reply.

Sorry actually I misread part of your comment in relation to the paper and confused δ and another parameter, K.

To clarify, δ is the number of tokens in the tokenized corpus and K is the size of the vocabulary.

So, if you are asking about why would they limit _K_, then my answer still applies (after swapping δ for K). But if you still mean "why do they pick some arbitrary δ as the limit of the size of the tokenized corpus", then I think the answer is just "because that makes it a decision problem".

mcyc··on Tokenisation Is NP-Complete
Hi, I'm Cognetta from the above Cognetta et al. I can't answer all of your questions (and I can't speak for the authors of this paper ofc), but I will try to answer some.

> Is a tokenizer that maximizes the compression of text (e.g. by identifying longer tokens that tend to be used whole) necessarily a better tokenizer, in terms of overall model performance? Compression might be a useful property for an objective function to consider... but then again maybe not, if it makes the problem NP-hard.

Compression isn't necessarily the best metric for language modeling quality [1][2][3], but there are some papers that find a correlation between it and quality [4] and also it has one important benefit: it reduces inference time by making the input sequences shorter (this is particularly important for transformers, because the runtime is quadratic in the sequence length).

If you imagine that with enough data, basically any reasonable tokenization algorithm would be ok (I think this is mostly true; there are definitely bad and "better" tokenizers and you see this very clearly in small data settings, but once you get into the trillions-of-tokens and 10s-of-billion-of-parameters setting, other things are going to matter more), then optimizing the tokenizer for compression is a good choice as it will provide tangible, practical benefits in the sense of reduced inference time.

> I'm also not sure how realistic the limitation to "at most δ symbols" is. [...] But why not just keep adding tokens as needed, rather than imposing any preordained limit?

This is a pretty realistic limitation imo. Of course you can arbitrarily increase the vocabulary size, but there is a tradeoff between modeling quality, parameter count, and inference time. If you increase the vocabulary a bunch, your inference speed will probably improve (although now you have a much larger softmax at the end of your model, which isn't usually a bottleneck anymore, but still not great), parameter count will increase (due to the larger embedding table), and your modeling quality will go down (in that you have tokens which are so rare in the corpus that they are massively undertrained; this can cause big problems [5]).

So by constraining it to δ, you are basically setting a parameter budget for the vocabulary, and this is a pretty reasonable thing to do.

> IIRC OpenAI's tokenizer has a vocabulary of around 52k subword strings.

Yeah, the size of the vocabulary varies a lot across models, but it isn't unusual to see significantly larger vocabularies these days (e.g., gemma has ~256k). However, these are still finite and very small compared to the corpus size.

> How could you possibly choose a meaningful δ from first principles?

This is a really great question, and something that we don't know how to answer. A lot of work has tried to answer it [6][7], but it is very much an open question.

[1]: https://arxiv.org/abs/2310.08754

[2]: https://aclanthology.org/2023.acl-long.284/

[3]: https://aclanthology.org/2024.emnlp-main.40/

[4]: https://arxiv.org/abs/2403.06265

[5]: https://aclanthology.org/2024.emnlp-main.649/

[6]: https://aclanthology.org/2023.acl-long.284/

[7]: https://aclanthology.org/2020.findings-emnlp.352/

mcyc··on Bit-permuting 16 u32s at once with AVX-512
It's just a convention for SIMD functions and types.
mcyc··on Call for Developer Projects
Suggested projects for the AT Protocol from the Bluesky devs.
mcyc··on Writing Month
A monthly writing challenge. See also: https://alpha.polymaths.social/@amin/statuses/01JBB8JKJMZ5ES...
mcyc··on Grandmaster-level chess without search
A bit late, but I wrote up an example inference script here: https://gist.github.com/mcognetta/7a98e50859664b8efbb4ec094a...

It is a bit roundabout, since it involves converting maia models to onnx before loading into pytorch and some outdated versions of libraries (maia/lc0 are a little old). We were using this for transfer learning for a competition, so we needed some flexibility that we didn't know how to do quickly/easily in TF.

Hope this helps.

------------------

Personal note: given your interest in chess ai and your starcraft username, I think we would have a lot of shared interests. Feel free to reach out (info is in my profile).

mcyc··on Double-Sided Gboard
See also: https://landing.google.co.jp/double-sided/ (in Japanese)
mcyc··on Finding a random seed that solves a LeetCode problem (2023)
You may enjoy this article: https://arxiv.org/abs/2109.08203

The author treats the seed as a hyperparameter and searches for the one that performs best for training a CV model.

mcyc··on Bridging empirical-theoretical gap in neural network formal language learning
If you are interested in this field (exploring the limits of neural models using formal language theory), I help run a weekly seminar on it, Formal Languages and Neural Networks:

https://flann.super.site/

We have had many great speakers (most of them are recorded and available on YouTube) and have a welcoming Discord.

mcyc··on Dean of Engineering at University of Nevada wrote a paper that’s bad
The Dean responded in the comments of this post.

https://statmodeling.stat.columbia.edu/2024/02/06/its-bezzle...

mcyc··on Show HN: Qwertle, yet another daily word game
Ah, the other color schemes look really good (especially I like the heatmap one). However, the default stoplight one is nearly impossible for me to disambiguate. For example, in [1], I can't tell which of mine are closer at all (I don't think I am colorblind, but maybe I'm just in for a surprise today).

Anyway, overall, a very fun variant. Thanks for sharing!

[1]: https://shorturl.at/cBJQ1

mcyc··on Show HN: Qwertle, yet another daily word game
Interpolating with the game's specified final green value (34, 139, 34) "forest green" provides a nicer effect than true green (0, 255, 0).

https://twitter.com/good_in_theory/status/175088137157541934...

mcyc··on Show HN: Qwertle, yet another daily word game
I also found the color scheme to be difficult to understand. A friend suggested interpolating between red and green depending on how close you are to the correct key. This is not too hard, since red is (255, 0, 0) and green is (0, 255, 0), so you can compute a distance (normalized to [0, 1]) and output (255 x d, 255 x (1-d), 0) to get the interpolated color.

It looks quite nice visually.

I wrote a small thread on it: https://twitter.com/good_in_theory/status/175079370720771734...

https://sigmoid.social/@mc/111821291510126156

Page 1 of 2Next →