It’s pretty good at understanding my broken Mandarin, though.
707 karma · joined June 20, 2012
blog: simedw.com X: https://x.com/simedw
It’s pretty good at understanding my broken Mandarin, though.
The only alternative version of Go I’ve played before was magnetic Go, where all the rules are the same except that, along the axis lines where you place a piece, like colours attract and opposite colours repel.
On Voronoi, I did feel like the continuous space made it really hard to judge whether there was enough room, whether the opponent could get in, etc. Maybe it would work better if it were discretized a bit? Or maybe I just need to play it more.
For the current model I’m using Core ML, which optimizes the kernels the first time you run it. I haven’t actually spent that much time tuning performance beyond that.
I can probably squeeze out quite a bit more than 100 notes/sec as well. I haven’t spent much time optimizing inference yet.
Pretraining was obviously a a lot slower, the 125M model took roughly half a day.
NOTE(
pitch,
delta_onset,
duration,
attack_velocity,
release_velocity,
...
)
More global things might be better modelled as a control event: NOTE(...)
| CONTROL_SUSTAIN(delta_onset, value)
| CONTROL_TEMPO(delta_onset, bpm)
| CONTROL_PROGRAM(delta_onset, instrument)
The harder part might actually be finding enough good training data with all of those attributes represented consistently.Yes, I think I’ve gotten it to roughly a GPT-2 level: good enough to share, but with a lot of room left to improve. I think adding some kind of bar/measure token might help with rhythm, and perhaps some form of longer-term planning for the overall composition.
That said, for `createAliasMap`, don't you think you could create a deterministic mapping from and to UUIDs <-> word chains? That way, no additional state would be needed. [Might require fairly long word chains...]
I noticed that if you go from training to watch and then back, the training temporarily drop significantly in score.
The generator first needs a depth map, and then derives the repeating pattern from that. A normal RGB image would be far too noisy; the fine texture variations would break the repetition needed for the brain to fuse the patterns correctly.
simedw ~ $ claude -p "random number between 1 and 10"
7
simedw ~ $ claude -p "random number between 1 and 10"
7
simedw ~ $ claude -p "random number between 1 and 10"
7
simedw ~ $ claude -p "random number between 1 and 10"
7I have just added sandhi support, please let me know if it's working better.
The other two are probably things that could be fixed with a bigger and more varied dataset.
I had a quick look at Farsi datasets, and there seem to be a few options. That said, written Farsi doesn’t include short vowels… so can you derive pronunciation from the text using rules?
Since you’ve been so consistent and are using your own software, have you experimented with different resurfacing rates? Did you notice a material difference in recall?
1. How does it know which words I already know? It doesn’t automatically. You provide that set. For example, if you’ve completed HSK 1, you can paste the HSK 1 word list into LangSeed and mark those as "known". From there, new explanations are constrained to that vocabulary. You can also paste in real text and mark the easy words as known, though that’s a bit more manual.
2. How much might I misunderstand word meanings? Depends on how advanced the vocab is and how large your known-word set is. I think of this as building intuition rather than giving dictionary-precise definitions. As you see words in more contexts, that intuition sharpens. This is just my experience from testing it over the last couple of weeks.
3. How inaccurate are the explanations? I tested it on Swedish (my native language). There are occasional awkward or slightly odd phrasings, but it’s rarely outright wrong.
I recently had some trouble converting a HF transformer I trained with PyTorch to Core ML. I just couldn’t get the KV cache to work, which made it unusably slow after 50 tokens…