HNHacker News
TopNewBestAskShowJobs

simedw

707 karma · joined June 20, 2012

CTO/co-founder of V7, before that Aipoly.

blog: simedw.com X: https://x.com/simedw

submissionscomments
simedw··on Gemini-3.5-Transcribe
For some reason, in almost every test I did with Gemini 3.5 Live, it classified my Swedish (native) as Dutch. There are similarities between the languages, but still.

It’s pretty good at understanding my broken Mandarin, though.

simedw··on Show HN: Voronoi Go
This was neat, well done!

The only alternative version of Go I’ve played before was magnetic Go, where all the rules are the same except that, along the axis lines where you place a piece, like colours attract and opposite colours repel.

On Voronoi, I did feel like the continuous space made it really hard to judge whether there was enough room, whether the opponent could get in, etc. Maybe it would work better if it were discretized a bit? Or maybe I just need to play it more.

simedw··on Show HN: I trained a 125M model to autocomplete piano on-device
Luckily I have access to 4x RTX 4090s, so I didn’t have to pay cloud GPU prices directly. If I had, it probably would have added up quite a bit given how many training runs and experiments I ended up doing.
simedw··on Show HN: I trained a 125M model to autocomplete piano on-device
The biggest speed improvement came from changing the note representation when I switched to compound note events: roughly 5× fewer autoregressive passes per note.

For the current model I’m using Core ML, which optimizes the kernels the first time you run it. I haven’t actually spent that much time tuning performance beyond that.

simedw··on Show HN: I trained a 125M model to autocomplete piano on-device
Yes, some kind of planning step is on my TODO list. Another thing I want to try is generating a few continuations in parallel, picking the one that looks best, and then continuing from there. Maybe the picking could be automatic.

I can probably squeeze out quite a bit more than 100 notes/sec as well. I haven’t spent much time optimizing inference yet.

simedw··on Show HN: I trained a 125M model to autocomplete piano on-device
For DPO I only had around 700 preference examples, so not much data at all. That took about 12 minutes to train on a single GPU.

Pretraining was obviously a a lot slower, the 125M model took roughly half a day.

simedw··on Show HN: I trained a 125M model to autocomplete piano on-device
If the attribute describes the current note, I would first try adding it as another field/head on the note event. For example:

  NOTE(
    pitch,
    delta_onset,
    duration,
    attack_velocity,
    release_velocity,
    ...
  )
More global things might be better modelled as a control event:

  NOTE(...)
  | CONTROL_SUSTAIN(delta_onset, value)
  | CONTROL_TEMPO(delta_onset, bpm)
  | CONTROL_PROGRAM(delta_onset, instrument)
The harder part might actually be finding enough good training data with all of those attributes represented consistently.
simedw··on Show HN: I trained a 125M model to autocomplete piano on-device
Thank you.

Yes, I think I’ve gotten it to roughly a GPT-2 level: good enough to share, but with a lot of room left to improve. I think adding some kind of bar/measure token might help with rhythm, and perhaps some form of longer-term planning for the overall composition.

simedw··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
Nice package, not only is using words more token-efficient [saving time and money], but weaker models are also less likely to make mistakes when providing the key, at least in my tests.

That said, for `createAliasMap`, don't you think you could create a deterministic mapping from and to UUIDs <-> word chains? That way, no additional state would be needed. [Might require fairly long word chains...]

simedw··on Show HN: Watch a neural net learn to play Snake
Cool project!

I noticed that if you go from training to watch and then back, the training temporarily drop significantly in score.

simedw··on Building a Magic Eye Generator and Decoder
No offense, but are you a bot?
simedw··on Building a Magic Eye Generator and Decoder
Agreed, it almost feels like we have a visual processing unit with special “opcodes” for operations like depth matching and pattern repetition.

The generator first needs a depth map, and then derives the repeating pattern from that. A normal RGB image would be far too noisy; the fine texture variations would break the repetition needed for the brain to fuse the patterns correctly.

simedw··on AI-generated password isn't random, it just looks that way
I think this speaks for itself:

  simedw ~  $ claude -p "random number between 1 and 10" 
  7
  simedw ~  $ claude -p "random number between 1 and 10"
  7
  simedw ~  $ claude -p "random number between 1 and 10"
  7
  simedw ~  $ claude -p "random number between 1 and 10"
  7
simedw··on Show HN: I trained a 9M speech model to fix my Mandarin tones
Great suggestin, added a toggle to see pinyin.
simedw··on Show HN: I trained a 9M speech model to fix my Mandarin tones
Thank for the great feedback!

I have just added sandhi support, please let me know if it's working better.

simedw··on Show HN: I trained a 9M speech model to fix my Mandarin tones
Hi, thanks for the feedback. The 了 issue was a bug on the JavaScript side; that should be fixed (training did thankfully handle it correctly).

The other two are probably things that could be fixed with a bigger and more varied dataset.

simedw··on Show HN: I trained a 9M speech model to fix my Mandarin tones
It’s fairly sensitive to background noise at the moment. I’m planning to train an improved version with stronger data augmentation, including background noise.
simedw··on Show HN: I trained a 9M speech model to fix my Mandarin tones
For accents, I’ve mostly tested with a few friends so far. I’m wondering whether region should be a parameter, because training on all dialects might make the system too lax.
simedw··on Show HN: I trained a 9M speech model to fix my Mandarin tones
Thank you.

I had a quick look at Farsi datasets, and there seem to be a few options. That said, written Farsi doesn’t include short vowels… so can you derive pronunciation from the text using rules?

simedw··on Coding Agents Are Good First-Time User Testers
It would be neat if it had a headless mode.
simedw··on Ask HN: Share your personal website
https://simedw.com personal site, mostly posts regarding various experiments
simedw··on Yearly analytics on my spaced repetition results
First of all, big kudos for not missing a single day. When I used flashcards in the past, missing even a couple of days led to an avalanche of cards to review.

Since you’ve been so consistent and are using your own software, have you experimented with different resurfacing rates? Did you notice a material difference in recall?

simedw··on Show HN: Learning a Language Using Only Words You Know
Thanks for the questions. Very fair concerns. Take all of this with a fairly large pinch of salt; this is still an experiment.

1. How does it know which words I already know? It doesn’t automatically. You provide that set. For example, if you’ve completed HSK 1, you can paste the HSK 1 word list into LangSeed and mark those as "known". From there, new explanations are constrained to that vocabulary. You can also paste in real text and mark the easy words as known, though that’s a bit more manual.

2. How much might I misunderstand word meanings? Depends on how advanced the vocab is and how large your known-word set is. I think of this as building intuition rather than giving dictionary-precise definitions. As you see words in more contexts, that intuition sharpens. This is just my experience from testing it over the last couple of weeks.

3. How inaccurate are the explanations? I tested it on Swedish (my native language). There are occasional awkward or slightly odd phrasings, but it’s rarely outright wrong.

simedw··on Show HN: Learning a Language Using Only Words You Know
This is a simplified version: Journey to the West in Easy Chinese by Jeff Pepper and Xiao Hui Wang. Otherwise, I would definitely have waited a bit before biting off something like this.
simedw··on Show HN: Learning a Language Using Only Words You Know
Surprisingly easy. If the language has a lot of conjugations (e.g., polite past verb forms), running each word through Snowball first makes the process a bit easier.
simedw··on Show HN: Learning a Language Using Only Words You Know
That's a really cool concept. Naively replacing words might work, but sometimes the context is needed. Maybe a model like gemini 2.5 flash lite would be fast enough but still maintain better context awareness?
simedw··on Prompt caching for cheaper LLM tokens
Thanks for sharing; you clearly spent a lot of time making this easy to digest. I especially like the tokens-to-embedding visualisation.

I recently had some trouble converting a HF transformer I trained with PyTorch to Core ML. I just couldn’t get the KV cache to work, which made it unusably slow after 50 tokens…

simedw··on Show HN: Learning a Language Using Only Words You Know
Thanks! I think getting comfortable with characters fairly early is important, as it helps shift your mindset into the right place. That said, I don’t think this project really works until you’re comfortable with at least ~60 characters.
simedw··on Gemini 2.5 Flash Image
I noticed that I get far fewer refusals when I set my VPN to the USA.
simedw··on Gemini 2.5 Flash Image
The model is only available in AI Studio when I set my VPN to the USA (I’m located in the UK).
Page 1 of 2Next →