HNHacker News
TopNewBestAskShowJobs

lunixbochs

1,848 karma · joined April 25, 2012

@lunixbochs // hn@bochs.info // talonvoice.com - don't hurt your hands!
submissionscomments
lunixbochs··on Solo – a .so loader for static Linux binaries
similar energy to my https://github.com/lunixbochs/crossldso
lunixbochs··on Speech Recognition and TTS in less than 500kb
I think Google's Conformer paper is SOTA at the <30M model size, where I think they put an incredible amount of flops into a 10M param model to reach around 2% lsc clean (the whole model and RNN decoder were trained domain specific to librispeech here).

I think my small Talon models are next, around 3% lsc clean at ~28M (greedy CTC decoding, no external encoder, no LM, not trained in a domain specific way). I reached around 6.5% at 10M.

I've been working on some new baselines I want to release soon as public artifacts. This article is inspiring me to try pushing the param size down a bit. I suspect we can do large vocabulary end to end in the <5M range.

lunixbochs··on Preparing for KDE Plasma's Last X11-Supported Release
Thank you.
lunixbochs··on Preparing for KDE Plasma's Last X11-Supported Release
I don't think any of your points reflect what I was trying to communicate.

> You're not interested in addressing customers' needs

I would love to support Wayland, but it is my position that it is impossible to "support Wayland" for Talon. I can only support a subset of the features and only on specific compositors, and it would be a lot of work.

> or giving them ways to address their needs themselves

As I said at the top of the message you are replying to, I believe today users already have the tools to address their needs themselves with about the same level of jank I'd be able to provide on Wayland. If this is a veiled hard line on open source being the only way for users to address their needs themselves, we have a philosophical difference that won't be sorted out in this thread.

> If slack interactions are so unpleasant, why do you direct all support through it?

That's a whole new sentence. I was specifically referring to the support requests for Wayland, which in the long tail have been more hostile toward me than is likely warranted.

> "yes this is a hack, but we'll live with it for now until an actually good solution is available"

The hack is switching to X11, which is fully supported, or working around it in your user scripts, which has already been done by some users for their specific environment.

lunixbochs··on Preparing for KDE Plasma's Last X11-Supported Release
Hi, I'm the developer of Talon.

It's possible to do the simple compositor specific hacks from Talon's scripting system to give yourself partial Wayland support at roughly the quality I'd be able to provide myself, and I know of a couple efforts to do this.

The tentative plan for "dropping support for X11" is just to do one more public Linux X11 release, stop there, leave it available to download, and make it very clear what to expect when you download for Linux or run on Wayland. I plan to continue supporting X11 on the paid version indefinitely.

Most requests relating to Wayland on the Slack have not been offers to help, and way too many have ended up being unpleasant conversations.

I have a standing offer to reconsider my stance if someone can show the vast majority of the necessary APIs are available and well supported without compositor specific hacks.

For some of the other points above, consider me disappointed but not surprised.

lunixbochs··on I/O is no longer the bottleneck? (2022)
naive c is just a memcpy. non-temporal uses the streaming instructions.
lunixbochs··on I/O is no longer the bottleneck? (2022)
your single core numbers seem way too low for peak throughput on one core, unless you stipulate that all cores are active and contending with each other for bandwidth

e.g. dual channel zen 1 showing 25GB/s on a single core https://stackoverflow.com/a/44948720

I wrote some microbenchmarks for single-threaded memcpy

    zen 2 (8-channel DDR4)
    naive c:
      17GB/s
    non-temporal avx:
      35GB/s

    Xeon-D 1541 (2-channel DDR4, my weakest system, ten years old)
    naive c:
      9GB/s
    non-temporal avx:
      13.5GB/s

    apple silicon tests
    (warm = generate new source buffer, memset(0) output buffer, add memory fence, then run the same copy again)

    m3
    naive c:
      17GB/s cold, 41GB/s warm
    non-temporal neon:
      78GB/s cold+warm

    m3 max 
    naive c:
      25GB/s cold, 65GB/s warm
    non-temporal neon:
      49GB/s cold, 125GB/s warm

    m4 pro
    naive c:
      13.8GB/s cold, 65GB/s warm
    non-temporal neon:
      49GB/s cold, 125GB/s warm

    (I'm not actually sure offhand why asi warm is so much faster than cold - the source buffer is filled with new random data each iteration, I'm using memory fences, and I still see the speedup with 16GB src/dst buffers much larger than cache. x86/linux didn't have any kind of cold/warm test difference. my guess would be that it's something about kernel page accounting and not related to the cpu)
I really don't see how you can claim either a 6GB/s single core limit on x86 or a 20GB/s limit on apple silicon
lunixbochs··on Python numbers every programmer should know
I'm confused why they repeatedly call a slots class larger than a regular dict class, but don't count the size of the dict
lunixbochs··on What makes you senior
I'm not familiar with C# compile at runtime. Are you saying your change was to do an AOT compile locally?
lunixbochs··on Compressing Text into Images
I did a silly experiment to compress word embeddings with jpeg - to see how it collapses semantically as you decrease the quality.

https://bochs.info/vec2jpg/

This was a very basic experiment. I expect you could perform the DCT more intelligently on the vector dimensions instead of trying to pack the embeddings into pixels, and get higher quality semantic compression.

lunixbochs··on StableLM: A new open-source language model
Are you using https://github.com/EleutherAI/lm-evaluation-harness?
lunixbochs··on Box64 – Linux Userspace x86_64 Emulator Targeted at ARM64 Linux Devices
> made gl4es

Neverball was working in the original glshim project before ptitseb forked it to gl4es. (Not to discount the significant work he's put in since, including the ES2 backend)

lunixbochs··on Numen: Voice Control for Handsfree Computing
The Talon model is fairly accurate, but it can be confusing for new users to use the command system correctly. I posted a sibling reply about this, but the most common reason for Talon users to complain about the recognition is that they are in the strict "command mode" and say things that aren't actually commands.

If you encounter what feels like poor recognition in Talon, I recommend enabling Save Recordings and zipping+sharing some examples on the Slack and asking for advice.

The current command set is definitely harder to learn than a system designed for chat/email where "what you say is what you get", but it's much more powerful for tasks like programming once you learn it.

I'm dubious about what kind of general command accuracy Numen is able to get with the Vosk models, as Vosk to my understanding is more designed for natural language than commands.

lunixbochs··on Numen: Voice Control for Handsfree Computing
Fixed commands are fast, precise, and predictable.

Assuming you mean speaking in natural language, that's slower to say, and likely less precise and predictable if you want to be able to just say "anything" any have a result.

You need a command system either way. If you want to express some precise intention, you need to understand what the command system will do.

There is a combined "mixed mode" system I've been testing in the talon beta where you can use both phrases and commands without switching modes.

lunixbochs··on Numen: Voice Control for Handsfree Computing
Talon's eye tracking functions as a mouse replacement. Is there a specific demo you'd like to see? I can record one.
lunixbochs··on Numen: Voice Control for Handsfree Computing
Depending on when that was: in 2018 the free model was the macOS speech engine, in 2019 it was a fast but relatively weak model, and as of late 2021 it's a much stronger model. I'm currently working on the next model series with a lot more resources than I had before.

It's also worth saying that if you only tried things out briefly, there are a handful of reasons recognition may have seemed worse. Talon uses a strict command system by default, because that improves precision and speed for trained users, but the tradeoff there is it's more confusing for people who haven't learned it yet.

For example, Talon isn't in "dictation mode" by default, so you need to switch to that if you're trying to write email-like text and don't want to prefix your phrases with a command like "say".

The timeout system may also be confusing at first. When you pause, Talon assumes you were done speaking and tries to run whatever you said. You can mitigate this by speaking faster or increasing the timeout.

The default commands (like the alphabet) may also just not be very good for some accents, and that will be the case for any speech engine - you will likely need to change some commands if they're hard to enunciate in your accent.

I recommend joining the slack [1] and asking there if you want more specific feedback. I definitely want to support many accents and even have some users testing Talon with other spoken languages.

[1] https://talonvoice.com/chat

lunixbochs··on Numen: Voice Control for Handsfree Computing
A handful of the datasets I tested are fully held out (I have reason to believe none of the models have trained on them), and talon was trained on none of the dev or test data of any of the datasets in question.

Due to whisper's weakly supervised training on a large amount of automatically scraped data and reliance on a bigger language model, it's far more likely whisper had seen some of the test data before.

lunixbochs··on Ask HN: Vision Models to Parse UI
Android voice access blogged about how they use a model to detect and classify buttons:

https://ai.googleblog.com/2021/01/improving-mobile-app-acces...

lunixbochs··on OpenAI quietly launched Whisper V2 in a GitHub commit
Can you elaborate on how you see this working?
lunixbochs··on OpenAI quietly launched Whisper V2 in a GitHub commit
In large scale tests, I observed hallucinations from Whisper in speech regions of audio.
lunixbochs··on OpenAI quietly launched Whisper V2 in a GitHub commit
Nice catch. I'll run my test suite [1] on this and report back.

[1] https://twitter.com/lunixbochs/status/1574848899897884672

lunixbochs··on OpenAI quietly launched Whisper V2 in a GitHub commit
If the target is the words "A A A" and you produce "B B B B", you have more errors than there were words in the target. 3 replacements and 1 insertion.
lunixbochs··on TinyGL 0.4.1
I used to work on a project called glshim (which ptitseb still maintains as gl4es, used by box86). glshim implements an ABI-compatible OpenGL 1.x fixed function API (as libGL.so.1) on top of an OpenGL ES 1.x driver. This allows you to accelerate completely unmodified OpenGL programs on mobile Linux devices such as the OpenPandora.

This approach even works with x86 user-mode emulation, by serializing render commands into a shared memory ring buffer and running the driver in native code out-of-process (which as a bonus was also generally faster than rendering in-process).

I built glshim after hand porting a bunch of open-source Linux games from OpenGL to OpenGL ES. The last game I hand-ported like this was Introversion's Uplink Hacker Elite. The diff was thousands of lines, and the resulting port had buggy rendering and blurry fonts, so I got fed up with hand ports and started working on a universal solution.

After a few years, I had ironed out most of the bugs in glshim, and ports to the Pandora were much easier. Uplink, however, was still in a sorry state. It made heavy use of glBitmap for crisp font rendering, and glBitmap was very slow to emulate on top of OpenGL ES 1.x. glBitmap basically manipulates framebuffer pixels directly, and the easiest way I found to do that was to upload a texture and render it with GL_NEAREST. Texture uploads seemed to be extremely slow and CPU-bound on the Pandora's Cortex A8 / PowerVR SGX 530. I ended up batching subsequent glBitmap calls, trying to coalesce them into a single texture upload per frame, and with that I still only managed single-digit fps.

With glshim, I had a moderately complete OpenGL implementation that only relied on a small subset of the OpenGL ES API. Then I stumbled on TinyGL. I realized that despite TinyGL not implementing much of the full OpenGL API, it was very fast and implemented basically everything I needed from OpenGL ES for glshim to work. So I forked it as "TinyGLES", deleted the extra GL parts, and optimized the rendering paths for ARM NEON. It ended up making a really good software rendering backend for glshim. I used this to finally release a good version of Uplink for the Pandora, which ran at around 80fps! With TinyGLES, glBitmap could blit directly to the software framebuffer, avoiding the slow/buggy texture emulation path.

It had been a few years and I couldn't find the modified source for my Uplink port, so I just swapped out the libGL.so file when I updated Uplink to use TinyGLES. But some users found an annoying bug in the port. Anytime you restarted the game, it put you directly in the "new game" menu, instead of allowing you to pick between new game or loading a save. I ended up fixing this with a binary patch to the game executable (which forced the game to the main menu screen on every launch), and that's the version currently available for the Pandora: http://repo.openpandora.org/?page=detail&app=uplink

lunixbochs··on Why is Rosetta 2 fast?
> To see ahead-of-time translated Rosetta code, I believe I had to disable SIP, compile a new x86 binary, give it a unique name, run it, and then run otool -tv /var/db/oah///unique-name.aot (or use your tool of choice – it’s just a Mach-O binary). This was done on old version of macOS, so things may have changed and improved since then.

My aotool project uses a trick to extract the AOT binary without root or disabling SIP: https://github.com/lunixbochs/meta/tree/master/utils/aotool

lunixbochs··on Whisper – open source speech recognition by OpenAI
Just posted results here: https://twitter.com/lunixbochs/status/1574848899897884672
lunixbochs··on Whisper – open source speech recognition by OpenAI
Do you have a demo audio clip for this? I'd be interested to see how it looks in practice.
lunixbochs··on Whisper – open source speech recognition by OpenAI
If the Whisper models provide any benefits over the existing Talon models, and if it's possible to achieve any kind of reasonable interactive performance, I will likely integrate Whisper models into Talon.

Talon's speech engine backend is modular, with Dragon, Vosk, the WebSpeech API, and Talon's own engine all used in different ways by users.

lunixbochs··on Whisper – open source speech recognition by OpenAI
For an offline (non-streaming) model, 1x realtime is actually kind of bad, because you need to wait for the audio to be available before you can start processing it. So if you wait 10 seconds for someone to finish speaking, you won't have the result until 10 seconds after that.

You could use really small chunk sizes and process them in a streaming fashion, but that would impact accuracy, as you're significantly limiting available context.

lunixbochs··on Whisper – open source speech recognition by OpenAI
If you count the GPU component and memory bandwidth, the Apple M2 is slightly weaker on paper for 16-bit inference than the NVIDIA A2, if you manage to use the whole chip efficiently. The A16 is then slightly weaker than the M2.

Sure, the Whisper Tiny model is probably going to be fast enough, but from my preliminary results I'm not sure it will be any better than other models that are much much faster at this power class.

Whisper Large looks pretty cool, but it seems much harder to run in any meaningful realtime fashion. It's likely pretty useful for batch transcription though.

Even if you hit a realtime factor of 1x, the model can leverage up to 30 seconds of future audio context. So at 1x, if you speak for 10 seconds, you'll potentially need to wait another 10 seconds to use the result. This kind of latency is generally unsatisfying.

lunixbochs··on Whisper – open source speech recognition by OpenAI
Ok, my test harness is ready. My A40 box will be busy until later tonight, but on an NVIDIA A2 [1], this is the batchsize=1 throughput I'm seeing. Common Voice, default Whisper settings, card is staying at 97-100% utilization:

  tiny.en: ~18 sec/sec
  base.en: ~14 sec/sec
  small.en: ~6 sec sec/sec
  medium.en: ~2.2 sec/sec
  large: ~1.0 sec/sec (fairly wide variance when ramping up as this is slow to process individual clips)
[1] https://www.nvidia.com/en-us/data-center/products/a2/
Page 1 of 19Next →