HNHacker News
TopNewBestAskShowJobs

willwade

152 karma · joined September 28, 2023

submissionscomments
willwade··on Show HN: S3mini – Tiny and fast S3-compatible client, no-deps, edge-ready
I have a good suspicion this has been written with help from an LLM. The heavy use of emojis and strong hyper confident language is the giveaway. Proof: my own repos look like this after they’ve had the touch of cursor / windsurf etc. still doesn’t take away if the code is useful or good.
willwade··on Comparing Docusaurus and Starlight and why we made the switch
Particularly like the honest take. I wouldn’t say reading this I’d go for starlight either

> When I tried to create marketing pages with Starlight in addition to the technical documentation, I nearly gave up. Coming from the Docusaurus world, this wasn't an issue as the starter template comes with a front page and a blog out of the box. You can even create multiple documentations on different paths. We used /docs, for example.

> Starlight, on the other hand, is only built for documentation and not for marketing pages. It even took an ugly hack to make sure the default path is /docs and not /.

> Please don't look at this custom script configured in our astro.config.mjs we need to execute on every page to make sure that the redirection works properly

Love the straight talking. Refreshing in the period of ai slop blogposts

willwade··on JavaScript-TTS-Wrapper – A Node and CommonJS TTS Wrapper
Its a bit rough with SherpaOnnx and eSpeak - So PR's welcome.

Idea here is this is a unifed TTS wrapper so getVoices is the same for all engines, and bytestream is the same and word events are the same - even if they don't support it we estimate it.

Maybe useful since I see a lot of TTS hackers out there!

NB: I will say JS isnt my bag - so feedback all welcome.

willwade··on Show HN: Aqua Voice 2 – Fast Voice Input for Mac and Windows
You’re real market you need to go hard on is the assistive tech market. You know the biggest companies in this space are those solving problems for dyslexia where govt grants in eg UK fund pretty much all their work? I had an access to work assessment and they recommend like sweets stuff from texthelp. It’s then paid for by the government following these assessments. But it’s crap. It literally is a crap tool for adhd or dyslexia because these users literally CANT remember or deal with barriers like learning how to dictate correctly. Aqua voice solves this. I’m your biggest fan. I recommend it in my AT assessments all the time :)
willwade··on Show HN: Aqua Voice 2 – Fast Voice Input for Mac and Windows
Aqua voice is nothing like talon. I wouldn’t bother trying to compare. It’s a dictation tool. Just entry. Not commands. But it’s bloody impressive. You don’t need to learn anything - you just talk like you would talk to someone across the way from you
willwade··on Coqui TTS: Free Text-to-Speech
Piper isn’t dead. I know because I’ve been working with the team recently
willwade··on Coqui TTS: Free Text-to-Speech
Is there anything new here? I thought this pretty much is dead. Most models available in Sherpa onnx and piper right?
willwade··on The Alexa feature "Do Not Send Voice Recordings" you enabled no longer available
Now is the time to look at Open Home Foundation and their Voice Assistant device which does Speech Processing offline.. Donate! Buy their stuff!
willwade··on Show HN: I Built a Telegraph Simulator
Nice. If you want to learn Morse try https://morse-learn.acecentre.net/ (I must throw cursor at this to fix some issues like working nicely on mobile).
willwade··on Show HN: While the world builds AI Agents, I'm just building calculators
Ok. I need a fully accessible sci calculator that is keyboard accessible for all functions. Check out desmos site. Their calc isn’t bad but clear button isn’t keyboard accessible.
willwade··on GibberLink [AI-AI Communication]
This reminds of me chirp https://archive.is/HEC29

https://audioxpress.com/news/data-over-sound-pioneer-chirp-a...

willwade··on I ditched my Pi-hole but still block ads with NextDNS
Does anyone out there have a working Apple Shortcut that can toggle on/off a denied domain like YouTube? That's one feature I had in Pihole that I can't seem to replicate.

Update.. as usual I was trying this all day and only after posting this does this work - Here's a bash script https://gist.github.com/willwade/251fa791da27267b5470c75a7b5... - a shortcut for this is way more complicated

willwade··on I shouldn't of made this. Another direnv alternative for no reason
Maybe of interest to someone out there. I'm waiting to get flamed as its totally useless for many many people (I could of solved it with a simple powershell script but in the process I learnt go.. ) - but it solved my problem on my ridiculous Windows machine. Maybe someone else out there..
willwade··on One Text to Speech wrapper for Python for all (?) TTS engines
sherpa-onnx. use one of these models - https://k2-fsa.github.io/sherpa/onnx/tts/pretrained_models/r... - if you use py3-tts-wrapper you can get a list of these using the get_voices method
willwade··on One Text to Speech wrapper for Python for all (?) TTS engines
fyi we currently support macos with systemtts (so sapi on windows and nss on Mac). But I really need to build a new replacement as nsss is deprecated.. nb: our engine does more than say - we do word events and standardise languages/calls etc
willwade··on One Text to Speech wrapper for Python for all (?) TTS engines
Coqui works using Sherpa onnx.
willwade··on One Text to Speech wrapper for Python for all (?) TTS engines
Well all, its a bit endless.. this is a hard task since few of the new players (11labs, PlayHT) bother with SSML or word timings.. but you know, we try and deal with workarounds. Feel free to get on board if TTS is your thing..
willwade··on Show HN: I built a full mulimodal LLM by merging multiple models into one
You know - the word "multimodal" i think is being used badly here. Its Multi-Model - not Multimodal - which certainly suggests a completeley different thing
willwade··on Edge TTS
https://ttsvoicesavailable.streamlit.app

Acapela, Nuance - but its around 75 languages.

willwade··on Wi-Fi and the Problem with Radar (DFS)
Interesting this. It reminds me of when my dad used to work at NATS /CAA (uk air traffic body). I remember him getting complaints in the early 90s that people’s car locking systems failed to work on the car park next to the radar tower at Gatwick. Turned out the radar (or the huge electrical field created by the motors in the radar) created havoc. If I remember right it ended up being a question on the back page of the new scientist..
willwade··on Markov Keyboard: keyboard layout that changes by Markov frequency (2019)
Well I knocked something together thanks really to @Lemaxoxo for his blog post inspiring me to do this.

https://scanningmvp.netlify.app/ - it’s more for my needs of an AT keyboard for scanning than your backlit one but you get the idea.

willwade··on Markov Keyboard: keyboard layout that changes by Markov frequency (2019)
Heard of dasher ? https://www.inference.org.uk/dasher/
willwade··on Markov Keyboard: keyboard layout that changes by Markov frequency
thats wonderful. id like that to not change the order of the letters - but change the highlight order. Do a round 1 of frequency order first (just do first say 6 letters) then do a round 2 which is standard order..

i probably am not making much sense. Look at where I'm coming from in the world of Assistive Tech - https://docs.acecentre.org.uk/products/echo (go to around 5 min mark in the vide)

willwade··on Markov Keyboard: keyboard layout that changes by Markov frequency (2019)
Nice. Some thoughts.

Look at PPM. Your prediction model would work better with personalised data. PPM is efficient (nb. I see you are using python - look at this https://github.com/willwade/pylm - although be warned - i think my code is not quite right..)

Layout shifting for finger movement - well its great if you didnt have to look. The time for visual processing the letters adds a significant lag (its why typical word prediction isnt used that much and when it is - not over 3 predictions (I have papers on this if you are interested). But its not all bad..

Switch users who need next letter prediction this could dramatically support their rate of input. (view https://youtu.be/Bhj5vs9P5cw?si=VnytfH_vdEUWuLok&t=73 - now note how the keyboard blocks the scan up. But imagine if it just scanned each letter first by next most likely - or heck - like this repo - actually changes button position and kept the scan pattern the same. It would be a ton more efficient)

(and a bit of a rabbit hole.. What if keys had word predictions on them? This is basically the end result of ACE-LP: https://discovery.dundee.ac.uk/en/publications/ace-lp-augmen...)

willwade··on A BBC navigation bar component broke depending on the external monitor
Back in html 4 days we did this shenanigans all the time. I worked on very over the top sites that played with multiple windows talking to each other and moving in synchrony. I’ve tried looking for examples on archive.org (eg I know we did this a ton on flash heavy sites like design museum in London ) but alas the ones I was looking for a broken in that archive.
willwade··on A BBC navigation bar component broke depending on the external monitor
See I work in accessibility. Like I provide and create solutions direct to end users with complex needs. Not regular web accessibility. I get the view of this. It’s the same idea of universal access. But actually I don’t fully agree. Yes. If you can stick to this principle - and do try / but I promise you edge cases - which in itself is what accessibility users are all about - cause headaches. At some level you have to do custom stuff. It’s the best way. Take for example switch users. Yes. If your ui is tab able - great. But what if you need your items scannable in frequency order. Your tab index needs to change to meet the end users needs. Or eye gaze users. The accuracy level changes. Add in cognitive issues. You can’t just make a one size fits all interface. At some stage you need to significantly customize it. You can’t rely on a user just learning a complex system level interaction technique- if they can’t do that you have to customise on an individual level.
willwade··on Phonetic Matching
Im intrigued.. Is this not done just with a phonemizer?

    from phonemizer.phonemize import phonemize

    text = "hello world"
    variations = [
        phonemize(text, backend="espeak", language="en-us", strip=True),
        phonemize(text, backend="espeak", language="en-gb", strip=True),
        phonemize(text, backend="espeak", language="en-au", strip=True),
    ]

I mean, espeak isnt the best but a lot of folks in the ASR/Speech world still are using this right?

(NB: If you are on iOS check out the inbuilt one - Settings -> Accessibility -> Spoken Content -> Pronounciations. Adding one it has the ability to phonemize to IPA your spoken message. If someone can tell me where that SDK/API is they use in that I'd love to know) for i, variation in enumerate(variations, 1): print(f"Variation {i}: {variation}")

willwade··on Control iPhone with the movement of your eyes
Probably not. Typical eyegaze using IR has two main ways of working - bright eye or dark eye. This though wont care as much. Its based on a inference model..
willwade··on Show HN: Bluetooth USB Peripheral Relay – Bridge Bluetooth Devices to USB
Nice. Check out this guys repos for stuff using nrf chips. It’s generally the other way round. Really nice. https://github.com/gdsports/ble-usb-devices

https://github.com/gdsports/usbhostcopro

willwade··on A CC-By Open-Source TTS Model with Voice Cloning
It was used quite a bit of speech to text - but tts it’s not that great.
← PreviousPage 2 of 4Next →