HNHacker News
TopNewBestAskShowJobs

sosodev

2,960 karma · joined January 28, 2019

I like software.
submissionscomments
sosodev··on Umbrel – Personal Cloud
It’s all subjective. Personally I think it would border on useless for local inference but maybe some people are happy with low quality models at slow speeds.
sosodev··on Tesla US sales drop to nearly 4-year low in November
Even if they fire him he'll still have a huge amount of ownership in the company...
sosodev··on Framework Raises DDR5 Memory Prices by 50% for DIY Laptops
Thank you for sharing this. Their point about the 128GB desktop mainboard being a bargain while their prices remain low rings true. I bought one a couple weeks ago because I've been wanting to build a beefy, efficient home server and I think this might be the last window of affordability for quite a while.
sosodev··on GPT-5.2
Yes, a sufficiently advanced marrying of TTS and LLM could pass a lot of these tests. That kind of blurs the line between native voice model and not though.

You would need:

* A STT (ASR) model that outputs phonetics not just words

* An LLM fine-tuned to understand that and also output the proper tokens for prosody control, non-speech vocalizations, etc

* A TTS model that understands those tokens and properly generate the matching voice

At that point I would probably argue that you've created a native voice model even if it's still less nuanced than the proper voice to voice of something like 4o. The latency would likely be quite high though. I'm pretty sure I've seen a couple of open source projects that have done this type of setup but I've not tried testing them.

sosodev··on GPT-5.2
It specifically says in the architecture docs for the agents platform that it's STT (ASR) -> LLM -> TTS

https://elevenlabs.io/docs/agents-platform/overview#architec...

sosodev··on GPT-5.2
You can test it by asking it to: change the pitch of its voice, make specific sounds (like laughter), differentiate between words that are spelled the same but pronounced differently (record and record), etc.
sosodev··on GPT-5.2
Does elevenlabs have a real-time conversational voice model? It seems like like their focus is largely on text to speech and speech to text. Which can approximate that type of thing but it's not at all the same as the native voice to voice that 4o does.
sosodev··on GPT-5.2
Qwen's voice chat is nowhere near as good as ChatGPT's.
sosodev··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
Weirdly, I just tried it again and it seems to understand the difference between record and record just fine. Perhaps if there's heavy demand for voice chat, like after a new release, they load shed by using TTS to a smaller model.

However, It still doesn't seem capable of producing any of the sounds, like laughter, that I would expect from a native voice model.

sosodev··on Mistral releases Devstral2 and Mistral Vibe CLI
I just got my framework mainboard today. I haven't had a chance to set it up yet but from the research I've been doing it seems like Minimax M2 might be the best coding model for it at the moment. Similar performance to Devstral2 with only 10b active params.
sosodev··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
Huh, you're right. I tried your test and it clearly can't understand the difference between homophones. That seems to imply they're using some sort of TTS mechanism. Which is really weird because Qwen3-Omni claims to support direct audio input into the model. Maybe it's a cost saving measure?
sosodev··on Apple Services Experiencing Outage
Aren't all text messages routed through a server? I guess it's more decentralized when it comes to telecom though.
sosodev··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
Is your work open source?
sosodev··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
That's unfortunate but not too surprising. This type of model is very new to the local hosting space.
sosodev··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
I did find this app: https://github.com/gabber-dev/gabber

The creator posted a little demo of it working with Qwen3 Omni that is quite impressive: https://www.youtube.com/watch?v=5DBFVe3cLto

He didn't include any details regarding how the model was running though

sosodev··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
It does for sure. I did some more digging and it does real-time too. That's fascinating.
sosodev··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
I think it's because they've crammed vision, audio, multiple voices, prosody control, multiple languages, etc into just 30 billion parameters.

I think ChatGPT has the most lifelike speech with their voice models. They seem to have invested heavily in that area while other labs focused elsewhere.

sosodev··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
Does Qwen3-Omni support real-time conversation like GPT-4o? Looking at their documentation it doesn't seem like it does.

Are there any open weight models that do? Not talking about speech to text -> LLM -> text to speech btw I mean a real voice <-> language model.

edit:

It does support real-time conversation! Has anybody here gotten that to work on local hardware? I'm particularly curious if anybody has run it with a non-nvidia setup.

sosodev··on Mistral releases Devstral2 and Mistral Vibe CLI
And it can do it, right? I think AMD AI Max line the first realistic offering for this type of thing.

The Apple offerings are interesting but the lack of x86, Linux, and general compatibility make it hard sell imo.

sosodev··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
If you just want to run a local LLM you could download ollama and do it in minutes. You'll be limited to small models (I would start with qwen3:1.7b) but it should be quite fast.
sosodev··on State of AI: An Empirical 100T Token Study with OpenRouter
The open weight model data is very interesting. I missed the release of Minimax M2. The benchmarks seem insanely impressive for its size. I would suspect benchmaxing but why would people be using it if it wasn’t useful?
sosodev··on Are we repeating the telecoms crash with AI datacenters?
I so desperately wish it weren't abandoned. I hate that it's almost 2026 and I still can't get a fiber connection to my apartment in a dense part of San Diego. I've moved several times throughout the years and it has never been an option despite the fact that it always seems to be "in the neighborhood".
sosodev··on Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
I’ve noticed that the open weight models have a lot of issues on OpenRouter. You get a lot of inconsistency in quality due to varying quants at least. I’ve had some seriously nonsensical responses from models that I can’t replicate at all when I switch providers. Lots that just randomly fail to handle requests too. I would recommend finding a provider that works best for your needs and pinning it.
sosodev··on Arcee AI Trinity Mini and Nano – US based open weight models
Because at that point you don't know where the data came from. You could be training on foreign propaganda without realizing it.

Presumably they wouldn't be training on synthetic data produced by anything less than a open frontier model and those are almost exclusively Chinese

sosodev··on Arcee AI Trinity Mini and Nano – US based open weight models
If the performance is comparable to Qwen3 in practice that's quite impressive.

Half the dataset being synthetic is interesting. I wonder what that actually means. They say that Datology needed 2048 H100s to generate the synthetic data. Does that mean they were generating data using other open weight LLMs? Seems like that would undermine the integrity of a "US based" dataset.

sosodev··on I don't care how well your "AI" works
I'm fairly certain that it's happening right now. There is no threshold that LLMs need to "break through" to see adoption. The number of non-technical using them to write software is growing every day.

I was working freelance through late 2023 - mid 2025 and the shift seemed quite obvious to me. Other freelancers, agency managers, etc that I talked to could see it too. The volume of clients, and their expectations, is changing very rapidly in that space.

sosodev··on I don't care how well your "AI" works
How do you know their skills and knowledge are declining rapidly? Does using an LLM cause one to suddenly forget everything?
sosodev··on I don't care how well your "AI" works
While that is part of the equation it's not at all that simple. If the average business owner wants a custom piece of software for their workflow how are they getting it now? For decades the answer would have been new hires, agencies, consultants, and freelancers. It didn't matter that most software boiled down to a simple CRUD backend and a flashy frontend. There was still a need for developers to create every piece of software.

Now AI makes it unbelievably easy to make those simple but bespoke software packages. The business owner can boot up Lovable and get something that is good enough. The non-software folk generally aren't scrutinizing the software they use. It doesn't matter if the backend is spaghetti code or if there are bugs here and there. If it works well enough then they're happy.

In my opinion that's the unfortunate truth of AI software development. It's dirt cheap, fast, and good enough for most people. Computer's couldn't write software before and now they can. Obviously that is real devaluation, right?

sosodev··on Implications of AI to schools
Ironically the practically of such instruction goes down as the status of the school goes up. I got a lot of 1:1 or 1:few time with my community college professors.
sosodev··on New magnetic component discovered in the Faraday effect
But there’s no quantum explanation of gravity, right?
← PreviousPage 6 of 25Next →