884 karma · joined April 18, 2018
Not sure these guys realize that the quality and latency of those Apple services in MacOS is way lower than SOTA and not too many people use them because of that…
And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...
0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
From time to time, Google ads sales reps call me and are trying to convince me to increase the ads spend. I complain about bots, and none of them had any suggestion on what I can do to appeal it etc. I mean I can appeal on some specific instance, once - but there's no process to communicate back to Google the cases where I'm 99% certain about bots on regular basis.
And these bots get increasingly more sophisticated. They used to just click on ads. Then they started installing the app. Then running the app once and do nothing in it. Then they started tapping on the app screen & actually try to go through some initial app steps. When they do it in a spiky way, it's easy to detect them. But when they do it through some distributed device farms, below noise, it's very hard.
As for training, I'd be highly surprised if they use the data they get from APIs for training. But the conversation itself & thinking trajectories probably will, unless you opt out, like you pointed out.
Eventually, though, Muse Confidential VM is an even more locked-in option coming, per that blog post...
My personal anecdote - we had a complex family travel in the summer, with 6 ppl and 6 separate flights booked (some multi-leg). SAS messed up on one booking - they decided to cancel the flight we had from Oslo, and rebooked us to an earlier date (2 days shift). That wouldn't work for the rest of our schedule, so I had to rebook, overpay extra money for now longer & more expensive flights - but to add the insult to the injury, when I did this through their website, they didn't transfer the food that I paid for the earlier reservation. Calling them in roaming and trying to get it fixed with a human support agent didn't help - the agent said he just can't fix it on his end, and suggested that we just file for reimbursement on their site.
I never had time to file it since July.
But today I had a reason - to test Muse - and asked it to handle it for me. To my surprise, it dug through all of the SAS emails and actually found the emails confirming that the food fees were reimbursed exactly the same date when SAS moved our original flight. Apparently I didn't notice those notifications since they were spamming with all kinds of notifications that day. But, if not some tool like Muse, I wouldn't even find time to handle this and either file a claim or figure out my miss on their earlier reimbursement.
Quick question - does it really need external SSDs, or if the local SSD fits the whole model - how fast the model would be? e.g. on your machine, M5 Max 128GB, with 4TB SSD? maybe it'd be good to add "0 external SSD" column on your graphs?
For some of us here, it's just what we love to do, no matter what tooling is available. When I first started building my own software long time ago, it was a very slow Basic and fast raw machine codes (in octal system, PDP-11 like CPU). I enjoyed it not because of tooling, but despite of it.
Over the years, the tooling was getting better in general, which allowed us to build increasingly more complex systems.
With AI, we will still be creating & debugging. It's just that before AI, I had to spend 90% of my work on mechanical not-so-fun things to get things to work, and only 10% on fun algorithmic-intensive parts. But with AI tools, this ratio seems to change, and all kind of boilerplate code & algorithms can be written much faster by AI, hopefully leaving more time for us to work on creative part of the work.
there're different dimensions for "scale" - like handling large monorepos, orders of magnitude more commits, tighter requirements for latencies (for agentic use, e.g. for agentic history navigation)...
Just give people some benefit of doubt. There're much simpler ways to explain certain things that suspecting some universal evil in every move...
Most of the glitches I heard with OpenAI's Voice were not WebRTC related - but rather, to my ear, they sounded more like realtime issues with their inference - which is a very different component to optimize.