249 karma · joined January 20, 2022
What I meant was really about the output tokens, on the API pricing level of interaction with an LLM.
As for token-wise inefficiency, this is based on my sporadic observations and discussions with friends and colleagues from the past three years. I can’t give you a fresh benchmark with latest models in a reproducible way. But I’m happily willing to accept your assertion at face value. This makes me curious to see for myself how the latest models perform, will perhaps set up a quick evals just out of curiosity.
From my past experience, the latest I’ve seen were Opus-4.6 and GPT-5.5 struggle a lot with standard GHC Haskell, no extensions, nothing fancy. Not that they produced impeccable Python or Rust. But it appeared to take several turns for obvious expressions, while at Java and TypeScript they were much better - fewer turns, time, and cost for the same verified results. I had assumed that the training set wasn’t large and diverse enough for Haskell and Scala, and this was the go-to explanation with everyone I talked about it. Generated F# code was good enough but not OCaml. With C++ there’s still quite some struggle, expectedly.
To me it resorted to the question, if we’re mostly generating code through LLMs now, which ones and what will it cost in total terms per task/capability completed. I’m wondering now if DeepSeek/MiMo or GPT-6 Luna can benefit from the terser, more (forgive the pun) load-bearing expressions.
In this new AI-driven world, is there still place for such a luxury as functional programming?
I mean few people still code by hand, few read the generated code, models aren’t trained on functional languages, it’s inefficient token wise to use functional languages - while a lot become self-proclaimed software engineers overnight by just prompting LLMs.
Agreed on the “being dumbed down” observation. It appears they’re most powerful at release time and then are gradually “optimized” so every new model feels more powerful. But there’s no evidence on routing to a deployment with other weights. It would be plausible to do so though at least at peak times.
I get way more usage for way less money without any quality or performance degradation. My $200 Codex Pro plan allowance is depleted in 2-3 days. Sometimes Tibo announces a usage reset. But GPT-5.6 models are really not good for coding. Sol has been making increasingly more mistakes in the past two weeks even in the reviewer and advisor roles. Astra is usable for coding but slow and very expensive. In the past two days I’ve used up over 70% on simple copy editing, with dedicated short specs and short sessions. Really little one can do to make it more efficient. Similar work took 20% at most just a month ago. I’m looking to use Astra for milestone reviews/advisory. Perhaps a $100 Pro downgrade will be enough. But my main work is now on open-weight models. And you don’t need to depend on someone to send you a reset. And it’s cheaper by the end of the month too.
With Claude the limits are not even fun anymore - my weekly $100 Max plan quota is gone in one day on merely review invocations, no coding. And my $200 Pro quota is gone in two with some coding. Sonnet 5 is not usable for coding. And Opus 5 tends to always make a couple avoidable mistakes on every task. Fable 5.1 is ok but tends to ignore skills and to work around explicit instructions. Completely canceled all my Claude subscriptions.
With Qwen 3.8, DeepSeek 4.1 Flash, GLM 5.3 I’ve been getting Opus 5-level performance, with less blah blah and no overengineered churn. Public benchmarks are really not telling the real story. The models are more dependable and more predictable. They have their own failure modes. Sometimes DeepSeek 4.1 Flash is quite stubborn but it fails in a good way. Bad for full autonomy - I need to intervene, but it sticks to the rails and instructions - other than Fable and Opus that try to outsmart you and your harness.
Grok is interesting but has been a bit underwhelming on Grok plans - my SuperGrok allowance is depleted in a single session overnight. SuperGrok+ gives more but it’s still about the same as with OpenAI, Claude is way less now.
Since the allowance volume has been shrinking with the major model providers, to me, open-weight alternatives are really necessary now to at least maintain the momentum and budget.
But at work it’s really an uphill challenge - it’s become impossible to convince the tech leadership once they got hooked on Anthropic. They No facts will help. Some people underestimate how expensive Claude really is after getting used to the subscription plans with allowance resets. OpenAI models are expensive too.
I used to rely on Fable for research when it was first out, today it doesn’t seem to be much better than Opus, and it uses up the quota exceptionally fast - 1h Fable in a single short session, and there’s little left for Opus to hit the 5h limit in a second session. With Opus I get about 3-5h of relaxed use with a couple subagents to save the context, but there’s usually quite some disagreement between the subagents and orchestrator - Claude does some model routing with default agents and picks Haiku and Sonnet for subtasks - only later to disagree with them and redo the work - and burn extra tokens. With Claude, it’s really either Opus or Fable if you want some quality.
That said, their marketing is exceptionally effective. Virtually all nontech folks consider only Claude.
Im not convinced to pay $200 for Claude’s models.
With Claude, I have to intervene every 15-20 minutes, it’s non-autonomous and it’s incredibly unreliable at self-correction. GPT is strong at self-correction but it tends to drift away from the plan to self-correct in a loop very often - a lot of tokens and time burnt on aimless churn. Opus tends to push its uninformed opinions and fake retrieval, drifting every turn increasingly farther from the intended and approved design. Opus skims over specs and makes too many mistakes.
As for closed frontier models, I prefer the GPT models over Claude’s.
I’ve started relying more on Grok, GLM, Kimi and DeepSeek models for subagents - I’ve ended up with a factory and am seeking to reduce my reliance on the closed frontier models - they’re just not SoTA on their own for development anymore.
It was discussed on HN a couple months ago. That one guy then went on Twitter to boast about his “high-impact PR”.
Now that impact farming approach has been mimicked / automated.
Another banking app has failed to identify me a couple of times (I attribute it to iPhone 17’s front camera distortion) and fell back to the snail mail id code as a 2nd factor. It arrived only several business days later. Instead of just letting me use my own 2nd factor such as a TOTP device or a physical security key. But maybe there are some legal requirements for that flow, I’m out of the loop.
So there’s a whole range between passkey-is-enough on one end and outsourced video id or snail mail for 2nd factor on the other. The latter can of course be misused to siphon as much personal information as possible out of you, even linking and scraping your other banking accounts for consumer profiling - designed as a requisite part of the authentication/authorization flow.
Sometimes it works with the front camera on one smartphone but doesn’t with another (iPhone 17’s distortion), sometimes it recognizes your face on one day, but desperately fails to recognize you on another. I had to repeatedly record videos for it only to fail over and over again. Anything their system flags as suspicious, anything, will trigger the same video identification flow again, which effectively blocks your money in the account.
I’m closing my accounts with a couple of banks with these video id flows. Simply because it’s way too easy to lose access to my money in the account with them. If their QA is not good enough for this vital requirement, I don’t want to know how they treat other requirements. They simply outsourced the id verification to some third parties that are way too unreliable.
I agree that the privacy controls on Apple systems are well-organized.
Still, it’s more important to have confidence that the privacy services are not smoke and mirrors with carefully carved-out loopholes. It’s one thing to provide something and hold the competitor as the litmus test, the other to sustainably live up to your promises, like the now pejorative “do no evil” slogan, with retroactive ramifications. There’s really little users can effectively validate about Apple’s privacy promises.
Then there’s a big disparity across all Android hardware vendors. Google must cater to that more or less federated topology of Android devices. It’s much harder.
Yet I don’t see any technical blocker for an opt-in for an Apple-grade ADP in Pixels and Galaxies.
It’s all quite weird. Even with Google Passwords, how do I know that it’s E2EE if I can unlock it from a browser with just a device PIN? Lots of loopholes.
Google pushes Gemini everywhere and wants to keep on to your interactions, with human reviews. While I applaud the transparency, having Gemini scrape my screen makes me uneasy. My frog’s not warm enough for that, yet.
And Gemini in Sheets and Docs is just a toy. Microsoft 365 Copilot is a step ahead but is wrong more often than not, at least from my interactions with them. Both very disappointing. No way to justify access to my personal or my company’s or clients’ information.
Apple promises something they call Secure Compute or so, don’t remember the exact name, which appears to be encrypted and randomized in their cloud compute, which is off-device. With iPhone being the most powerful to date (per GeekBench), Tensor Pixels will have to offload most of the edge compute to GCP, and Snapdragon Samsungs while being powerful (I have no idea but would assume) must follow the Pixel Android approach.
So AI features will exfiltrate even more personal information, occasionally, accidentally, or purposefully, and the user would have consented to that and the human reviews just to get access to the smart features.
That’s true. On Pixel Android, there’s several unrelated places in the various settings for the device and for the Google account to take care of and see that they do not collide. And for every function there’s always some sort of small print like “it’s all private to you unless you choose to share” - but to use any of the features/services you have to “share” like with Google Photos and Calendar and Tasks, you lose track of what you share with whom in the end. So essentially not only the metadata is collected but also the content and nothing’s private as a result, at least that’s what I got to understand. And even if you ask Google to delete your personal information, it will retain it for a while for compliance purposes.
As for
> - App developers are mandated to publish what they collect when publishing apps to the App Store.
I believe that’s still moot and rather a voluntary disclosure that no one vets. I’ve seen apps with no collection stated on App Store but deviating privacy policies, or app functions that contradicted their own privacy policy.
From what I heard and read, I understood that as a well-meant idea but still a misconception on the consumer part due to lack of enforcement by Apple.
Both promise security, Apple promises some degree of privacy. Google stores your encryption keys, and so does Apple unless you opt in for ADP.
Is it similar to Facebook Messenger (encrypted in transit and at rest but Meta can read it) and Telegram (keys owned by Telegram unless you start a private chat)?
There are things Pixels do that iPhones don’t, e.g., you get notified when a local cell tower picks your IMEI. I mean it’s meaningless since they all do it, but you can also enable a higher level of security to avoid 2G. Not sure it’s meaningful but it’s a nice to have.
What you’re describing about Windows is very reminiscent of what Pixel users describe on Reddit.
I’m totally with you, I wouldn’t use Windows voluntarily. I’m not in a position to tell whether it’s more or less ready though, just no recent experience with it.
This was the first time in two decades that my smartphone broke, and it could only be replaced.
In the end, to me it’s really too much maintenance with Pixels and Android devices in general. Really don’t get it why people prefer Android. It’s like desktop Linux. Not there yet.
What are all the use cases that let you avoid Play?