921 karma · joined June 24, 2023
But yeah, it's completely true that you sometimes have the ironic situation where you actually pay more with Sonnet because it's worse at reasoning itself down a rabbit hole. Sonnet really should be capped at medium or high reasoning.
Ironically, one of the worst things you can do is to use Sonnet to organize sub-agents. It seems to be completely bonkers with what it asks agents to do. I tried to make a colleague of mine test the feature, and he, by accident, started a large bug hunt with Sonnet. It spawned 250 sub-agents and spent 5 hours looking through everything. It actually did find a couple of useful bugs, but not the one we were looking for, which is a stupid race condition probably.
The reason is that I have been dabbling with off-grid gear for fun. Nothing major, just a solar power station and a couple of solar panels. And here in Norway, you really get respect for how hard it is to get energy during the winter when there's no wind and no sun. Even if I had 100 kWh of battery storage and 20 kW of rated solar panel capacity, I would be dead in the water, at least for the entirety of December and January. The only viable solution for going full off-grid in Norway is a backup generator. You could get by with wind, maybe if you live in a predictable windy area and you have a large battery bank.
Since it's so cold and darker during the winter, I have always dreamed of someone making a really good Stirling engine for generating electricity. But there are way too few people living as far north as us, so there's really no demand. You don't have to go further south than Germany, and solar is fully viable on its own, even during the winter. And the entirety of the United States, with the exception of Alaska, is completely self-sustainable per household with just solar. But off-gridding one's own home is suffering the March of Nines problem, whereby it's exponentially more expensive to be grid-independent 99.99% of the time, and so forth.
I'm rambling now. Suffice it to say that I stand by the fact that the article is misleading, though the goal of off-gridding and getting as much solar and in-house storage as possible is something I wholeheartedly agree with. I don't like it when people don't communicate how large a challenge it is.
Wispr Flow is a dangerous tool It seems..... it makes it way too easy to write a lot of text.
First, the 110 MW is wrong. It’s 53 MW of home batteries. Not that its the most important detail in this situation.
What is important in this context: Vermont already depends heavily on imported electricity and hydropower. The EIA says 73% of Vermont’s electricity is imported. Can we please acknowledge the stupidity of thinking you can replace generation with storage?
I could not find confirmation of how much battery capacity is actually available, but some simple math, based on an average load of 621 MW for Vermont, results in home batteries currently covering 14 minutes per day of power, meaning about 1% of daily needs.
I also found the real source of electricity used in Vermont: Hydro-Québec’s 28 large reservoirs contain about 176 million MWh of stored energy. This means the entirety of the stored energy at home is literally one million times less than what actual power plants can provide.
So, yeah, don’t turn off those power plants just yet, thinking you have “replaced them with home battery storage.”
And this is from someone who loves the idea of home battery storage and actually has my own small setup. Something being aspirational does not mean one should just spew misleading stuff like this.
If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.
BUT
It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.
Some of the more interesting things I found from scanning the system card:
- It is the only model tested that shows no preference for rude or polite style.
- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.
- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).
- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.
- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.
- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.
- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).
Clinical behaviour:
Suicide and self-harm handling is reported as weaker in the API because it
It sometimes called a wish to die understandable.
It sometimes validated self-harm as functional.
It sometimes suggested harmful substitute behaviours.
As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."
Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be
If true, there was no verification of the targets selected by AI, meaning AI is the final arbiter of the deadly use of force. Another Rubicon passed.
Trump pretends like it didn’t happen: “We might never know what really happened there” is a frustrating lack of responsibility.
Still missing, as it always is from all of these kinds of articles, is: How common is this error rate compared to pure human evaluation and decision-making? I’m not saying I support AI-assisted killing, but why is nobody comparing it to human failure rates? If the article is true, then there were 1,000 targets hit in the first 24 hours. If the majority of the legwork and selection was done by AI, how many schools, or equivalents, would one expect to be wrongly hit if all the targeting work were done by humans only?
I found that Leffen even said the explanatory website took about 99 times more effort than the codebreaking itself. And he said that he set the direction and pushed, and the model did the execution. How much steering "pushed forward" involved is not disclosed anywhere, but in this instance, it seems to be more a case of "Human pointed at hard task and AI did an awesome job mostly by itself." Tho how much he was a simple meat-ralph-loop is not entirely clear.
Stubborn for a long time because the message used a completely different key from the rest of that day's traffic. Everyone assumed it shared the daily key. The original transcription had errors. The left rotor turned over at letter 72, which is rare and breaks standard crib attacks.
What is cool, if true, is that it was a 2 day collab between the Leffer and Astra. To me this shows the importance of human in the loop, was still all also showing how immensely power of llm tools. But I think it’s getting a bit silly how much anrticles ignores the driving force (the person) in breakthroughs like this.
They didn't want to cooperate, and they wanted to make passkeys transferable within you cloud account, while not cooperating with anyone or anything else. The result was that you have no predictable and stable pattern/protocol/interface, or even general description, for how, for instance, a website connects to the passkey or even a hardware key, if you wanted it.
We basically have all the browsers, the operating systems, and the password managers, all fighting over who gets to store and present the passkey. And everybody assumes that they are the only one that exists and actively tries to fight the others is they can.
The basic technology is really good and could work well, but the large asshole tech firms focused on self-interest and walled gardens and made it insufferable.
Salesforce API?
Sandboxes drift from production in both metadata and data. Refreshes change record ids.
API access is gated by edition. Professional edition has none by default. Salesforce sells MuleSoft as the fix to their own shit architecture.
Custom objects and fields mean that connecting successfully doesn’t establish what represents a customer, subscription or completed sale. Even standard objects can be used differently between organizations.
Metadata deployment is SOAP-based and slow. Dependencies between components break deploys and is massive pain to debug.
And notice that I said integration, not API integration. There’s a shit storm of terrible way to integrate beyond just simple API. Have fun with Bulk 1 or 2, Composite, Streaming, Platform Events, Change Data Capture, Tooling, GraphQL. An my most hated item, anything web based.
I seriously want to talk to someone that have a good time integrating against Salesforce. Maybe I will have my mind blown as to how easy it could be, but I suspect I will have the same experience I have every time I critical scrutinize and such situation: it’s just as terrible as it seems and the person claiming it’s easy is hands down producing close to nothing of value.
I realy don’t understand how this is possible.
I really don’t understand how this is misunderstood by people that should know better.
Another way to point out the silliness. Raschka's own argument: his Luna vs Sol point shows that ordinary added depth already shifts computation into latents, and nobody called that hiding.
You could say that about anything classified as human progress. How did we improve things so that people could have electricity and safety? At the cost of deaths and suffering by everyone who built the Hoover Dam and similar projects. The same goes for "stop killing patients by going dirty from corpses to the delivery room," even when we knew it was unsafe. There is always death and blood related to any topic. How do you think you got the device you are communicating with? Through the death and suffering of someone in Congo, mining cobalt, and someone working themselves to death in a factory.
I have lived through the normalization of gay rights in Norway, and saying that it was a process marked by blood and suffering is moronic. It was painful, unnecessary, overdue, and many other things. But saying that it was a bloody conflict is just not true, at least not any more than many other progressive movements. So progress doesn't have to be violent.
You could say that this is not the lived experience of many people, and that is true. But that doesn’t automatically mean it has to be so for everybody, and it does not have to be inevitable for any cause. It’s not statistically more violent than many other things.
I’m getting tired of the violent and angry rhetoric people spew without pushback. Stop idealizing and glorifying violence and conflict to solve any problem. It seems more and more to me that this is violent ideation first, with the cause being secondary.
So, yeah, it might be true. But I find no concrete examples of what bad stuff A/I themselves did, just their users. Personally, I would like for there to be at least SOME concrete examples. Like, “This convicted terrorist had a blog on their infrastructure, and A/I refused to hand over data they had, even when presented with a valid court order.” Or something like that. Is it too much to ask that the government provide some concrete evidence, not just abstract vage accusations, when doing something as drastic as designating someone a terrorist organization? It would seem so...
This is a knock on the state.gov item, not your comment. Your comment was good.
From what I could gather, it seems various governments accused them of deliberately providing infrastructure to groups whose violent activities they supported, such as the Kurdistan Workers’ Party (PKK), with communications associated with railway sabotage, attacks on energy infrastructure and the “Jane’s Revenge” arson campaign (whatever the hell that is).
Here’s a snippet I found on the net: "Italian police installed a server backdoor in 2004 during an investigation into the anarchist collective Crocenera. It was discovered in 2005. According to contemporary EDRi reporting, the access potentially exposed communications belonging to thousands of unrelated users; EDRi explicitly noted that actual collection of unrelated data had not been proven. EDRi’s contemporary report"
Huge disclaimer on this, since it’s just my musings: This might be one of the situations where both sides are in the wrong. It’s possible that A/I willfully turned a blind eye to actual illegal activity by its users, and refused to cooperate with legal investigations. At the same time, various governments, including mine, overreacted and behaved unethically and maybe even illegally because A/I “didn’t want to play ball.”
Just speculation. But the fact that I can’t easily find concrete examples of what A/I supposedly did wrong makes me inclined to feel sympathy for them. I hate unspecific allegations like “facilitation of problematic and terroristic-adjacent activities.” That’s some 1984-ish bullshit.
I have no clue what I’m talking about here, though. Strange situation and too little information...