1,886 karma · joined September 12, 2014
Working on FOSS and user-friendly alternatives to things like khanacademy, anki, MathAcademy, Alpha School, etc.
Modern, open edtech tooling.
Also http://paritybits.me
With this framework, I'm making (among other things) an early literacy app at https://letterspractice.com. My aim here is to hit >= 75% efficacy of Mentava at <= 1% of the price.
The app is near to production readiness, and I'd be happy to share access now with anyone who has verbal but non-literate kids. Be in touch if interested at colin at letterspractice.com
Nobody believed or suggested that GPT2 could do longform or produce novel text that stood up against careful scrutiny as insightful or well informed. But because the capabilities were novel, people would have no strong alternative than to believe some person wrote it.
You current tripping over LLMisms is irrelevant. You have years of antibodies, both personal and herd-immunity (eg, the many, many articles and comments that describe LLMisms).
The "danger" was in terms of spam / misinformation proliferation, not the same category of capabilities adjacent risks current discussed.
You can hold your own opinions on spam/misinformation as a problem, but to say there was no credibly anticipated outsized downside to a sudden jump in human-passing text generation feels pretty off to me.
changing the CSS - $0.05
knowing which CSS to change - $30
The entire premise of a system is that it removes the need for careful attention.
system: signal lights tell me whether or not I can pass through an intersection, so that I do not have to attend to potentially high speed traffic from a variety of directions.
system: the side my knife blade sits on my arched guide fingers, so that I do not have to attend to the edge of the blade or the location of my fingers.
etc etc.
autonomous replication: https://en.wikipedia.org/wiki/Computer_worm
nb that writing your own quine remains in general terms a fun and challenging exercise in many programming languages, but not python.
I try to imagine myself long ago, on the outside looking in, with someone explaining to me that extreme pain, wondrous art, hunger, triumph, and despair would all unfold in due time where the rocks were wet and the lights bright enough.
I can imagine myself calling this clear nonsense.
Not to be confrontational, but the OP assumed no such thing. OP asserted that it's important for Claude to have the qualities - not that it's important for Claude to present as-if it had them.
Yes, Anthropic is compute constrained, even after the SpaceX Colossus deal.
But supply constraints are the normal operating mode of any market. Anthropic could choose to serve whatever models it pleases at whatever price points it chooses and let the market decide where the value is.
If Mythos at $X overwhelms their capacity, they could just charge $X+1. If still overwhelmed, there are larger prices as well.
to be clear, I don't mean that Claude refuses to read AGENTS.md files, but that Anthropic refuses to bake AGENTS.md into the harness as a first-class feature.
Claude the model by now has learned what an AGENTS file is, and itself is not petty enough to ignore it for partisan purposes, but claude-code the harness doesn't natively support it.
I doubt there is any large demographic of users paying subscription fees for the joy of abusive role play.
I saw a video a while back on one social media site or another where someone sitting in a car recorded three young men shotgunning some beers on an apartment balcony. The insinuation being that hanging out was cringe, and that the poster had caught some losers in the act.
It's hard to gauge "real" general sentiment from social media, but if having a beer in a slightly silly way is the level of vulnerability at which you can be recorded for public ridicule, it's not hard to empathize with a generation reluctant to reach out for connection.
We just meant in the city of Cannes.
Are the dividing lines around personality? Working domains? Opinionated software stuff?
Who knows?
I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5).
So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that I don't firmly grasp any capabilities improvements over my memory of 4.5, but it's all so fuzzy that it's truly difficult to tell.
Maybe my own tastes are saturated now (it's smarter than me?) and I'll never again perceive model progress. Maybe the incrementalism is such that I'd notice immediately if my 4.7 workflows were redirected now to 4.5.
Difficult spot for the labs to be in because, if they have a stronger product, I'd prefer they release it and that I can use it.
But as this dynamic continues, the improvements are going to be less and less legible for end-users, who will complain about the churn-without-payoff, even when the payoff may actually be real.
At a glance, they're providing an interface to YT sourced content with some value adds around tracking or categorizing listening.
A quick question for users: can the site itself be configured as a listener without streaming / displaying the video? In general, YT has a lot of music, but the perf hit of streaming typically high-quality video as well is a blocker when doing dev work on my main machine.
For a while I expected there could be a good return on a good implementation of this, but now as soon as a strong interface itself is created it seems easy to copy.
On balance, and via dictation, it feels likely to be faster overall to just enact the changes I want 'inline' of the conversation thread.
Is this stuff any better now? I think current harnesses probably do have things like file change listeners that automatically inform agents before they act on a file they've previously engaged with if it has changed in the meantime.
On average, I wouldn't expect the regulatory agencies to be very friendly toward outright fraudulent reporting from Waymo. On the very outside, maybe these 80-90% reductions are optimistic roundups from 50-65% reductions. Or do you believe that Waymo is secretly running people down and scooping corpses into their trunks?
What is a sedentary pace of driving?
I don't intend this to read as pure snark, but someone's abstract value isn't much good to them if the job market itself can't / won't recognize it.
We currently sit in the ballpark of 300,000 pedestrian deaths per year worldwide [2]. You should be relieved every time they deploy to a new city.
[1] - https://waymo.com/safety/impact/
[2] - https://ourworldindata.org/data-insights/more-than-a-million...
https://waymo.com/safety/impact/
Over a given driving distance, compared to humans, Waymos produce a 90% reduction in serious injury, 90% reduction in pedestrian strikes, 83% reduction in airbag deployments, 85% reduction in cyclist strikes.
The lack of compassion that people display here is shocking to me.
"Don't automate science, because there are junior scientists could be denied the thrill of specific discoveries."
Cancer patients are not accessories to anyone's self-actualization.