very curious to see more about what kinds of hardware you can run this on and the perf. characteristics… on the face of it, it seems like optimizing for inference speed might(?) be good for running on smaller hardware, but i suppose it could be the other way around and it is actually much resource-hungrier for the number of parameters, etc. …
my projection is that they are still gonna be pretty far behind, but they will sew it up in the next few releases. it feels like they were caught with their pants down on how much work OpenAI has put into that area, but i doubt there is some magical secret sauce that OpenAI has that Anthropic simply cannot catch up with.
completely agree. i will admit to taking llm phrases when i am really trying to refine every last detail and it just makes a suggestion that is too good to unsee. but in general it doesn’t feel like it helps directly with the “word choosing” part of the writing task at all (if you are someone who cares about word choice), which is… definitely a pretty big part of the job, lol.
it is great at analyzing the argument, finding inconsistencies, helping you think through what parts should be cut, helping you refine examples or fix the occasional “how do i get this phrase to work correctly in this transition?” kinds of stuff. but anything where LLMs are the prima materia… that stuff literally only makes sense _to me_. which makes sense, because it is written _for_ me, no matter what instructions i actually give it, bc of memories and a million other things. and i say this as a complete maximalist wrt. trying to use llms for absolutely every last thing they possibly can be used for, just to see what it’s like.
i guess i would say that it does very, very little to make the writing process meaningfully faster; it _can_ do _plenty_ to help make your output better though, which is definitely something— just isn’t the thing most people are looking for.
yeah am i crazy or could it not like already do this as well? i swear i’ve seen somesuch similar in the cloudflared options. maybe not. but i can second having had some problems over the year getting cloudflared to install/set up correctly. tbf the actual feature works amazing once you get it working.
jesus christ i don’t get this attitude at all. you do realize that software is like 100x easier to make than it used to be right? the world is severely idea-constrained in a way that it hasn’t been since, what, 1992 with the web? i love robust, “hard problem” engineering more than your average bear, but we are so much further down the alan kay “the right pov is worth 60iq points” road than we have ever been in my lifetime, and this is a very non-obvious, creative mashup of things that i directly found inspiring. not because i think the project is perfect or genius or extremely well-executed or anything like that— because it is a creative mix of things, and we live in an age where we can be exponentially more creative with what kinds of atoms we can smash together than ever before.
yeah, personally i think branching out and using something like a game engine for something like this is a fairly inspired choice. inspired meaning that it literally inspired me to immediately think “oh yeah, there really is a wider universe of UI stuff that can push good fps and not be electron”. so just for that alone i appreciate this project’s existence. :)
i will definitely be curious to see how you find that decision to be working out for you as the project evolves— will definitely be keeping my eye on it! actual refreshing, interesting take!
tell that to Doug Englebart or J.C.R. Licklider-- the idea of human-computer symbiosis predates basically all computers that we would recognize as such today. crazy how it is only really now becoming true in the way that they envisaged in the 1960s.
i think we will all look back on Fable as the start of the AGI inflection point. for all i know there are still multiple leaps between now and AGI (i personally am inclined to think that for all intents and purposes we are "already there", but reasonable people can still disagree on that point), but Fable was the first time that something felt genuinely magical about the results themselves, not just particular outputs. which is kinda funny in that i don't know anywhere near enough in terms of behind the scenes as to whether or not there was something meaningfully different, or if it is just the point at which the scale had finally accumulated such that i happened to notice that the output was fundamentally different.
i can't wait to dig in on 5.1 because while i have always been somewhat predisposed to think that openai's models have usually been "better" (my own subjective opinion, that) "on average", i have been kinda tired of the regime of late where it felt like Anthropic was miles behind while simultaneously clearly having models (Mythos) that are surely face-meltingly impressive-- it has just been very hard to square with the fact that i feel like Anthropic hit the "real" "critical point" first... i have no doubt that 5.1 will finally reset the ecosystem balance into a more healthy place.
i think there definitely is some truth to this in terms of embeddings spaces, which is why i believe they are implemented by OpenAI/Anthropic in roughly highest import => least import bit order-- an overwhelming majority of the variance is in the first few hundred vector bits. i haven't actually tested this myself by manually truncating vectors, but it is my understanding that they generally speaking have this property.
i read this before the article and thought “alright it can’t be that bad”… honestly, you are letting them off easy. it’s like pedantry, but the dunning-kreuger kind.
people who throw out “x is a fallacy” very seldom have anything interesting to say, and this point is no different. the thing is not released yet. there does not exist any easy way to communicate how strong any particular llm model is en toto, so people inevitably look for stories, metaphors, and the like to help explain the story. saying that it is “vastly oversold” is comically absurd for a model that has not been released yet, especially through the lens of a completely unfalsifiable framework for contextualizing it. here, i’ve got a “fallacy” for you: this whole post reeks of “no true scotsman”: it is bold to claim that the model is oversold in its abilities when it has racked up this many novel proofs before even being released widely, but hiding behind “that doesn’t mean it is good at ‘math, generally’” is absolute weasel language— it invites proof-by-example in a way more flagrant and devastating than anything the author points out about the discourse around the model itself.
yeah, the urge to turn this page into a skill and then just let it run loose on all my random vibe-coded side projects is very, very strong. seems great for that kind of thing.
wow this is such an excellent framing of something that i have not managed to think about or come across in the vast ocean of discourse about LLMs. i had not really thought about how explicitly they are designed to not have histories that are user-interpretable. it is really quite not unlike how Apple locks down iOS and MacOS-- there are UI/UX reasons for the choices the frontier models make, but that really is only part of the story, and these choices undeniably do conspire to make the resource more locked-down than it absolutely has to be. and the analogy to "people don't switch OSes every day, but the ability to switch OSes changes your relationship with the provider" is extremely apt.
trying to get off vaping, myself, and it is both "not that hard" and... a very sticky habit. good for you. was never a smoker, but that doesn't make it a good habit.
whoa. feel bad for the people of tuscon, because they didn't do anything wrong, but holy shit this case looks absolutely awful for the city and i hope they lose. big.
ah, so the soft implication that you are simply smarter than me. fine. there's at least a 1/100,000 chance of that. no one forced you to be that smug and dismissive, though-- seems like a few traits that are fairly antithetical to learning. but i guess you don't need to learn, since you have already ascended to Enlightenment.
surely this is closer to the common experience with LLMs... i don't know how people don't see it for the miracle it is. it's like if God himself came down from the heavens and said "no one will ever die of hunger ever again, but you have to run 5 miles a week" and we all got bogged down in how hard it is to run 5 miles...
that's a great point. i personally am toward the "i go to sleep and it does my work for me" side, but share the same confusion about how we went from 100x-0.1x programmers to 100000x-0.0001x ... vibe coders. to be clear: i am aware of what my side of the fence vs the other looks like, and i don't find anything about "meh iono it doesn't work that great for me" unreasonable at all.