Meanwhile my friends kids are turning 16 and inheriting the honda accord with 250,000 miles I remember getting shuttled to and from soccer practice when I was 12 in the 90s
3,720 karma · joined October 6, 2017
Meanwhile my friends kids are turning 16 and inheriting the honda accord with 250,000 miles I remember getting shuttled to and from soccer practice when I was 12 in the 90s
There is still a cost to the business owner, they have to find a place to hire a freelancer, interview and then explain what problem they need solved, go through several iterations with them (either via meeting or email, or both) etc etc. If you're a small business that's not setting up an online storefront, the time spent setting it up and building it yourself is possibly less time. And then there's the bigger headache: what happens when the freelancer stops returning your calls? For the average small business website/page it is 2-10 pages. The average business owner can probably crank out that site with AI in less time than setting up a single interview.
As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off.
Fake edit: there is a longer video here 1 min long: https://github.com/PhiloLabs/fable51-worlds/blob/main/union-...
I would be especially curious to see the NPC person/car logic and if they're on rails or what, that's a pretty good NPC density for a demo.
Qwen: Looking at you for a new ~35B MoE! Please and thank you
Wether or not the most recent example is the best example, doesn't matter. What matters is when the government says "jump" in legalese, google's lawyers say "how high?"
We blocked out ~2 months to build out a testing system for our internal agentic business processes, and I just found out there's a startup with $50 million in seed funding to answer this exact question. I'm sure others will be working on this in the future, as the preferred model for local LLM seems to change every ~6 weeks.
This appears to be using mjlab, which is is MuJoCo Warp plus rsl_rl, which uses a completely different design/technology heritage from Issac, so that's worth a lot to me, since I've basically sworn off Issac for robot gait training.
TL;DR why would I accept someone elses' vibe coded tool, when I can build my own to my exact specification?
Of course enormous batch jobs are different. I was explicit when I said consumer laptop.
Yes, but writing also helps you lens your thoughts in a completely different way. This is why so many people journal, or blog, or vlog, even though only a handful of people might actually audit their output. Obsidian sort of abdicates a lot of the value of journaling, but I think there's still some (although much less, perhaps 5% as much) value in the actual act of directing the journaling, even if the rest of the thinking is outsourced to the clanker.
we went from 62% completion to 92% using a claude code harness
Me: you exist so I can launch iTerm2, Chrome, and VS Code
MacOS: oh, my god
I can't ever imagine using walled garden graphics api in 2026. Particularly for work tooling