2,813 karma · joined January 1, 2022
Unless Dots is dramatically more capable than Muse, I'm also more bullish on Muse than Dots. I think Muse is a better consumer play because it can be forever subsidized by Meta ads and find distribution in family of apps while Dots is in a weird place between consumer & professional. From the release, it also sounds like you'll have to pay per Dot at some point which doesn't sound appealing.
I don’t think this is normalized at all.
Something about the ad-riddled website and way this is written makes me wonder if this is even a real article.
I find the most consistent tells are often just excessive descriptive text and a lack of visual hierarchy. It makes sense given these are language models.
Also, I find models are really quite good once given some design context to work from. Slop UI is lazy at this point, not a reflection of model quality.
Re: the captcha solver
> As far as we can tell, agents eventually abandoned this approach and were unsuccessful in generating Hugging Face user accounts from external endpoints.
I wonder how the swarm eventually decides to abandon an approach.
Model update Input Output Reduction
------------------------- ------------- ------------- ---------
GPT-5.6 Sol → GPT-6 Sol $4 → $2 $20 → $10 50%
GPT-5.6 Luna → GPT-6 Luna $0.20 → $0.10 $1.20 → $0.50 50%> Output tokens: FREE (too cheap to meter).
Insane. The video demos are really compelling, in particular the speed.
> Structured outputs slot into ordinary software as fuzzy decision rules: classify, route, score, extract, or branch where hand-written logic is too brittle. The surrounding code constrains their freedom, making them easier to compose into reliable systems.
I buy this vision. A lot of LLM integration I see these days is ultimately exactly this. OpenAI-style structured outputs works decently but this would be a great improvement in cost, latency.
And the agent found the token in Docker build history after finding a Baseten image repository.
I wonder how many of these kinds of agent-driven security exploits we're not hearing about these days (i.e. driven by bad actors), worrying.
https://andonlabs.com/market - $25,098 all-time revenue
https://andonlabs.com/cafe - $14,343 all-time revenue
The shops aren't doing particularly well (i.e. they don't seem to be turning a profit over time), but it's an interesting trial and benchmark.
Congrats to Temporal! They'd just previously raised in February, https://temporal.io/blog/temporal-raises-usd300m-series-d-at...
> Finland has emerged as an attractive location for data centres because of its cool climate, plentiful low-carbon electricity and relatively uncongested power grid.
> Four times a year, we put a hand-picked group of YC's enterprise-AI founders in a room with buyers who can actually run a pilot.
Seems like a win-win.
Definitely shows how important a user data flywheel is for RL and model improvement.
Too good.
Seems like reasonable criteria. "Regular" unemployment is
# unemployed/# employed
where unemployed is they do not have a job, have actively looked for work in the prior four weeks, and are currently available to work. But not much criteria in terms of what work people are finding.
> The slop-acalypse that so many of us fear, when people will stop looking at anything made by humans and only look at things made by machines—why hasn’t that happened yet?
Unfortunately, I'm pessimistic here. I don't think the slop-acalypse is held back by taste or an appreciation for "thickness". I think there's already very successful slop media in the world (e.g. see microdramas). In other words, it is happening, but like anything else, it's happening in a non-evenly distributed way.
Interesting, is Claude becoming the Kleenex of the agent world?
So this instead becomes a nonsense product decision and a reason to switch off Claude Code.
Could this be fixed with better harness restrictions/tool sandboxing?
Echoes of WhatsApp. Huge congrats to the team!