Clef seems to be a pretty strong attempt at a high difficulty UX. I've created an account and will be giving it a go. Wishing you luck, and thanks for sharing.
1,886 karma · joined September 12, 2014
Working on FOSS and user-friendly alternatives to things like khanacademy, anki, MathAcademy, Alpha School, etc.
Modern, open edtech tooling.
Also http://paritybits.me
Clef seems to be a pretty strong attempt at a high difficulty UX. I've created an account and will be giving it a go. Wishing you luck, and thanks for sharing.
But naive-compact is forced to just sort of guess at what is and isn't relevant from the prior work.
The harnesses have gotten better at some JIT ui stuff, throwing interview questions / forms at users. Compact is the ideal time for this:
Where are we headed here? (2-5 viable options, sourced from current context and imagination)
Then potential follow-up questions as required, but honestly I expect the single guiding answer there to improve post-compact performance pretty dramatically!
Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud, Opus 4.6 may end up hitting its conceptual ceiling and unable to contribute.
Winding the clock back on your statement gives:
> I'd gladly pay for a Claude Sonnet 3.5 in silicon and use it for 1-2 years.
Man, I dunno.
Token count is a less important factor in context pollution than idea count. The worst of the rot factors are when models latch onto irrelevant information, or over-index on some vague idea/suggestion as if it was a hard direction, and then go off course.
The names + one-line descriptions of 10 tools can do as much (or more!) to distract the focus and intentionality of an agent than a 30k token exhaustive API documentation of some tool.
No future for research mathematicians othet than as tastemakers / agenda setters?
This is a little dicier in post-agent AI, because it's easier for casual users to automate power-user consumption, but the providers have done decently in discouraging that.
> Increasing Corporate Skepticism: The news is full of stories of corporations that are throttling the employee use of AI since the costs to use the software are a lot higher than expected.
Real AI spend is out of control, with the news is full of stories about corporations trying to keep a lid on it, but also real AI spending is low and concentrated to a few firms.
I don't know. This doesn't feel very coherent to me, but rather like a collection of assertions that are adopted because they individually say something bearish about the industry.
Certainly some investments will have been overreaches, but I find it pretty unlikely that any of the compute build-out to date is going to be left sitting idle one or two or five years from now.
With respect to FLT, my hopes have modestly increased that a truly marvelous demonstration of this proposition does in fact exist, that Fermat actually had it, and that it may someday be recovered!
edit: some emphasis on modest. But let me be romantic here!
Claude code interacts with many system processes, files, etc, as well as external APIs. Processes audio via built in dictation. Manages a bunch of nasty auth. Etc etc.
What are the categories of features that wouldn't be exercised by this class of software?
Yes, the mechanics are straightforward if Anthropic (or Claude, if you want to ascribe the decision there) decides to burn a pile of your money. But the strategy fails basic game-theory of repeated games - you'll simply stop playing.
(this isn't to say it invalidates the incentive to inflate token count, but it overcomes in terms of weighing options and making long-term profit decisions.)
It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but I don't see them clearly getting tripped up by them as prior models used to. (eg: Since you're running Ubuntu in Canada, here are some drills you can try to help your kid hit a baseball more consistently.)
I made a lower effort but similar scaffold for LLMs to do iterative drawing in Nov 2024, with Sonnet 3.5 as the artist: https://paritybits.me/llm-drawing-with-eyes-open/
Quite a difference.
Yes this is feasible, but respecting it as a design problem, renting is more transient than ownership and tilts the floor away from deep communal relationships.
Comparing the neighborhood I grew up in with the one I now live in is night and day. My mom has had the same next door neighbors for 44 years. Up and down the street there are many similarly familiar persons.
By comparison, from my own front door, I can only physically see two houses that are owner occupied. There are good neighbors (and friends!) in the rental houses as well, but investing in those relationships pays off with lower certainty because circumstance is very likely to uproot them at any given moment.
Heads up: the "See it live in the showcase → " links in the API documentation do not go back to the showcase - they just reload the same current API section.
Question: maybe I've missed it, but what exists here wrt distribution / packaging / bundling / source availability? I see MIT listed, but no repo. I see src="https://jelly-ui.com/package.js" as a sourcing import, but obviously I'm not going to bundle foreign assets into my app.
Apologies for the bad example. Replace w/ gain of function / whatever else, or just brainstorm with your local model, ect.
Even at the level of, say, Opus 4.5+, open weight models give a quick turnaround to every Joe and Jane on earth having easy access to pretty high quality improvised weapons design, cyber / auto-fraud capabilities, etc.
All the existing models (closed and open) put up decent resistance to participating in activities like this, and especially behind API walls with content monitoring and account bans.
But the published open-weight models can be fine tuned or abliterated into arbitrarily sharp-edged tools. EG, if it's physically feasible to build a nuke in your garage, it may soon be the case that more or less anyone will have competent guidance to do so.
Although saying that out loud makes me question it - each per-user chat and growing cache would need eventually to own its own ~contiguous memory block.
Push notifications.
> Integration with OS features is what made the app ecosystem, because of utility.
This is true of some apps, like the beer-drinking one that uses the accelerometer / other orientation sensors.
It's not true of a large number of other apps, hence the "your app could have been a webpage" charge. This is distinct from "every app could be a webpage".
1. The post was obviously bullish / optimistic on the technical capabilities. Not in the least dismissive.
2. The economics extrapolation is obvious. See current precedent for paid access for purchased screen-casts of dev work: https://pdoom.org/open_calls/04_crowd_cast.html
Better than text-stripping the internet - this thing will soon be pulling the logits as well.
Funny that I read this as AuthLeft coded (specific to Youtube suppression of Covid truthing). But obviously the alignment is just a function of whatever specific information is labelled "mis".
But in general: agreed, and this is a good list.
I see PDF as a blessed output, but it seems mostly in context of longer form typesetting-heavy workflows (books, papers), rather than design-heavy.
I think there's probably some good juice to squeeze in terms of spacial awareness by doing a benchmark something like
- give 3d modelling task
- render and snapshot from a variety of angles
- feed to third-party vision model for a "what is this" type query
- grade on end-to-end accuracy
Bonus points for asking the vision model something like "how beautiful is this 1-10".
The interviewer asked something like "who is our competition here?", and the friend of friend listed off other places in the mall to get ice cream, candy, deserts, etc.
Wrong answer. The ice cream and chocolate store was in competition with every other store in the mall. Time or money spent at the GAP can't be time or money spent here.
---
Whether or not people are using LLMs for news specifically, any new, large eater of eyeball-time is going to hurt the business landscape for all other eyeball harvesters.
Your private fork doesn't meet the conditions described.
For package X, I should be able to present my npm (homebrew, apt, nuget, etc) credentials with publishing rights for the package.
If package X is of sufficient public interest (user count, nature/sensitivity of user data, downstream distribution, etc), then the public interest + cryptographic credentials should permit access to best-available security auditing.
Yes, we still are trusting trust, that the owner of the package itself is not malicious, but that's not a sharp degradation from status quo.