And yet, here we are.
425 karma · joined October 18, 2019
And yet, here we are.
Also their documentation was frequently just straight up incorrect (as in the described json schema for a response was violated. keys missing, different field names, etc.).
But it's been over 10 years, has it improved since then? I'm still in my impression of their stack from back then, although they were decently mature by then as well.
They didn't even need torches to subvert democracy. I think a majority of Americans are under the misconception that they live in a civilized democracy, and that they have been for a while. But if you look under the covers (or if you're a minority), or just learn history, then you'll find that it's not.
But that's not the worst of it. It took me some time to realize this, but when I wanted to see why my battery was draining so fast one day (I suspect Youtube has a long standing bug that causes it to go turbo drain mode for some reason, the app is incredibly poorly programmed and leaks memory over time, if you didn't know).
But lo and behold, I couldn't see the per-app battery usage breakdown! Why is that? Turns out Apple disables it if it can't detect a third party battery. Why? I'm sure someone will come and try to defend it, but if you do, you better come with an EE degree and math. To me, it looks like a punishment for not being able to go back to USA within my warranty period to replace my phone's battery which was down to 75% capacity.
This is why I'm not worried about being replaced for now or the forseeable future. For all of the improvements they've made, this part just never seems to change. They could slap another heuristic prompt for the edge case, but eventually it'll revert to the mean again.
I think there is a way to use LLMs to help with programming, but not when I'm not the driver in the seat writing the tests and deciding the architecture. Also I would never ship code written by them as the final product for anything I care about. Since I, like most people, find reading code to be arduous. The more fun thing to do is to force yourself to rewrite it all, treating the LLM's work as a rough draft.
- the addition and standardization (with incomplete coverage) of the solution of adding typing to Python
- how much people are re-discovering the value of performance + typing (e.g. Rust)
then I'm going to take a small leap and extrapolate that the trend will be similar here.
The equivalent of the "one off script in python" will be the LLM, and the long term stable and maintainable solution will be something much more structured and focused like Jev.
The nice thing about this dual setup is that I tend to only want the semantic one if I'm thinking more, so there's naturally a larger time budget for it.
One must always think about the experience they want in UX first, rather than the tools they want to build.
I suppose the scraping you're doing is a huge part of your value proposition, but I would like to gently nudge you in the direction of making the datasets available via p2p (e.g. a torrent) like how Wikipedia distributes its snapshots in the spirit of democratizing access to data that is becoming increasingly walled off. Also, I think another potential benefit that kind of bulk sharing would have is relieving the congestion from those doing the equivalent operation to extract data via the querying interface.
I'm slowly starting to come out of my shell again.
The internet for all of its activity is becoming harder to index because the classic dumb search engines were turned into personalized semantic people pleasers.
The classic model for knowledge discovery still exists, though. You find a blog you like, and use that as a thread of knowledge. You join a community and talk with someone and share resources together. Join a small group chat and post with those people. Your knowledge sharing comes from people you share interests with.
Go to a library, ask your coworkers. If you put in the ground-work, it's possible to do that. For seeds of knowledge outside of your local network, at this point you have to find the indexers that work for you.
Personally, I am in the process of writing the infrastructure for my own search engine, and maybe I will share the tech broadly one day. For now, I'll keep my inventions to my closer personal circles.
In general, I think people should be able to more easily maintain their own offline indices, and share knowledge graphs peer-to-peer rather than relying on a centralized one, in my opinion. This tech has no monetization opportunity, though, so I'm guessing that's why it's relatively under-developed, but luckily for me I have enough money and now enough time to develop and release it (and similar tools).
As someone who has used LLMs heavily at work and at home in various serious experiments, I agree. It still requires heavy babysitting and a lot of its limitations wrt context length are fundamental, not something that’s going to be easy to overcome.
It will be good to see what jj will look like with more funded dev work, but I'm always a little worried about financial incentives mixing with the tools I use for the long term. I guess the saving grace is that I don't really need more upgrades to jj or jjui as it stands. I'm pretty much content with the features and so I could just save this copy of the repo for future reference.
As far as large assets goes, I have my own VCS-ish system which I just integrate with jj, but it would be nice to see a non-git backend handle large assets better as well.
A big reason being that anyone who is using Fable seriously will run out of usage limits very quickly, and so will lean on the "Fable for review + design discussion, Opus 5 agents for implementation" paradigm. But an incredibly annoying UX problem is that the resulting report from the agents that Fable reads isn't surfaced to us in the main dialog, it's only summarized back to us (unless you idle in the agent's window to avoid it closing so you can read what it said directly). As a consequence of this game of telephone, the Fable agent will start using some "terms of art" that it and the agents invented, leaving out literally all context that would be useful in helping me understand what converged/diverged from the implementation attempt. It will often try to ask me for input or say that I have to deliberate on something while also referring to things I've never seen (from the agent result) and without providing any context.
I have to repeatedly prompt it to verbosely explain every time (putting it into the system prompt did little to improve this) and remind it that I can't see what the hell it's talking about.
I'm not sure I'll re-subscribe or even really use AI again because it's honestly more frustrating than it's worth, and so the net emotion I'm left with is frustration and without the satisfaction of learning + building something myself. But at the very least, I thought I'd give someone at the company a tip on what seems to me like a common and obvious UX/UI/workflow failing for using Fable, as some last bit of good will.
- the existence and severity of the financial AI bubble
- the claimed efficacy of AI in terms of its utility vs the actual observed utility
- whether the net good provided by AI outweighs its very heavy costs
I find the section of listing a bunch of selected "predictions" and just saying "Wrong" to elucidate very little. Not that a sentence is sufficient to provide explanation, but Dan stops even doing that bare minimum partway through and just saying "Wrong" full-stop. The reasoning is left up to the reader I guess?
How is it wrong? What was the actual thesis behind it? Is the underlying idea wrong or just the specifics on execution? Was there undetermined factors that mled to the wrong prediction? What can we learn from those factors in order to update our model?
We saw that even though the underlying financials in 2008 were trash and lots of people knew they were trash, things didn't quite collapse in the time frame or way we expected, because an unknown part is how much shenanigans companies can do to extend the runway.
As an example, credit ratings agencies didn't drop ratings to match reality because of customer relation incentives, which is a factor that is not easy to account for and strongly affects the timing of the collapse.
I find the positive reactions to this blog to be confusing. I feel like I learned nothing at all, which makes sense considering under "why write this?" he says "I got four hours of sleep and my brain wasn't good for much of anything and I saw someone posted a screenshot of a reddit post dunking on Ed Zitron's prediction record."