2,808 karma · joined March 18, 2016
Lived in Portugal, Ireland, UK, and now in the US / New Jersey.
Always happy to chat. You can reach me at hn@pocketarc.com.
Although:
"(idiomatic, etc.)" I don't think that should be part of the aim (or it should be under-weighted), because... you can just provide guidance on how to do things more idiomatically, rather than depend on that knowledge already being encoded in the model. I'd be more curious about verifying that it can work through gnarly bugs / features in an Elixir codebase. After all, what use is a model that by default does everything idiomatically if it can't figure out some small concurrency bug.
It’s not hard to imagine what 5-10 years of pressure to increase RAM will do to specs, on top of the normal tech improvements.
That’s worth bearing in mind when thinking about local models.
Plus, local models keep getting better and better; 2 years ago what you could get out of those 48GB of RAM was embarrassing compared to what’s doable today.
We’re getting there. Just takes time.
Nowadays, with the focus on agentic use and coding, it seems models have all been RLHF’d to death, it’s so incredibly hard to have them write in a different voice than their default. I put together a skill to review its writing and have it edit its own output (e.g. code comments), which does make a difference, but isn’t perfect.
What, if anything, do people do for writing? That feels like a neglected side of LLMs. They’ll make 100 Bash calls referencing ancient commands without batting an eye but heaven forbid they use something other than “load-bearing” while talking. For something trained on “all the human knowledge” it’s incredible how limited their default vocabulary seems to be.
That's not what's happening here, and it's worth remembering: A caveman from 200K years ago would have been just as intelligent as any of us here today, despite not having language or technology, or any knowledge.
In Carolyn Porco's words: "These beings, with soaring imagination, eventually flung themselves and their machines into interplanetary space."
When you think of it that way, it should be obvious that LLMs are not AGI. And that's OK! They're a remarkable piece of technology anyway! It turns out that LLMs are actually good enough for a lot of use cases that would otherwise have required human intelligence.
And I echo ArekDymalski's sentiment that it's good to have benchmarks to structure the discussions around the "intelligence level" of LLMs. That _is_ useful, and the more progress we make, the better. But we're not on the way to AGI.
"source available"[1] is a different thing, and you're right that this project is "source available".
[0]: https://opensource.org/osd
[1]: https://en.wikipedia.org/wiki/Source-available_software
Having a promised date lets you keep the opportunity going and in some cases can even let you sign them there and then - you sign them under the condition that feature X will be in the app by date Y. That's waaaay better for business, even if it's tougher for engineers.
Failures are fed back to the LLM so it can regenerate taking that feedback into account. People are much happier with it than I could have imagined, though it's definitely not cheap (but the cost difference is very OK for the tradeoff).
You're looking for your first DevOps person, so you want someone who has experience doing DevOps. They'll tell you about all the fancy frameworks and tooling they've used to do Serious Business™, and you'll be impressed and hire them. They'll then proceed to do exactly that for your company, and you'll feel good because you feel it sets you up for the future.
Nobody's against it. So you end up in that situation, which even a basic home desktop would be more than capable of handling.
I really hope that I never end up in a situation where someone tells me "well the conversion rate would be much higher if you just stopped fighting it and put up the damn banner".
Most of the changes are completely reasonable - a lot are internal cleanup that would require no code changes on the user side, dropping older browsers, etc.
But the fact that there are breaking API changes is the most surprising thing to me. Projects that still use jQuery are going to be mostly legacy projects (I myself have several lying around). Breaking changes means more of an upgrade hassle on something that's already not worth much of an upgrade hassle to begin with. Removing things like `jQuery.isArray` serve only to make the upgrade path harder - the internal jQuery function code could literally just be `Array.isArray`, but at least then you wouldn't be breaking jQuery users' existing code.
At some point in the life of projects like these, I feel like they should accept their place in history and stop themselves breaking compatibility with any of the countless thousands (millions!) of their users' projects. Just be a good clean library that one can keep using without having to think about it forever and ever.
You could take those, make the tools better, and repeat the experience, and I'd love to see how much better the run would go.
I keep thinking about that when it comes to things like this - the Pokemon thing as well. The quality of the tooling around the AI is only going to become more and more impactful as time goes on. The more you can deterministically figure out on behalf of the AI to provide it with accurate ways of seeing and doing things, the better.
Ditto for humans, of course, that's the great thing about optimizing for AI. It's really just "if a human was using this, what would they need"? Think about it: The whole thing with the paths not being properly connected, a human would have to sit down and really think about it, draw/sketch the layout to visualize and understand what coordinates to do things in. And if you couldn't do that, you too would probably struggle for a while. But if the tool provided you with enough context to understand that a path wasn't connected properly and why, you'd be fine.
Fully agreed. This might be one of those areas where people not knowing it's possible can lead them to far worse solutions. Generated columns in databases, computing things based on row data, and things like materialized views, are so, so useful.
> on large screens
> in dark mode
> when hovering
> bg should be red-500
The above is an unrealistic example, but, you can't achieve that with the style attribute. You'd have to go into your stylesheet and put this inside the @media query for the right screen size + dark mode, with :hover, etc.
And you'd still need to have a class on the element (how else are you going to target that element)?
And then 6 months later you get a ticket to change it to blue instead. You open up the HTML, you look at the class of the element to remind yourself of what it's called, then you go to the CSS looking for that class, and then you make the change. Did you affect any other elements? Was that class unique? Do you know or do you just hope? Eh just add a new rule at the bottom of the file with !important and raise a PR, you've got other tickets to work on. I've seen that done countless times working in teams over the past 20 years - over a long enough timeline stylesheets all tend to end up a mess of overrides like that.
If you just work on your own, that's certainly a different discussion. I'd say Tailwind is still useful, but Tailwind's value really goes up the bigger the team you're working with. You do away with all those !important's and all those random class names and class naming style guide discussions.
I used to look at Tailwind and think "ew we were supposed to do CSS separate from HTML why are we just throwing styles back in the HTML". Then I was forced to use it, and I understood why people liked it. It just makes everything easier.
There's always going to be a prevailing zeitgeist, and that's always going to dominate people's focus.
But by and large, HN is the most intelligent, thoughtful community I've encountered, with people going out of their way to be kind and thorough in their replies, discussing things maturely, disincentivizing unhelpful and low-effort behavior. Not to mention the dizzying wealth of knowledge and experience the people here have.
Obviously it's not 100%. There are plenty of harsh people and attitudes. But there's a lot of good discussion happening here, and that's what I try to focus on. I can only hope one day to be as interesting/inspiring a member of the community as some of the people I've seen on here.
They are talking about vitamin supplements, not literal vitamins that you need in order to live. Vitamin supplements do not survive in a budget reduction spreadsheet - they're easy to let go of for a while. On the other hand, if you're in pain, you need painkillers, and you're not going to be thinking about your budget, you're just going to go get some to get rid of the pain, even if it's just a temporary fix (and even better for the business if it's just a temporary fix - recurring revenue!).
That's the whole thing, the whole "solve a real problem" thing they keep talking about for startups.
Like, it's fine for you to use AI, just like one would use Google. But you wouldn't paste "here are 10 results I got from Google". So don't paste whatever AI said without doing the work, yourself, of reviewing and making sense of it. Don't push that work onto others.
LLMs are not encyclopedias.
Give an LLM the context you want to explore, and it will do a fantastic job of telling you all about it. Give an LLM access to web search, and it will find things for you and tell you what you want to know. Ask it "what's happening in my town this week?", and it will answer that with the tools it is given. Not out of its oracle mind, but out of web search + natural language processing.
Stop expecting LLMs to -know- things. Treating LLMs like all-knowing oracles is exactly the thing that's setting apart those who are finding huge productivity gains with them from those who can't get anything productive out of them.
Yeah, this is definitely one of the saddest things about modern UI fashion. We have the highest-resolution, highest-DPI, cleanest-looking extra-bright, extra-deep-black HDR OLED screens, and... we've got flatter UI than ever, UI that would've looked dull even on a 90s CRT.
We now praise dark mode as some big achievement, but... we -had- dark mode, before. The Mac Themes Garden has countless "dark mode" themes.
I absolutely agree with you. I've been very very keen on CSP for a long time, it feels SO good to know that that vector for exploiting vulnerabilities is plugged.
One thing that's very noticeable: It seems to block/break -a lot- of web extensions. Basically every error I see in Sentry is of the form of "X.js blocked" or "random script eval blocked", stuff that's all extension-related.