Still WIP. Feedback welcome.
1,412 karma · joined July 18, 2016
site: https://jm.dev email: j+hn@jm.dev Twitter (eng): @jmq_en
follow:
security: tptacek, moxie, nickpsecurity, strcat, agl, SwellJoe, drewcrawford, schoen, mirimir, secfirstmd, mjg59, userbinator, gorhill, rgovostes, lallysingh, malandrew, mikewest, jedberg, wtarreau, michaelaiello, segmondy, whitequark_, jsnell, salgernon, geofft, ericb
startup: pg, sama, garry, mbesto, coffeemug, davidu, rms, xal, malgorithms, Alex3917, jacquesm, jl, AndrewWarner, emmett, ig1, anateus, lpolovets, ivankirigin, sahillavingia, Joshua, ericflo, immad, rdl, joshfraser, gdb, grellas, gatsby, mayop100, bryanh, josephsunny, ayw, pbiggar, sytse, csallen, joshu, jhuckesteinm, pclark, whockey, sjtgraham, jenthoven, aresant, mayop100, cmdrtaco
code: aphyr, KiranDave, peterwaller, saosebastiao, amirmc, lhorie, judofyr, nikita, huhtenberg, pron, Animats, dmbaggett, chris_wot, aaronbrethorst, grey-area, drewg123, dom96, jashkenas, rich_harris, jordwalke, Homunculiheaded, mark_l_watson, kibwen, Sir_Cmpwn, KenoFischer, ahoyhere, scrollaway, trishume, dbaupp, skrebbel, cryptica, peterhunt, rauchg, TazeTSchnitzel
systems: bcantrill, brendangregg, DannyBee, steveklabnik, fsk, pcwalton, jbk, ajross, yosefk, netguy, minimax, munificent, ColinWright, beat, zwischenzug, derefr, jandrewrogers, shykes, lallysingh, dochtman, SamReidHughes, hnkimb3558, rurban
mods: dang
linux: rwmj, pdkl95
oth: phkahler, davidw, antirez, gwern, patio11, jgrahamc, darksaints, jamwt, nostrademons, plinkplonk, mikekchar, holman, mikeash, edw519, jrockway, noonespecial, staunch, petercooper, jmathai, tzs, jacques_chester, coldtea, peteretep, happy-go-lucky, aaronbrethorst, mtgx
os: vezzy-fnord, rbehrends, vardump, amirmc, pjmlp, rsync, waddlesplash
db: craigkerstiens, teraflop, ifcologne
graphics: pcolton
net: zx2c4, keithwinstein, bsder, walrus01
aws: colmmacc, _msw_, aligouri, illumin8, jcrites, socttlegrand2, openasocket, otterley, mslot
ai: jph00
Still WIP. Feedback welcome.
- I think the "if you use another model" rebuttal is becoming like the No True Scotsman of the LLM world. We can get concrete and discuss a specific model if need be.
- If the use case is "generate this function body for me", I agree that that's a pretty good use case. I've specifically seen problematic behavior for the other ways I'm seeing it OFTEN used, which is "write this feature for me", or trying to one shot too much functionality, where the LLM gets to touch data structures, abstractions, interface boundaries, etc.
- To analogize it to writing: They shouldn't/cannot write the whole book, they shouldn't/cannot write the table of contents, they cannot write a chapter, IMO even a paragraph is too much -- but if you write the first sentence and the last sentence of a paragraph, I think the interpolation can be a pretty reasonable starting point. Bringing it back to code for me means: function bodies are OK. Everything else gets questionable fast IME.
> Unlike prose, however (which really should be handed in a polished form to an LLM to maximize the LLM’s efficacy), LLMs can be quite effective writing code de novo.
Don't the same arguments against using LLMs to write one's prose also apply to code? Was this structure of the code and ideas within the engineers'? Or was it from the LLM? And so on.
Before I'm misunderstood as a LLM minimalist, I want to say that I think they're incredibly good at solving for the blank page syndrome -- just getting a starting point on the page is useful. But I think that the code you actually want to ship is so far from what LLMs write, that I think of it more as a crutch for blank page syndrome than "they're good at writing code de novo".
I'm open to being wrong and want to hear any discussion on the matter. My worry is that this is another one of the "illusion of progress" traps, similar to the one that currently fools people with the prose side of things.
I think AI slop is decidedly different, because it just doesn't have the charm. I don't know if I can yet decompose exactly why that is.
It's probably bad career advice to completely avoid politics (most places aren't doing great work) but it depends on what you're optimizing for.
The problem with everyone getting into the political game is that then we have everyone talking and noone building.
In my experience, the best strategy is to minimize your use of it — call out to binaries or shell scripts and minimize your dependence on any of the GHA world. Makes it easier to test locally too.
I think tons of interpersonal engineering issues boil down to a failure to apply this principle.
I only came down hard on that quote out of context because it felt somewhat standalone and I want to broadcast this “fluency paradox” point a bit louder because I keep running into people who really need to hear it.
I know you know what’s up.
This is the thing with LLMs. When you’re not an expert, the output always looks incredible.
It’s similar to the fluency paradox — if you’re not native in a language, anyone you hear speak it at a higher level than yourself appears to be fluent to you. Even if for example they’re actually just a beginner.
The problem with LLMs is that they’re very good at appearing to speak “a language” at a higher level than you, even if they totally aren’t.
Source: live in Japan, have asked Japanese people around me if they know about this concept (that is popular in USA). Usually hear: へ〜、全然知らない。
1. Queues are actually used a lot, esp. at high scale, and you just don't hear about it.
2. Hardware/compute advances are outpacing user growth (e.g. 1 billion users 10 years ago was a unicorn; 1 billion users today is still a unicorn), but serving (for the sake of argument) 100 million users on a single large box is much more plausible today than 10 years ago. (These numbers are made up; keep the proportions and adjust as you see fit.)
3. Given (2), if you can get away with stuffing your queue into e.g. Redis or a RDBMS, you probably should. It simplifies deployment, architecture, centralizes queries across systems, etc. However, depending on your requirements for scale, reliability, failure (in)dependence, it may not be advisable. I think this is also correlated with a broader understanding that (1) if you can get away with out-of-order task processing, you should, (2) architectural simplicity was underrated in the 2010s industry-wide, (3) YAGNI.
I don't think sharing a prefix/root implies that they're the same thing.
Also, I don't think the suggested "permissions" and "login" terminology would work for all AuthN/Z schemes. For example, when exactly do you "login" when calling an API with a bearer token? Doesn't work for me.
I think the SKIP LOCKED part is really only useful to avoid contention between two workers querying for new work simultaneously.
It seems to me that if the listener dies, notifications in the meantime will be dropped until a listener resubscribes, right? That seems prone to data loss.
In the SKIP LOCKED topic-poller style pattern (for example, query a table for rows with state = 'ready' on some interval and use SKIP LOCKED), you can have arbitrary readers and if they all die, inserts into the table still go through and the backlog can later be processed.
I'm happy to see this, and it should be all goodness, but... the posturing... I don't want to be negative for the sake of being negative, but I don't understand how anyone can write that first paragraph with a straight face and publish it when you're announcing ARM chips for cloud in 2024(?, maybe 2025?).
FWIW, you can configure this in Cargo.toml:
[profile.release] strip = true