4,410 karma · joined March 26, 2016
So given the choice between Cognito (easy right? It's a managed service!) and Keycloak (a large, complex security Swiss Army knife) should have been a no-brainer in favour of Cognito. But we absolutely hated it. I think it's one of the worst AWS services I've used.
Keycloak is only hard because security is hard. The terraform provider is amazing, worked remarkably well. We never had a single hitch, and I could spin up new tenants in seconds, or deploy Keycloak to other regions in a few minutes.
Personally, I just hate the idea of running that type of thing in production unless it's my full time job (and it wasn't) and I have other team members who also knowit backwards. But for a sufficiently well staffed org that feels like moving this expertise inhouse and dumping a paid auth service, Keycloak really is amazingly good.
Now that I've moved on to another job, my time spent in the belly of Keycloak wasn't wasted either. I can help our security team figure out auth issues blindfolded (ie: without access to their Okta admin interface, and in fact having never seen one)
I was finally able to stop managing and rotating TLS key pairs, replicating private keys globally to my apps.
So: who really reads those little boring announcements, and connects the dots? Seemingly, not even AWS themselves!
Companies are no different. During company growth, it's all about hiring, and giving tasks and some responsibilities to subordinates, setting up the right channels for communication etc. But at what point do company leadership actually say "You know what,that department would work better without direct instructions from us. Let's facilitate communication, act as conflict adjudicators and make company-wide 'North Star' decisions"
I think that's what OKRs are supposed to do, but I've never seen that work well (for very long anyway). It always gets up-ended somehow. Start of quarter optimism, over-promising, not spotting that someone else's stretch goal is actually your blocker. Then mid-quarter you're dealing with production emergencies, database code red and running out of time. At the end of the quarter, two of your team go off on PTO while you desperately try not to "fail your quarter", working late, and delivering shoddy work.
Does true bottom up exist at large companies? Hmmm. Some. We've all heard the legends of IBM's "build the first PC!" skunkworks. And we know about research-only departments at Xerox that led to so many modern computing inventions. But these were brief projects.
I find these days that people have the desire to get out of the way and let us work, but the tiling is in the way. I waste so much fixing time in Jira, slack, doing performance reviews, interviewing...
However coming up with a prompt that didn't turn out total garbage was impossible. After wasting over an hour and I ended up getting Qwen side by side with Llama 3.2 3B, just to see if I was being stupid. Nope, it just looks like Llama is orders of magnitude better at this specific task for some reason).
If you think I'm doing it wrong, you're probably right, I don't know a ton about local LLMs. But I hand selected 50 songs, set up ollama with both LLMs, and for each iteration on the prompt text, ran both LLMs 10x times per song. Side-by-side comparisons showed that Qwen 3 4B was so bad that I actually downloaded Qwen again, thinking there must have been some mistake and I accidentally grabbed an old 1B model.
Claude has got better at "just fucking doing it" by asking if it's ok to go read the latest github issues and pull the README, which means that people will likely get lazier and lazier.
There's no consistency over multiple runs of the same prompt on the same model.
Also, for the purpose of music lyric analysis Qwen 4B is laughably bad. Like it's going out of its way to be extremely wrong, misunderstand the prompt. When it does correctly understand what I asked for it ALWAYS tells me that the mood is Angry. Sometimesiit just gives back all the lyrics. Sometimes it claims that it doesn't have the list of moods or the lyrics and tells me I should look them up on the Internet first.
All with the same prompt every time.
Are models being overtuned for coding tasks?
Also...
Place one rounded teaspoon of tea per cup into the pot.
My rule is usually one per person and "one for the pot".Yes this is real. And it's hell. And now all of those Jenga players are AI powered.
I work at a company where the biggest problems are not 'writing code', they are:
- Organising teams
- Designing the system
- Prioritisation of work
The fuckups that we make on a daily bases are not 'code errors' they are failures in THOSE three things. I'll go into detail if anyone cares.
I find that nobody really knows how to do this. Machine learning can detect some song attributes well (bpm, ez right?) but it's inconsistent with some things (eg mood, spotify valence)
I prefer to only add metadata that I can rely on: track credits & instruments (when available), lyrics, bpm / "energy" and genre. At least that's what I've got for now. I'm not adding anything unreliable.
So far I'm able to pick a genre, artist or even better, song and it gives me a list of tracks that are similar. I can alter the weights of "era", "instruments", "genre".
So far i haven't run old school NLP on the lyrics but that's the next step. It's likely to be far more informative than "valence"
Anyway, not public, still very alpha but I like it and find it useful.
I did set up tailscale, way back. After using it a few times to test, it failed me when I really needed it (I was out of the country and it failed - can't remember exactly what went wrong but it wwas 100% 'in my tailscale account'). I immediately dropped it and went back to OpenVPN (shit but reliable) before building my current setup.
So these days I almost exclusively use Go wherever bash and friends aren't enough. Compile to Linux x86 & ARM and Darwin ARM, or just "go run" if it's simple enough.
This means that I can always use public DNS servers like 1.1.1.1, 8.8.8.8, nextDNS etc
This is not "done right" by any stretch but it's extremely low effort to set up and has never once failed me, unlike countless complex meshy things.