HNHacker News
TopNewBestAskShowJobs

aleksiy123

871 karma · joined November 2, 2021

submissionscomments
aleksiy123··on MicroLLM Lab – Try 7 tiny LLM's in the browser
To be fair this is worse.

It’s verbose with a layer of looking legit play first glance sloppy UI.

I don’t mind vibe code ui but at least some ui efffort would be nice. Not just 1 shot.

aleksiy123··on Using LLMs to trace alchemical knowledge and decode 17th century letters
I also have been doing this every round of models releasing.

Thanks for reminding me to try this new round of models.

I wonder if there are some good tools/apis that can help here.

A lot of information is often login gated

aleksiy123··on Claude's Load-Bearing Seams
Woosh?
aleksiy123··on Show HN: SQLBraid – Write SQL directly in TypeScript without a query-builder DSL
As a sidenote from the slop language, which is awful.

Feel like latest models have really got hung up on “facts” and “evidence” in a really unnatural way.

And especially Claude seems to have a lot of trouble separating it’s internal thoughts, from prompts/instructions, to output.

So you get all of these weird comments and artifacts in the output.

“neighboring versions do not inherit certification. No tag, npm publication, Pages deployment, or release authorization is implied.”

aleksiy123··on AI-generated posters don’t have to be horrible
Tyty I was thinking about something similar
aleksiy123··on LLM Classification Is Feature Engineering
You can also get LLM to optimize rules for a rules engine iteratively against some dataset.

It’s sort of like memoizing or distilling the knowledge. Works really well for certain type of problems.

aleksiy123··on Show HN: Real-time AI news aggregator with daily digest
The font, the colour, and the dot on the tags all scream Claude slop ui.
aleksiy123··on Sony makes bold claim about game ownership
PlayStation store not allowing returns for games that have been downloaded is the most bs policy.

In a world where steam allows instant returns with under 2 hours play time.

I have a few games now on ps where I barely played or found out the controls were terrible and have no way to return.

aleksiy123··on It's OK to hardcode feature flags (2025)
You can do much more with feature flags like ab testing/experiments, integrations with analytics.

There can be whole UIs and tooling and infrastructure to manage around them and that’s what the sass offer

aleksiy123··on Coordination Headwind: How Organizations Are Like Slime Molds
Managing complexity while scaling is a concept just keeps reappearing.

Funnily enough it’s extremely relevant to agents/subagents.

aleksiy123··on Show HN: Every push-up becomes an attack in a camera-counted RPG game
Huh this is cool. Was thinking about earlier that it must be a lot easier to make motion capture games now.

Are there any new dance style games?

As an aside wonder if there is something to adding game/gambling mechanics to exercise.

aleksiy123··on Quick impressions: A week of using Codex more than Claude
“Extreme defensiveness” is a good characterization.

I wonder if it’s an artifact of OpenAI’s values or rl training approach.

Also, it prob does make it perform better just not more efficient.

Great for the OpenAI employee working on security scanning who doesn’t have to pay for their tokens.

Not so much for the dev building their web app who is trying maximize their subscription.

Like hiring an aerospace engineer to build you a shed.

aleksiy123··on A week of using Codex more than Claude
As a follow up.

I feel like codex/sol is better at well scoped hard technical problem.

Where it can sort of run this brute force analytical loop.

Like doing performance optimization or other search type problems. I think the math proofs are good examples of this.

aleksiy123··on A week of using Codex more than Claude
Agree with most of these.

One thing I don’t love about codex/sol is I find it tends to overengineer and be overly cautious.

I was using it to do create some scraping + data processing.

It went kind of crazy on the provenance, need at least 3 sources of consensus before promoting facts type bullshit.

defined a bunch of enums and gates.

I just wanted scrape some site data and put it into a SQLite dB. Like chill codex.

I feel like Claude is better at that.

aleksiy123··on Claudette: Make Claude stop talking like a BuzzFeed article
Hooks is the way.

Intermittent nudges

aleksiy123··on Opus 5.0 drives incoherence into the stratosphere
I use a hook. Checks for banned words and patterns.

Injects a reminder to use ASD-STE100 Simplified Technical English which I picked up from a suggestion in another thread.

Honestly it’s working pretty well. Except for I need to check how often it’s actually firing.

https://en.wikipedia.org/wiki/Simplified_Technical_English

I feel like hooks aren’t utilized enough. Really nice for being the sort of auto steering as long as you can encode some pattern to detect the bad behaviour.

Prompts and skills just don’t cut it.

aleksiy123··on Composable Tests
¯\_(ツ)_/¯

You can believe whatever you want to believe. Or maybe you live in a magical world where all your tests are pure and side effects nor IO exist.

In the other thread you think running tests in parallel doesn't count because it has documented gotchas...

aleksiy123··on Composable Tests
You’ve never ever had to be mindful about isolation of tests? Also, Unit tests aren’t the only tests that exist.

You say assured? You think it’s impossible to write 2 tests that interact each other?

All you gotta do is google “test isolation”. But if you want to pretend it’s not a thing the I have no idea what to tell you?

Here’s one example since junit came up

https://martinfowler.com/bliki/JunitNewInstance.html

aleksiy123··on Composable Tests
This
aleksiy123··on Composable Tests
All I mean that there is no way for a framework to prevent your tests from interfering with each other, or to solve isolation for you.

It’s the implementation of each test that is responsible for its isolation. Every time you write or make changes to your tests. You have to think, can they mess with each other.

aleksiy123··on Composable Tests
I guess it all depends, isolation isn’t binary.

you can share some things and be isolated across other dimensions.

You could share a Postgres dB connection but just isolate the data logically.

You can isolate or share across individual tests, across suites, across envs, across runs, across time.

More isolation is generally better unless the complexity or cost is too high.

But I guess I was just confused as to how a test requires isolation rather than have it.

aleksiy123··on Composable Tests
I don’t think that’s right?

The isolation comes from the test implementation not the framework. There isn’t any framework out there that can guarantee/give you isolation.

If I create a new in mem dB in the test there’s nothing stopping me from running it in parallel?

Nothing about that “requires” isolation. It is isolation.

aleksiy123··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Is this only for realtime tts use cases?

Wondering if you also support some non realtime models.

aleksiy123··on How Claude's text watermarking works
Curious if you can prompt Claude to sue some scrambling scheme and then unscramble to defeat this.

E.g. prompt Claude to write all sentence in reverse, or swap every 2 words etc. Then use a script to put reorder in the right ordering?

aleksiy123··on Don't classify, hallucinate
Jeopardy clustering

Pretty cool technique honestly. You could do it the other way as well right?

If you had a list of categories you have the model to generate a sample query and then do embedding on that?

aleksiy123··on Maximizing the value of your Claude Code sessions
Is it possible to have some kind of script to keep your cache warm, or auto compact or something.

I sometimes just leave some goals or something running before I go to bed or out and I don’t want to pay the cache text when I come back.

aleksiy123··on How AI text watermarking works
but is it even possibly to selectively apply it? I guess some kind of tag or indicator token that its code or not code?
aleksiy123··on How AI text watermarking works
In curious how does this work with tool calls or CLI scripts etc?

like if you have a long cli command or something will it still try to watermark it ?

Is there some way you can know which tokens are required to be syntactically correct vs not?

aleksiy123··on Process as a Proxy for Motivation
I can make a decision to go to gym tomorrow.

My process/playbook for determining if I go to the gym is I go if I didn’t go the day before.

This is both a premade decision and a process that requires discipline to execute on.

So why isn’t a decision made in advance an example of a simple process?

aleksiy123··on Process as a Proxy for Motivation
This is a good one. Sometimes it doesn’t even matter if the made decision is suboptimal. It’s better than no decision.

Opinionated linters is an example.

Sometimes you just gotta turn your brain off and move forward.

Page 1 of 15Next →