HNHacker News
TopNewBestAskShowJobs

k__

15,843 karma · joined March 11, 2013

Author that tries to smash the gates that keep you from building great things.

Website: https://kay.is

Blog: https://fllstck.dev

GitHub: https://github.com/kay-is

submissionscomments
k__··on DeepSeek peak/off-peak pricing update
Yeah, None of them offers the same cache hit prices and my cache hits are >95%.
k__··on DeepSeek peak/off-peak pricing update
I didn't get the impression that anyone competed with the old prices before.
k__··on DeepSeek V4 Pro 0813
I get like 80.
k__··on DeepSeek V4 Pro 0813
I wouldn't exactly call it snappy, but faster than Pro, yes.
k__··on DeepSeek V4 Pro 0813
I tried the previous Pro model and in the end it was 50% more expensive than the previous Flash.

Wasn't worth it.

k__··on DeepSeek V4 Pro 0813
Around 5 percentage points better. (E.g., 87% instead of 82%)
k__··on Go is an ideal language for AI-assisted software engineering
Had the same impression about TypeScript and Rust.

Not as fun to write as Python and Nim, but I don't have to write it.

k__··on Zero-Mem: Zero-Token Memory Operations for LLM Agents
Care to elaborate?
k__··on Zero-Mem: Zero-Token Memory Operations for LLM Agents
"A local Qwen3.5 retrieval model attends over the indexed memory"

Hrm.

k__··on DeepSeek V4 Flash 2.98x faster, lossless
Seems like this only helps with parallel workloads.
k__··on Ask HN: Who wants to be hired? (August 2026)

  Location: Germany 
  Remote: Yes
  Willing to relocate: No
  Roles: Technical Writer, Software Engineer 
  Homepage: https://kay.is
  LinkedIn: https://www.linkedin.com/in/kay-plößer
k__··on LLMs reward expertise
Prompt an image or video generator without knowledge in photography or art skills and your results will look sloppy.
k__··on Can an AI Model Obey a Strict Style Guide?
I did a bit of research on that in my last job and got the impression that encoder models might help.

They are well suited to check input for rule violations and are much cheaper to train and run than decoder models.

The downside is, they can't fix issues by themselves.

k__··on Qwen3.8-Max: A New Bar for Coding and Cowork
Yeah, I'd assume it's possible to extract all languages as steering vectors from a model and then substract the ones you don't need from its weights.

However, that would just change the weights values and not their dimensions.

k__··on More German than many Germans
I think, the main issue with German rules is that we haven't embraced digital technology 100%.

All these rules would be way less cumbersome if they didn't come with a bunch of literal paperwork.

k__··on Generative AI floods and dilutes the market for books
The ratio of good books to slop (AI or not) was far to bad for humans to reasonably filter 20 years ago.
k__··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
Haha, and I almost felt bad after seeing this chart yesterday.
k__··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
Half OT:

Why do the cache hit rates seem to vary so much between harnesses?

I use pi, which is very minimalist, and I get a hit rate of ~99%. Paying like $1 a day for Flash. Yet, the hit rate mentioned on OpenRouter is only ~79%.

k__··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
On OpenRouter it's 93 TPS.
k__··on DeepSeek-V4-Flash Update
I'm using pi and my caching is ~99%.
k__··on DeepSeek-V4-Flash Update
Yeah, it needs quite some hand holding.

I didn't do much agent coding and had a mix experience.

1. It would build something that was in the spirit of what I wanted, but unusable in practice.

2. It would build something quite useful, but only the public APIs were nice, the deeper code layers would get more and more convoluted.

3. It would built what I wanted and it would have okay-ish code.

However, for 3. I also had to add a custom AGENTS.md, many more code example, extra repos as subtrees, and review any code that had new concepts.

Much more work, but still much less than typing it all by hand.

k__··on DeepSeek-V4-Flash Update
I'd take more throughput while everything else stays the same.
k__··on Kimi K3-256k
While I tend to clear my session after every task, I feel less stressed when my context is as 10% than when it's at 30%
k__··on Kimi K3-256k
They remove potentially irrelevant details.
k__··on Kimi K3-256k
As I understand it, they would have to train a whole knew model to hard cap it's context to different lengths. That would be cheaper to train and had cheaper inf, but still a huge investment.

So I'd guess it's API level.

k__··on Kimi-K3 Releases on HuggingFace 7/27
Right
k__··on The front end framework for correctness: built on Effect, architected like Elm
Yeah, I think the Effect team tried building a compiler for once (TS++ or something) but they abandoned it, as it was too much work.
k__··on Writing by hand is good for your brain
Do crosswords count?
k__··on Qwen 3.8
My 2 weeks with DeepSeek V4:

Pro is ~50% more expensive than Flash.

Both need babysitting.

Plan, split in small tasks, give it docs, types, tests, linter, best practice examples, etc.

Always start a new session when starting a task.

Do regular manual sanity checks, and tell it to find issues in the codebase.

I pay like $1,50 per day for Pro.

k__··on The Kimi K3 Moment
Thanks!

What is the parento frontier?

← PreviousPage 3 of 34Next →