HNHacker News
TopNewBestAskShowJobs

markab21

180 karma · joined March 31, 2014

submissionscomments
markab21··on Alternatives to MinIO for single-node local S3
Yup, same. We use RustFS in a production system with no issues. We switched over after MinIO license changes and after evaluating a few different stacks.

We have heavy concurrent usage, and we haven't had a single issue yet.

markab21··on DeepSeek-V4-Flash Update
We use it at a moderate scale, self-hosted on B300 hardware. It's great :D

QA analysis of voice transcriptions. Napkin math: we operate at 2-5% of the cost of running on Equiv Frontier, though this changes near-weekly because pricing is so volatile.

It took us about a month to get the inference configured to achieve these numbers. But if you can get your hands on a pair of B300 GPUs and the context works, it's untouchable for price/performance.

(B200 would work, but you don't have the B300's memory, which lets you run it on 2xGPU instead of 4xGPU... with Dspark, it's like magic)

On a side note, for tasks that don't require the intelligence of DS v4 flash, we're using Nemotron-3-super with incredible success. I'm shocked we're not seeing more adoption of this model, given how easy it is to fine-tune and how blisteringly fast the nvfp4 version is. (A single B200 GPU can produce an insane amount of throughput with Nemotron 3 Super.)

markab21··on Claude Opus 5
Yes - THIS! I can't even believe how exhausting it is to read. I'm not sure why or what changed in Fable. Did they do this writing-style output to give it more token compression during/for training or to prefer output for less money?

I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.

I'm guessing it's just not enough time doing RL on human feedback.

Check out the anouncement of Inkling (https://thinkingmachines.ai/news/introducing-inkling/)... the section in the middle

"Early in RL verbose, grammatical" (if you search) :

We need to understand the operator. The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π]. The internal coordinate is periodic. The background is a warped product: metric g_{MN} where M,N = 0..4. The internal direction has metric e^{2A(x)} dx²? Wait, the ds² is e^{2A} (ds²_4d + dx²). So the internal metric is e^{2A(x)} dx². Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².

vs. Post RL

We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x. For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}? Wait internal metric is e^{2A} dx²? Actually ds² = e^{2A} [ds_4² + dx²]. So internal metric is e^{2A} dx²; warp factor same for 4d and internal? Yes. We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization. …

I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.

markab21··on Apertus – Open Foundation Model for Sovereign AI
I'm mildly surprised that more people aren't using Nemo models for this reason. We've moved most of our processing to a combination of Nemo Ultra and Super, with some support for multi-model-specific tasks on Omni. The setup is working REALLY well for us, and I'm comfortable with the more measured pace of improvements. We work with many long-context problems, and the ecosystem is great.

There were a number of use cases where we needed to use Gemini (audio modality), and Ultra has been a VERY cost-effective alternative once we got through the nuances.

markab21··on Google releases Gemma 4 open models
I'll pipe in here as someone working on an agentic harness project using mastra as the harness.

Nemotron3-super is, without question, my favorite model now for my agentic use cases. The closest model I would compare it to, in vibe and feel, is the Qwen family but this thing has an ability to hold attention through complicated (often noisy) agentic environments and I'm sometimes finding myself checking that i'm not on a frontier model.

I now just rent a Dual B6000 on a full-time basis for myself for all my stuff; this is the backbone of my "base" agentic workload, and I only step up to stronger models in rare situations in my pipelines.

The biggest thing with this model, I've found, is just making sure my environment is set up correctly; the temps and templates need to be exactly right. I've had hit-or-miss with OpenRouter. But running this model on a B6000 from Vast with a native NVFP4 model weight from Nvidia, it's really good. (2500 peak tokens/sec on that setup) batching. about 100/s 1-request, 250k context. :)

I can run on a single B6000 up to about 120k context reliably but really this thing SCREAMS on a dual-b6000. (I'm close to just ordering a couple for myself it's working so well).

Good luck .. (Sometimes I feel like I'm the crazy guy in the woods loving this model so much, I'm not sure why more people aren't jumping on it..)

markab21··on If DSPy is so great, why isn't anyone using it?
I think the entire premise that the prompting is the surface area for optimizing the application is fundamentally the wrong framing, in the same way that in 1998 better cpam will save CGI. It's solving the wrong problems now, and the limitations in context and model intelligence require a tool like Dspy.

The only thing I'd grab dspy for at this point is to automate the edges of the agentic pipeline that could be improved with RL patterns. But if that is true, you're really shorting yourself by giving your domain DSPY. You should be building your own RL learning loops.

My experience: If you find yourself reaching for a tool like Dspy, you might be sitting on a scenario where reinforcement learning approaches would help even further up the stack than your prompts, and you're probably missing where the real optimization win is. (Think bigger)

markab21··on Gemini 3.1 Pro
You just articulated why I struggle to personally connect with Gemini. It feels so unrelatable and exhausting to read its output. I prefer to read Opus/Deepseek/GLM over Gemini, Qwen and the open source GPT models. Maybe it is RLHF that is creating my distaste from using it. (I pay for Gemini; I should be using it more... but the outputs just bug me and feel more work to get actionable insight.)
markab21··on Experts Have World Models. LLMs Have Word Models
And I think you basically just described the OpenAI approach to building models and serving them.
markab21··on Orchestrate teams of Claude Code sessions
Shaking fist at clouds!!
markab21··on Qwen3-Coder-Next
It's getting a lot easier to do this using sub-agents with tools in Claude. I have a fleet of Mastra agents (TypeScript). I use those agents inside my project as CLI tools to do repetitive tasks that gobble tokens such as scanning code, web search, library search, and even SourceGraph traversal.

Overall, it's allowed me to maintain more consistent workflows as I'm less dependent on Opus. Now that Mastra has introduced the concept of Workspaces, which allow for more agentic development, this approach has become even more powerful.

markab21··on TimeCapsuleLLM: LLM trained only on data from 1800-1875
Basically looking for emergent behavior.
markab21··on Show HN: Mysti – Claude, Codex, and Gemini debate your code, then synthesize
I love where you're going with this. In my experience it's not about a different persona, it's about constantly considering context that triggers, different activations enhance a different outcome. You can achieve the same thing, of course by switching to an agent with a separate persona, but you can also get it simply by injecting new context, or forcing the agent to consider something new. I feel like this concept gets cargo-culted a little bit.

I personally have moved to a pattern where i use mastra-agents in my project to achieve this. I've slowly shifted the bulk of the code research and web research to my internal tools (built with small typescript agents).. I can now really easily bounce between different tools such as claude, codex, opencode and my coding tools are spending more time orchestrating work than doing the work themselves.

markab21··on Apps SDK
The skepticism is understandable given the trajectory of GPTs and custom instructions, but there's a meaningful technical difference here: the Apps SDK is built on the Model Context Protocol (MCP), which is an open specification rather than a proprietary format.

MCP standardizes how LLM clients connect to external tools—defining wire formats, authentication flows, and metadata schemas. This means apps you build aren't inherently ChatGPT-specific; they're MCP servers that could work with any MCP-compatible client. The protocol is transport-agnostic and self-describing, with official Python and TypeScript SDKs already available.

That said, the "build our platform" criticism isn't entirely off base. While the protocol is open, practical adoption still depends heavily on ChatGPT's distribution and whether other LLM providers actually implement MCP clients. The real test will be whether this becomes a genuine cross-platform standard or just another way to contribute to OpenAI's ecosystem.

The technical primitives (tool discovery, structured content return, embedded UI resources) are solid and address real integration problems. Whether it succeeds likely depends more on ecosystem dynamics than technical merit.

markab21··on Mistral NeMo
It looks like it was built jointly with nvidia: https://huggingface.co/nvidia/Mistral-NeMo-12B-Instruct
markab21··on How Mandelbrot set images are affected by floating point precision
I bet you have been waiting years to pull that one out of your pocket.

Well played sir! Nice shot man! :D

markab21··on Memory and new controls for ChatGPT
I've found myself more and more using local models rather than ChatGPT; it was pretty trivial to set up Ollama+Ollama-WebUI, which is shockingly good.

I'm so tired of arguing with ChatGPT (or what was Bard) to even get simple things done. SOLAR-10B or Mistral works just fine for my use cases, and I've wired up a direct connection to Fireworks/OpenRouter/Together for the occasion I need anything more than what will run on my local hardware. (mixtral MOE, 70B code/chat models)

markab21··on Mistral CEO confirms 'leak' of new open source AI model nearing GPT4 performance
For Llama-based progress - Reddit - /r/LocalLlama has been my top source of info, although it's been getting a little more noisy lately.

I also hang out on a few Discord servers: - Nous Research - TogetherAI / Fireworks / Openrouter - LangChain - TheBloke AI - Mistral AI

These, along with a couple of newsletters, basically keep a pulse on things.

markab21··on Gulf Stream weakening now 99% certain, and ramifications will be global
[flagged]
markab21··on Obsidian 1.4.10 Desktop (Public)
Yeah, slow news day.
markab21··on Debian celebrates 30 years
Linux distributions, including Debian, offer a variety of desktop environments, each with its own design philosophy and user experience. If one environment doesn't suit your preferences, others might be more to your liking. It's worth exploring different desktop environments to find one that aligns with your expectations.
markab21··on Debian celebrates 30 years
Debian and Ubuntu have similarities, but keep in mind - Ubuntu is derived from Debian, not the other way around. However, they differ in areas like release cycles, package management, and default configurations.
markab21··on Ask HN: Should HN ban ChatGPT/generated responses?
I used it as a consultant on a development project to help me organize some of the milestones and design goals in some documentation.

It wasn't that I didn't know the stuff, I do, but more helpful with quickly organizing and presenting information in a clean and well-written way. I did have to go through and re-write parts of it specific to our domain.. but it saved me many hours of work doing tedious organization of data.

I also tested it with helping create some SOP's for a new position in our very small company, even breaking down the expected tasks into daily schedules.

It's not that it's perfect, but it generates a bit of a boiler-plate starting point for me which then I can work with from there.

markab21··on Google and Facebook execs allegedly approved dividing ad market among themselves
I'd be surprised if they don't have async mechanisms.
markab21··on FBI's ability to legally access secure messaging app content and metadata [pdf]
Assume anything sent over a cellular network carrier via normal SMS can not only be retrieved, but intercepted.
markab21··on Oracle vs. PostgreSQL: First Glance
Why anyone would use Oracle for anything other than supporting legacy systems is beyond me.
markab21··on BMW is accelerating electrification plans
The claim is dead-right.

I own a Tesla Model 3, my wife drives a BMW i3, my daughter has a leaf.

The ONLY car we can effectively travel outside of the greater Tampa area without major headache is the Tesla.

The ONLY car that I would try to drive to New York from Florida in is the Tesla. (Yes, we've done it.. but would only try it in the Tesla)

markab21··on BMW is accelerating electrification plans
Care to site/link these two studies?
markab21··on The age of electric flight is finally upon us
As someone who regularly pilots a piston-single, I'm looking forward to what hybrid or full electric can do for us little guys. The takeoff phase of flight is a high-risk situation in events such as fuel contamination which could prove fatal.

Having a small buffer, even a few minutes of battery power to rely on while trying to get back to the field to land the "impossible turn" [1] would make me feel a lot better and could be the difference between life and death.

There are various groups (Pipstrel[2], Diamond[3]) that I'm aware of that are working on electric GA aircraft. For young pilots that are looking to train, the cost of jumping in a Cessna 172, the gold standard in GA trainers will cost at best $120-200/hr. Electric costs should be 1/5 (or better) of that in reality due to the absolute bargain of replacing a TBO electric engine, scheduled maintenance and relatively low level of complexity.

[1] https://www.aopa.org/training-and-safety/air-safety-institut...

[2] https://www.pipistrel-usa.com/electric-propulsion/

[3] https://www.flyingmag.com/diamond-da40-hybrid-electric-proto...

markab21··on Tesla Model 3 becomes best-selling car in Switzerland
> I'd like to get one but it doesn't fit my use case (namely long road trips in the summer).

I own a P3D, facing the same issues of long road trips 1-2x a year I concluded that I'll just spend the few hundred dollars and rent a car for the deep edge cases of my driving and the other 99% of my time I'll enjoy driving my Tesla.

99% of my normal day to day driving I just charge at home at night, I've used the supercharger network on a long road trip and it was surprisingly little-hassle.

markab21··on Swansea Uni study: African wild dogs 'sneeze to vote' (2017)
I've noticed my dog (black lab) doing similar behavior consistently in the mornings, also after she had her breakfast. She'll walk away from the kitchen and her dog bowl with her tail wagging and doing some kind of sneezing sniffing-sound, also when we come home from being out of the house she'll do it too.

It's cute as hell.

Page 1 of 3Next →