HNHacker News
TopNewBestAskShowJobs

redmalang

43 karma · joined June 27, 2017

say hello: hikes-thane2p@icloud.com
submissionscomments
redmalang··on Pi 1.0
For the past several (months now!) I have been slowly working to open source a proxy we built for pi internally. Its been super helpful for us, helping us centralise session logging, hook and model insights as people work on stuff. There is a bunch of other interesting things (such as runtime model evals) we now adding. If this would be of interest -- open source of course -- please drop me a note here, it will incentivise me to finally extricate it from our broader system:

Slopsite here: https://piproxy.latchlabs.dev

redmalang··on Benchmarking coding agents on Databricks' multi-million line codebase
This is in fact what we do (with higher order abstractions now built on top of this). This builds self evolving interactive knowledge base and puts it into a QMD searchable index. The indexer is already open source: https://github.com/jibs/duffel
redmalang··on Benchmarking coding agents on Databricks' multi-million line codebase
Aside 2: Anecdotally we found that Pi performs more or less on par with native harnesses at lower cost on decently specified prompts. It is also phenomenal at context cacheing especially on Deepseek models (its hard to precisely attribute credit here are my understanding is this is a DS speciality). But it fails much worse on poorly drafted prompts. I'm generalising but native harnesses seem to be better kind of flailing along on those.
redmalang··on Benchmarking coding agents on Databricks' multi-million line codebase
Yeah.
redmalang··on Benchmarking coding agents on Databricks' multi-million line codebase
We have an internal proxy (that I've been meaning to open source for ages) that routes all llm usage at our company, which allows us to see data in realtime. Its been fascinating how rapidly Pi has been adopted. Moreover since its pretty hackable, we've been able to automatically aggregate context from pi sessions, which has resulted in Pi efficacy being higher as more people use it, putting in place a interesting virtuous loop. I didn't expect this outcome: for whatever reason I assumed proprietary harnesses fine tuned to work with a companies' models would work better? ps/random aside: there is something slightly off about Pi's edit command, we are planning to investigate this further and patch this as we have quite a few session traces now..
redmalang··on GLM-5.2 is a step change for open agents
We have switched approx 80% of our work to deepseek, and it works great. Our setup is a bit unconventional though, we upload all cot / sessions to shared storage and generate centralised project level context. We've found this is helpful in directing and working with these slightly less sota models and getting great value for ai spend.

I'm planning to open source all this infra soon, hopefully useful for others too.

redmalang··on Running local models is good now
Try llama.cpp it seems to be a lot more performant and a lot more hackable. Also I'm surprised how substantial the impact of some of the inference configs (beyond just temp) can have, though this is much more model specific.
redmalang··on Ollama is now powered by MLX on Apple Silicon in preview
i've found llama.cpp (as i understand it, ollama now uses their own version of this) to work much better in practice, faster and much more flexible.
redmalang··on Show HN: I'm making an open-source platform for learning Japanese
Anybody aware of anything like this for mandarin ?
redmalang··on The infrastructure behind ATMs
In the UK at least, afaik the Consumer Credit Act applies to Credit Card purchases.

IANAL etc.

redmalang··on Borges: Recommendations from a life of lectures and essays
I love these lectures, his sonorous, lilting voice and surprisingly acute comedic timing, like a native english speaker. Never never quite realised it until you put it down like that, just how he weaves a whole cloth out of these cross cultural threads.

If there are any other hidden gems, these dialogues for example, please do share!

redmalang··on Borges: Recommendations from a life of lectures and essays
These lectures by Borges are seriously under-appreciated: https://www.youtube.com/watch?v=YSLV7t9DvN8

The fact that Borges was, apparently at this point blind and doing these lectures essentially from memory is just mind boggling to me.

redmalang··on Matt Levine makes sense of Wall Street like none other
Thank you :)
redmalang··on Matt Levine makes sense of Wall Street like none other
Thank you :)
redmalang··on Matt Levine makes sense of Wall Street like none other
I really enjoy his writing, which is why I really want to know what this might be:

"Asked about the Etruscans, Mr. Levine said he thought Mr. Mystal might be referring to one of his favorite anecdotes from Herodotus. It was actually about the Persians, he said. He fetched his copy of “The Histories” and read it to me.)"

redmalang··on Delivering Billions of Messages Exactly Once
BQ doesn't have primary keys. Perhaps you are thinking of the id that can be supplied with the streaming insert? This has very loose guarantees on what is de-duplicated (~5m iirc)
redmalang··on Delivering Billions of Messages Exactly Once
Another possible approach: https://cloud.google.com/blog/big-data/2017/06/how-qubit-ded...
redmalang··on A Slow-Motion Trainwreck Facing the Meal-Kit Industry
Whilst I'm not the target market (I enjoy cooking from scratch way too much), I think there are a few things wrong with the analysis:

- Surely this isn't the first business to make a loss on acquisition in expectation of breakeven. That in of itself doesn't make this untenable. I also feel that the Groupon analogy is a bit underbaked: Groupon (and the ilk) were also I suspect totally unsustainable for the end businesses.

- There is substantial scope for curation. This in & of itself may not not be enough of course, but this doesn't seem like total commodity land either. Also it isn't just about getting the ingredients bundled together. It is harder than you'd think making sure standard measures of stuff in the supermarket doesn't go to waste from spoilage when trying to cook whilst keeping repetition low. Optimising for this is probably an interesting and not super simple problem.

I've only been exposed to the UK variants. The above may not hold true for other places though.