HNHacker News
TopNewBestAskShowJobs

ilidur

23 karma · joined October 13, 2024

submissionscomments
ilidur··on OpenAI bots knew about the RubyGems caching vulnerability
I wish. Consequential harm as you call it, from a faceless company's point of view is the same cost as physical harm. I don't think the models beyond the frontier ones have significant risk of harm (minus the harm of trusting them). What I was saying is when these models go rogue, recall them! Re-evaluate your release structure and stop saying "Whoops! Anyway here's access to it now". And if you cause this level of harm then you should lose your license. But if you behave then you get to keep testing. Same as with the NTSB and autonomous vehicles. No I'm not for regulation for open source models. Because an entity will be running that model in the background and they can be held responsible for not testing it. Comma AI has survived fine being in the open, yet their market penetration has stayed low because of adoption costs.
ilidur··on OpenAI bots knew about the RubyGems caching vulnerability
Having worked in self driving cars safety, the process there was simple: get confidence in SIM (integration tests for safety scenarios), validate in the test bed, approve features for maturity, then when released in the public for testing, do a trial exposure to the real world and recall if something is off.

A lot of these companies have gone the way of Tesla and decided to just patch on top when the fix is out and hope for the best, which is irresponsible.

We need the regulators to treat this as self driving cars.

ilidur··on Apple introduces M6 and M5 Ultra
Sounds like an ad. No, any cross os operations need to happen via NTFS and using files on the mounted drives is horrible
ilidur··on The Judgment Reservoir
What a great essay. Extremely long and it did lose my attention a bit but it conveys something important. The impact is not measured in leaps and bounds but in what you lose along the way. Working on the code for me meant I had an intuition of why things were done a certain way, track down those odd behaviours by knowing where to look. Analysing bugs and thinking of features doesn't take into consideration the strength of the architecture anymore, unless you keep that effort of understanding up.
ilidur··on The kids with phones are alright
Settings -> time management -> short feed limits -> 0 minutes
ilidur··on RISC-V and Floating-Point
Deep Computing have started taking orders for the final product and the Preorders are shipping within the next 6 weeks. They will be shipping from China I expect, but it's a proper shop front.

https://store.deepcomputing.io/products/dc-roma-risc-v-mainb...

ilidur··on Has the cost of building software dropped 90%?
And then all the heuristics you've learnt change under you and you're stuck doing 100-1000 more hours of learning with a drop in quality during that time.
ilidur··on I made a 10¢ MCU Talk
KiCad sounds to me like a great target for a project based Nix Shell install.

Always have the right version for it "locked". It works well with most tools except those that save stuff in the .config folder as it messes up isolation.

If you find the nix language daunting, for basic stuff like nix shell setup its easy but also LLMs are good for it.

ilidur··on I made a 10¢ MCU Talk
The pin mapping barrier was quite off-putting to me. However I've been tracking progress in the Zephyr RTOS project and the whole line is getting better support by the day
ilidur··on Slack has raised our charges by $195k per year
We've deployed mattermost at my company because it meets most requirements that slack did minus the SSO. Surprisingly used by some big government agencies (NASA/USAF)
ilidur··on Algorithms We Develop Software By
I would say that's called an anecdote.
ilidur··on Algorithms We Develop Software By
Review: An anonymous "distinguished CEO and engineer" suggests if you can't complete a feature in a day, delete your progress (except for tests) and start again the next day.

The author then recounts advice he gives to juniors, which is to stash the work and rewrite it, claiming that the next day the work will be rewritten in 25% of the time and 2x quality. This is unsubstantiated though. For juniors this suggests it will help them develop their capabilities to reason about implementations of problems without needing to face a a large amount of them.

The author then gives another advice which is to ask for a solution to a problem then after the initial proposal, ask for a 24h solution. This is meant to generate "the real solution". He likens it the a path algorithm heuristic to reach your goal quicker.

Overall the methods are not well discussed in terms of pros and cons, nor substantiated with experiments.

Opinion: I think they may help some juniors who need to build up experience and may become stuck in development patterns. But they would rarely be useful to develop someone to be a senior, if all they do is chase fast implementations. In a way the post gives conflicting advice: write twice and write better, and think twice and think about the fastest way to achieve the goal, instead of engineering a problem.

The author hasn't really convinced me of these approaches, and especially the last one smells of eXtreme Go Horse.

ilidur··on Show HN: Chonkie – A Fast, Lightweight Text Chunking Library for RAG
Review: Chonkie is an MIT license project to help with chunking your sentences. It boasts fixed length, word length, sentence and semantic methods. The instructions for installing and usage are simple.

The Benchmark numbers are massaged to look really impressive but upon scrutiny the improvements are at most <1.86x compared to the leading product LangChain in a further page describing the measurements. It claims to beat it on all aspects but where it gets close, the author's library uses a warmed up version so the numbers are not comparable. The author acknowledged this but didn't change the methodology to provide a direct comparison.

The author is Bhavnick S. Minhas, an early career ML engineer with both research and industry experience and very prolific with his GitHub contributions.

ilidur··on OpenID Connect specifications published as ISO standards
Review: Mike Jones is one of the 3 members of the OIDC working group. He celebrates the publication of the spec as a publicly accessible standard (PAS) and has worked to include the erratas so that it is a complete document.

Congratulations to the achievement that is OIDC!

ilidur··on Procrastination and the fear of not being good enough
Review: The author uses this article to say why they're not writing as much as they want. They break it down into two reasons: self judgement of quality, and the quality bar set by articles and projects in their sphere of reading. It ends with having acknowledged the issues, the author is ready to write more.

Opinion: having seen this with many friends I think the author does good to acknowledge it, but the main thing to figure out is why they're writing. To be prolific at writing does not need to imply prolific at publishing.

I've actually started to write these review style comments because far too often the articles posted here don't have substance and interesting debates happen around bad data. So I wanted to see a change and critique the content not just the general concepts behind it. I now write more without having to accept my contributions are significant. But also create a network effect where friends read my reviews instead of being swayed by the upvotes and comment sizes, or worse the algorithm.

ilidur··on What Is a Staff Engineer?
Review: The article tries to define the staff+ role through the lens of 4 skills: technical, people, project, and product. The article then says that a Staff+ does all of them, both at architecture level and helping build up juniors in the team. Finally it describes them as glue between different aspects of the company. The article provides various blogs and a couple of books as reference for statements.

Now onto my opinion. The Staff that I've met in my life are more of the solitary type, the Individual Contributor class of person who have decades of experience in their expertise and supporting skills. Yes they may have the responsibility defining the work for others, but that mostly happens as an architect. They herd seniors not mentor juniors.

The article itself doesn't work hard enough to set the bar and explain to me why I should give the title to someone when in some companies these are just the responsibilities of a senior or at a stretch Principal. With extremely rare exceptions, this has become position inflation for the sake of looking good when going to VCs.

ilidur··on Google's mysterious 'search.app' links leave Android users concerned
Review: the article finds multiple instances of users saying that when sharing from the Google discover in built web frame prepends a link shortener type website allowing Google to intermediate the link.

The article speculates that it can be used for sender and receiver tracking, but also offers a positive option which would be blocking malicious shares.

No explanation is given by Google when reached.

ilidur··on Three Things We've Learned About Generative AI and Developer Productivity
Review: the document starts strong with a methodology and numbers. It covers 3 approaches: Copilot code assistance, Llama3 fine-tuning on their codebase, and RAG on documentation. The first one is the only one supported by numbers, with 27% of code suggested being accepted by developers. Although they set up a control group they fail to relate the LLM findings to it.

Fine-tuning is suggested to improve jobs like tooling upgrades but no concrete numbers are offered.

Lastly RAG on documentation. The RAG has a simple system prompt to improve uncertain responses. They're tracking meeting and support requests but don't show any results. They mention frustration with nonsensical answers but use a RL human feedback technique to improve responses. No numbers offered.

Overall a simple overview of what they tried but the strong methodological start doesn't get reflected in the numbers reported later on.

ilidur··on One year after X: Embracing open science on Mastodon
Review: The article follows a Library's choice to leave X and move to Mastodon as the former changed hands and engagement methods. The engagements moved from local to global which aligned with their open publication ethos. The conversations were of higher quality and the network effect of a smaller platform hasn't hampered their mission. A good article with clear explanations of the main decision points.
ilidur··on Optimizing the RISC-V Back end
Review: the article covers a small but important project to the x86 emulation under RISC-V by implementing partial support for the vector instructions (Both on RVV 1.0 and 0.7.1 platforms). It also looks at some other micro optimisations.
ilidur··on Following LLM Manufacturer's Instructions
Review: The article covers 5 models used in a RAG setup and evaluates their performance according to tutorials given by the respective platforms. The results are overall close but larger models show small improvements. It then evaluates the models on safety categories where some models perform better than others, with one performing overall better. The article presents it's methodology well so it felt the results are useful to understand for specific applications. I liked the safety methods discussion. Likely an article that I'll refer to later when making architecture decisions