HNHacker News
TopNewBestAskShowJobs

boyter

8,190 karma · joined February 16, 2010

websites: https://searchcode.com/ https://boyter.org/ https://bonzamate.com.au/ email: ben@boyter.org twitter: @boyter
submissionscomments
boyter··on Sloc Cloc and Code 4.0 (scc) – Finding the files that need the most attention
Thats what this should be able to do for you. Feel free to hook it in and see if it works. If so please let me know!
boyter··on Sloc Cloc and Code 4.0 (scc) – Finding the files that need the most attention
Probably? You could look at hotspots using it and determine a human in the loop must happen.

If you do hook this up please let me know, id love to write about it.

boyter··on Sloc Cloc and Code 4.0 (scc) – Finding the files that need the most attention
Not one I have come across. Now in my list of reading.
boyter··on Sloc Cloc and Code 4.0 (scc) – Finding the files that need the most attention
Its what I used to migrate to the 6 TB sqlite database... so yes? Depends on your definition of backup. In my case I wanted a way to restore in the case of data loss, so yes. Im sure some pedantic sql admin will disagree with me here.
boyter··on Sloc Cloc and Code 4.0 (scc) – Finding the files that need the most attention
You are welcome.

Id like to see that fork. Perhaps I can fold some of it in.

boyter··on Sloc Cloc and Code 4.0 (scc) – Finding the files that need the most attention
Thanks mate.
boyter··on I regret migrating to Codeberg
Tooting my own horn here for search but give https://github.com/boyter/cs a go. You can have a http server up if you like (even modify the templates) and have it sync your repos for you on a schedule you define.

Let me know if it fails in some area and I’ll fix it.

boyter··on John C. Dvorak has died
It came to a head with that story about the kid making a clock and getting in trouble. Leo was all in on one side that it was about racism and such, with John suggesting he fell for a media lie.

I think John turned out to be right in the end. They did seem to patch things up with John appearing on a episode years later I think "a christmas miracle" was the title, but that was the last time he was on as far as I know.

boyter··on John C. Dvorak has died
When Leo cut off Dvroak and CJal I turned out. You really do need that sort of personality for people to bounce off otherwise its less entertaining.
boyter··on Is Grep All You Need? How Agent Harnesses Reshape Agentic Search
Came here to post than and you already did. Thank you!
boyter··on Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep
You are welcome. Glad to hear its working for you. I have a few ideas I am working on to improve its relevance too that I hope pan out.
boyter··on Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep
Grep prints out every matching line. For some searches a LLM might do it will get a lot of noise, and it might have to make that search because it cannot be specific. Targeted search can reduce the number of tokens.

I suspect this comparison is against reading the whole codebase though compared to just getting the bits you need.

boyter··on Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep
Interesting. I too have been working in this space, though I took a different approach. Rather than building an index, I worked on making a "smarter grep" by offering search over codebases (and any text content really) with ranking and some structural awareness of the code. Most of my time was spend dealing with performance, and as a result it runs extremely quickly.

I will have to add this as a comparison to https://github.com/boyter/cs and see what my LLMs prefer for the sort of questions I ask. It too ships with MCP, but does NOT build an index for its search. I am very curious to see how it would rank seeing as it does not do basic BM25 but a code semantic variant of it.

This seems to work better for the "how does auth work" style of queries, while cs does "authenticate --only-declarations" and then weighs results based on content of the files, IE where matches are, in code, comments and the overall complexity of the file.

Have starred and will be watching.

boyter··on Ask HN: What are you working on? (May 2026)
Been working on https://searchcode.com/ again which I bought back, albeit as code search tool for LLMs. It solves the “should I use this library” by allowing the LLM to inspect search and analyse it before integration. Can use it to compare multiple repositories before downloading. It comes with a large amount of token savings and can be really useful when wanting to learn about a codebase.

Since it does it anyway I added dossier pages to it as well https://searchcode.com/repo/github.com/rust-lang/rust Which is useful for humans, and shows what the system is creating.

Best part is that I get to use the tools I have built, so https://github.com/boyter/scc and https://github.com/boyter/cs to improve it which benefits anyone using those tools.

boyter··on I quit drinking for a year
Many Muslims drink anyway. A lot of those in Iran brew wine/beer in their house.

Tobacco in Australia has been taxed to the point we have a huge black market for it now. You would have thought people would have learnt from prohibition.

You cannot police morals.

boyter··on I quit drinking for a year
Id be ok with that if wine had the same taste. No alcohol free wine tastes even close, and none of them are good in their own right.

Some of the non-alcoholic beers are pretty good though and I am happy to drink them.

boyter··on Ask HN: What Are You Working On? (April 2026)
I reimagined https://searchcode.com/ since I realised LLMs have issues when it comes to understanding code you want to integrate. It’s useful for looking though any codebase, or multiple without having to clone it.

I use it when I have candidate libraries to solve a problem, or I just want to find out how things work. Most recently I pointed it at fzf and was able to pull the insensitive SIMD matching it uses and speed my own projects up.

I can’t find it right now, but there was a post about how ripgrep worked from a someone who walked through the code finding interesting patterns and doing a write up on it. With this I get it over any codebase I find interesting, or can even compare them.

boyter··on Fast regex search: indexing text for agent tools
I read this when it came out and having written similar things for searchcode.com (back when it was a spanning code search engine), and while interesting I have questions about,

    We routinely see rg invocations that take more than 15 seconds
The only way that works is if you are running it over repos 100-200 gigabytes in size, or they are sitting on a spinning rust HDD, OR its matching so many lines that the print is the dominant part of the runtime, and its still over a very large codebase.

Now I totally believe codebases like this exist, but surely they aren't that common? I could understand this is for a single customer though!

Where this does fall down though is having to maintain that index. That's actually why when I was working on my own local code search tool boyter/cs on github I also just brute forced it. No index no problems, and with desktop CPU's coming out with 200mb of cache these days it seems increasingly like a winning approach.

boyter··on Ripgrep is faster than grep, ag, git grep, ucg, pt, sift (2016)
The disk cache has a huge impact. However they claim it’s for multiple searches so it should be in it.
boyter··on Ripgrep is faster than grep, ag, git grep, ucg, pt, sift (2016)
Such a good read. I actually went back though it the other day to steal the searching for the least common byte idea out to speed up my search tool https://github.com/boyter/cs which when coupled with the simd upper lower search technique from fzf cut the wall clock runtime by a third.

There was this post from cursor https://cursor.com/blog/fast-regex-search today about building an index for agents due to them hitting a limit on ripgrep, but I’m not sure what codebase they are hitting that warrants it. Especially since they would have to be at 100-200 GB to be getting to 15s of runtime. Unless it’s all matches that is.

boyter··on Floci – A free, open-source local AWS emulator
Its pretty easy to step over those limits.

Also localhost and presumably this are good for validating your logic before you throw in roles, network and everything else that can be an issue on AWS.

Confirm it runs in this, and 99% of the time the issue when you deploy is something in the AWS config, not your logic.

boyter··on Ask HN: How are you all staying sane?
All of the above.

Stop reading the news. It makes you depressed or angry. Go hiking. Walk on the beach. Play with a dog or your children. Climb a tree.

Leave the slave slab phone at home, or delete every news and social app. Do not browse the web. Take a book and read.

It will be hard at first. Then it gets easier. Best thing I ever did.

Reminder. What passes for news today wouldn’t have registered for most people 100 years ago.

boyter··on Intel XeSS 3: expanded support for Core Ultra/Core Ultra 2 and Arc A, B series
Chased the wrong thing. It’s the 1% lows that matter more generally.
boyter··on Intel XeSS 3: expanded support for Core Ultra/Core Ultra 2 and Arc A, B series
If you have a high frame rate to start with it’s pretty nice and feels smoother. But a low frame rate turned into a high one looks good but feels laggy.

So arguably you never need frame gen for a game, since it only really works when it’s already pretty nice.

boyter··on Show HN: CS – indexless code search that understands code, comments and strings
Not familiar with that tool. What follows is my best guess based on what I am seeing.

Serena looks to be a precision tool. Since it uses uses LSP its able to replicate a lot of what a IDE would allow and IDE for LLM's.

cs by contrast is more of a discovery tool. When you're trying to find where the work actually happens it can help you, and since there is no index involved you can get going instantly on any codebase while they are index.

You could use cs for instant to find where the complexity lies, and then use Serena to modify it.

boyter··on Show HN: Lightwave – Real-time notes app, 3.5 years of hand-rolled JavaScript
Seems to be a load issue, hopefully easily resolved

    Request URL https://lightwave.so/api/register/ephemeral
    Request Method POST 
    Status Code 429 Too Many Requests
boyter··on What they don't tell you about maintaining an open source project
People who give away things like this tend to be good people. As such when someone comes asking for help or new things they are inclined to help.

Your response is where it should go when things get rude, but you don't want to start there.

boyter··on A critical look at NetBSD’s installer
I had never seen this before. Although I am now trying out lagrange and seeing what it can do.

Sorry for hijacking the thread on what is a great post, but is gemini common these days? This is literally the first time I have ever seen it, and it seems fairly interesting, albeit deliberately limited.

Annoyingly the learn more about it link https://gemini.circumlunar.space/ is now dead, as possibly might be protocol.

boyter··on Show HN: Minimal MCP server in Go showcasing project architecture
Perhaps adding some guide on how to hook this up to... well anything would be good :)

There is a lack of guidance for https://github.com/mark3labs/mcp-go/ which this is using as well so while everything is there, its hard to know how to make it do anything.

boyter··on Open-Source Is Just That
Ignoring the open-source vs free software discussions that are bound to come about from this well said. Large companies exploiting developers and abuse towards the maintainers is probably my biggest bugbear when it comes to this.

In fact I have a similar post https://boyter.org/posts/the-three-f-s-of-open-source/ which I redirect people towards if they become aggressive towards me when I am trying to help them. Thankfully I have only had to use it a handful of times.

Page 1 of 34Next →