HNHacker News
TopNewBestAskShowJobs

michaelbuckbee

13,187 karma · joined October 12, 2007

https://expeditedsecurity.com

Twitter: http://twitter.com/mbuckbee

email: mike@expeditedsecurity.com

submissionscomments
michaelbuckbee··on Dots: Always-on agents
There's still bounds to all of this. I _heavily_ use AI for support tasks but it's all on the investigation, root cause categorization and initial response generation which posts I draft to the helpdesk software which I tweak and approve (often just hitting send).
michaelbuckbee··on Phyllotaxis: An audio-reactive LED display
Neat trick to use mulberry paper under a clear printed part. I wonder if you could simulate stained glass like that.
michaelbuckbee··on AI-generated posters don’t have to be horrible
To me, they look very much like Canva style designs. Some pre-existing design template that someone took and then smooshed their own dates and information into.
michaelbuckbee··on Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA
I feel like even TPS is becoming less of a good metric as we're seeing certain models handle similar problems while burning far fewer tokens.
michaelbuckbee··on Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute
Movie production (editors, sound design, and a portion of fx work).
michaelbuckbee··on The entire city of San Francisco as a video game
I feel like the closest we've gotten to that so far has been the flight sims.
michaelbuckbee··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
It also distorts the testing as it encourages non-typical behavior.
michaelbuckbee··on New MCP Roadmap
I'm very bullish on MCP (or at least MCP "like" implementations), as they solve a lot of problems for non-devs as they're much easier and safer to add into ChatGPT, Claude and other desktop + web apps.

They provide a set of tools and a context when to use those tools (much like a packaged version of a CLI+API and a skill) which makes them more discoverable than other options.

I've got a few folks using my open source data storage MCP - https://github.com/ExpeditedProjects/hutchdb - now and it makes a lot more sense than any other implementation for what they're doing

michaelbuckbee··on OpenRouter is joining Stripe
Nice. Happy to be mistaken on that point as that seems very useful.
michaelbuckbee··on OpenRouter is joining Stripe
I've been using OpenRouter since relatively early (I build evvl.ai - an eval platform on top of it) and think that they're selling at a good time.

An underpinning of their model is that API calls / inference are similar across providers allowing for commodity tokenization cost comparisons, but the providers are beginning to shift to non commodity features that don't easily shift.

For instance, calling Gemini with "search grounding" isn't something that OpenRouter can do (they sub their own web search in), last I checked they weren't doing real time voice models, etc.

michaelbuckbee··on Cursor launches Origin, GitHub alternative
Mostly it seems like greatly increased usage pushing their systems to the limit.
michaelbuckbee··on Cursor launches Origin, GitHub alternative
https://www.githubstatus.com/history
michaelbuckbee··on Choosing an AI model: one prompt, 11 models, different results
The eval world is split into:

1. Long form task based examinations like this that test the ability of the model+harness to remain on task, tool calling, overall effectiveness and taste.

2. More direct 1:1 and qualitative comparisons that you might get with a tool like https://evvl.ai/ - which also uses OpenRouter and does similar one off model comparisons (or lets you use it as a MCP from your dev env to be like: "take the prompt from this loop and try it against these other models")

michaelbuckbee··on Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
Since most analytics is done with JS (Google Analytics, etc.) very little of this shows up in site visit stats.
michaelbuckbee··on U of Michigan drops first-semester grades to ‘curb mental health crisis’
Parent of a middle and a high schooler and we try very hard to not pressure the kids, but a lot of it is coming from their peers and the school.
michaelbuckbee··on DeepMind's WeatherNext model achieves breakthrough forecasting cyclones
The 2 vs 3 days makes less of an impact on personal decision making but has massive benefits for decision making at the country wide response level.
michaelbuckbee··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
To your slow point: I did a quick eval for a data viz task to do a qualitative comparison and Fable was much faster (by 4x), but I struggle with the "token inefficiency" as a sort of whatever metric.

Kimi was half the cost and produced a near identical output.

https://fr4geiw93g.evvl.io/

On a fairly simple coding task Kimi was 2x the cost and 7x as slow (and gpt-5.6-sol was even cheaper).

https://r26pakjmhv.evvl.io/

michaelbuckbee··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
It's kind of ridiculous how good these are getting. 3.5 Flash lite is pretty comparable to Opus 4.8 (at least for the couple tests I did) while simultaneously being 6x faster and 19x cheaper.

https://fy2zp1ri90.evvl.io/

michaelbuckbee··on Kimi K3, and what we can still learn from the pelican benchmark
It's a free site, so I was trying to limit both the privacy and risk exposure.

Making the content auto expire after a short period of time greatly decreases the attractiveness of the site to lots of SEO spammers and other types of abuse, and if someone were to get something malicious or vile posted it will clean up after itself without me having to wade into things.

michaelbuckbee··on Kimi K3, and what we can still learn from the pelican benchmark
Like Simon concludes the article, the main use of this isn't to say which model is "better", but to try and poke at the model to sort out things like quality vs cost vs speed.

So I put together a quick comparison of the last couple iterations of Opus, Fable and now Kimi.

Kimi is cheapest by 5x but also slowest by 2x

https://9gpyw4uxr2.evvl.io/

michaelbuckbee··on Why eval startups fail (2025)
I built a simple (free) eval tool for my own uses (Github Gists + Model Outputs) after not being able to find a suitable one in the market.

The market's being split into

1. Longitudinal LLM observability tooling

Most eval startups have gone down the route of something more like being an observability platform for LLM inference. They want to be in your stack and running the inference to collect data on performance of it.

They collect things like how often a model returns JSON that's out of spec or returns values that aren't expected as well as general timing and cost info.

2. Safety Limiting / Pentesting

Say you're doing something in the medical field or that's sensitive in some way and you want to figure out what model has the best outputs for your task that won't fly off the guardrails.

3. Simple cost + performance + quality swapping

This is what my tool does, basically lets you test if you _really_ need to be running that frontier model in a loop across a million records or if you'd be better with an older model or something else.

https://evvl.ai/

Example eval: https://giyd8stidy.evvl.io

michaelbuckbee··on John Jumper to join Anthropic
Vesting schedule?
michaelbuckbee··on If your product is Great, it doesn't need to be Good (2010)
Something that's improved my life has been buying a sticker sheet of those LED darkening dots. They're only a couple bucks and look much cleaner than other solutions I've tried while still allowing for _some_ light to come through.
michaelbuckbee··on Loreline – Tools for writing interactive fiction
This, more than anything else I'd ever read about Inform, really makes me want to give it a try.
michaelbuckbee··on Openrouter Fusion API
I ran a quick eval to see what this looks like qualitatively vs just calling Opus 4.7 or GPT 5.5 directly.

As expected, Fusion was 7x slower and 4x the cost.

This isn't a knock against it, just that it I think this places Fusion into a "use it only when you need it" category.

https://3fpi5avcqq.evvl.io/

michaelbuckbee··on No, everyone is not using AI for everything
A counterpoint to this is that we have some real different definitions of AI.

If you consider things like the machine learning filters in your smartphone camera and Google's AI Overviews for searches it's entirely plausible that the US is currently at 75%+ of AI usage.

michaelbuckbee··on Solar generates more energy in US than coal for first time
I think it's your last point that's actually the strongest.

There's always gaps between theoretical and practical, but to see China investing so hard in the future while the US digs in it's heels is infuriating.

michaelbuckbee··on Ask HN: What are tools you have made for yourself since the advent of AI?
I thought it was more implied, but let me be more explicit:

- This is something I made for myself without a lot of commercial thought, so I still haven't thought through pricing + usage + limits + operational limits. In it's current wildly unoptimized state it's still very cheap to run.

- For the specific concern about API Key leakage there's not a lot I can do about that (that I'm aware of) as the logic of what gets sent is handled by the client AI. It is possible to pull down and audit both the tools + instructions that are published by the MCP server if there are concerns on that side.

michaelbuckbee··on Ask HN: What are tools you have made for yourself since the advent of AI?
This is all very fair criticisms. This thread asked: "What are tools you have made for yourself?" and that's genuinely what this is. I wanted it so I made it and then AI makes it so easy to just throw up a marketing page.

I've a a handful of dev friends that have started to use it as well and give their feedback and it's been slowly growing as I've added sharing/invites.

I would absolutely not recommend putting big production data into it currently.

My vision for it was something more like how the #1 use of spreadsheets is actually people making lists and not actually people doing lots of calculations.

Given the uptake today (thanks everybody!) and your feedback (thanks Mystery-Machine) I'm going to work at addressing your concerns.

michaelbuckbee··on Ask HN: What are tools you have made for yourself since the advent of AI?
The funniest thing I've made is a free utility called "Moniker" that contextually renames files based on their contents.

Uses local AI models and I was able to snag this great domain name.

https://finalfinalreallyfinaluntitleddocumentv3.com/

But hands down the most useful thing I've made is HutchDB, which is a MCP service that you can call from any AI chat or Agent setup to store data for you.

Literally from your AI you just say "save that to Hutch" and then it figures out:

- The schema + fields - Builds nice webviews (Kanban, Timeline, Grid, Calendar) - Lets you share the output with people

So people use it for all kinds of things like time tracking ("every hour save a summary of my activities to Hutch"), for Agent to Human handoff ("Every day check social media for mentions of my company and save them to Hutch").

I use it for things like recording all of our marketing activities and then having my AI compare those to signups for rough attribution, etc.

Dead useful and at https://hutchdb.com

Page 1 of 34Next →