13,187 karma · joined October 12, 2007
Twitter: http://twitter.com/mbuckbee
email: mike@expeditedsecurity.com
They provide a set of tools and a context when to use those tools (much like a packaged version of a CLI+API and a skill) which makes them more discoverable than other options.
I've got a few folks using my open source data storage MCP - https://github.com/ExpeditedProjects/hutchdb - now and it makes a lot more sense than any other implementation for what they're doing
An underpinning of their model is that API calls / inference are similar across providers allowing for commodity tokenization cost comparisons, but the providers are beginning to shift to non commodity features that don't easily shift.
For instance, calling Gemini with "search grounding" isn't something that OpenRouter can do (they sub their own web search in), last I checked they weren't doing real time voice models, etc.
1. Long form task based examinations like this that test the ability of the model+harness to remain on task, tool calling, overall effectiveness and taste.
2. More direct 1:1 and qualitative comparisons that you might get with a tool like https://evvl.ai/ - which also uses OpenRouter and does similar one off model comparisons (or lets you use it as a MCP from your dev env to be like: "take the prompt from this loop and try it against these other models")
Kimi was half the cost and produced a near identical output.
On a fairly simple coding task Kimi was 2x the cost and 7x as slow (and gpt-5.6-sol was even cheaper).
Making the content auto expire after a short period of time greatly decreases the attractiveness of the site to lots of SEO spammers and other types of abuse, and if someone were to get something malicious or vile posted it will clean up after itself without me having to wade into things.
So I put together a quick comparison of the last couple iterations of Opus, Fable and now Kimi.
Kimi is cheapest by 5x but also slowest by 2x
The market's being split into
1. Longitudinal LLM observability tooling
Most eval startups have gone down the route of something more like being an observability platform for LLM inference. They want to be in your stack and running the inference to collect data on performance of it.
They collect things like how often a model returns JSON that's out of spec or returns values that aren't expected as well as general timing and cost info.
2. Safety Limiting / Pentesting
Say you're doing something in the medical field or that's sensitive in some way and you want to figure out what model has the best outputs for your task that won't fly off the guardrails.
3. Simple cost + performance + quality swapping
This is what my tool does, basically lets you test if you _really_ need to be running that frontier model in a loop across a million records or if you'd be better with an older model or something else.
Example eval: https://giyd8stidy.evvl.io
As expected, Fusion was 7x slower and 4x the cost.
This isn't a knock against it, just that it I think this places Fusion into a "use it only when you need it" category.
If you consider things like the machine learning filters in your smartphone camera and Google's AI Overviews for searches it's entirely plausible that the US is currently at 75%+ of AI usage.
There's always gaps between theoretical and practical, but to see China investing so hard in the future while the US digs in it's heels is infuriating.
- This is something I made for myself without a lot of commercial thought, so I still haven't thought through pricing + usage + limits + operational limits. In it's current wildly unoptimized state it's still very cheap to run.
- For the specific concern about API Key leakage there's not a lot I can do about that (that I'm aware of) as the logic of what gets sent is handled by the client AI. It is possible to pull down and audit both the tools + instructions that are published by the MCP server if there are concerns on that side.
I've a a handful of dev friends that have started to use it as well and give their feedback and it's been slowly growing as I've added sharing/invites.
I would absolutely not recommend putting big production data into it currently.
My vision for it was something more like how the #1 use of spreadsheets is actually people making lists and not actually people doing lots of calculations.
Given the uptake today (thanks everybody!) and your feedback (thanks Mystery-Machine) I'm going to work at addressing your concerns.
Uses local AI models and I was able to snag this great domain name.
https://finalfinalreallyfinaluntitleddocumentv3.com/
But hands down the most useful thing I've made is HutchDB, which is a MCP service that you can call from any AI chat or Agent setup to store data for you.
Literally from your AI you just say "save that to Hutch" and then it figures out:
- The schema + fields - Builds nice webviews (Kanban, Timeline, Grid, Calendar) - Lets you share the output with people
So people use it for all kinds of things like time tracking ("every hour save a summary of my activities to Hutch"), for Agent to Human handoff ("Every day check social media for mentions of my company and save them to Hutch").
I use it for things like recording all of our marketing activities and then having my AI compare those to signups for rough attribution, etc.
Dead useful and at https://hutchdb.com