HNHacker News
TopNewBestAskShowJobs

TofuLover

38 karma · joined September 20, 2023

submissionscomments
TofuLover··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
Er, yes, good point! We'll make sure billing is a bit more discreet than that!
TofuLover··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
These are great questions, thanks! I'll answer them in turn, in a list because I like lists.

* Confidentiality: no logging, third party analytics, or anything like that. More details in our Privacy Poilicy [1]. Our hosting providers will have their own policies, but we're not running a super private service like Proton or similar. Might do some kind of secure tenancy in the future if there's demand.

* Price: I think Runpod vs per-token are very different beasts and for different purposes. I really can't make a direct comparison, as it'll be based on use case, but we're going for convenience over price, so all else being equal I'd expect us to be more expensive for most users anyway (edit: i meant "than other API providers"! We'd definitely need to be cheaper or at least competitive with spinning up your own cloud infra. We'd have parallelism and economies of scale on our side for this). We have a lot of experience with running and optimising open models though, so that's part of the value proposition too.

* Subscriptions: Only API for now. Maybe subscription later but honestly we prefer simplicity. My own experience with subscription plans is that they're usually sold at a huge loss at first, then the price creeps up as the service is enshittified. That feels like a bit of a scam to get users, and that's not really what we're about. We want to provide something specific, and aren't really concerned about scaling as fast as possible. Maybe we'll provide subscriptions if there's a real demand for it, but no plans at the moment to do so.

* Methodology: we use abliterated models, but I've been advised to hold off talking about that for now. Might make a blog post about this though (when we have a blog).

* Cache: yeah about 90/10 for pricing. We're still trying to find the sweet spot for tuning eviction. Running LRU with no guarantee/storage at the moment, could probably be less aggressive with retention, but that also has privacy surface area implications. Ongoing conversation.

* Quantisation: my brother in christ, everyone runs quantised. :) We're initially targetting FP8 on most models, but have had great results with MXFP4 though. If we can pack more concurrency onto nodes without losing quality, we'll reflect that in pricing. Or we'll offer as a separate model for cheaper and give users the choice. Edit: I see you were asking specifically about speed, which MXFP4 doesn't improve, but maybe if there's demand we'll run other qaunts for speed increase, especially on the larger models.

* Engine: vLLM gang all the way! For now at least, as it's what we have most experience with, and we find it the most flexible. We've been experimenting with SGLang though, and there's definitely some interesting optimisations we could do with it.

Hope this answers your questions, at least the ones I could! The irony of that hasn't escaped me!

[1]: https://violentdelights.ai/privacy

TofuLover··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
We look forward to providing many headaches to our lawyers going forward.
TofuLover··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
More seriously though, I think we should be fine: we don't host any content, and what people do with the models is their own responsibility (legally speaking, in our jurisdiction, at least according to Claude -- we're talking to a real lawyer next week). Like any other provider, we offer no guarantees of sane, safe, or accurate results.
TofuLover··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
I guess we'll burn that bridge when we get to it!
TofuLover··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
Completely coincidentally, we're just about to launch a service that does exactly this (API access to uncensored open models)! We have a waitlist at the moment but will be live very soon!

https://violentdelights.ai

TofuLover··on To become a better writer, read as much as you can
This reminds me of an article [1] (that I think I may have even seen initially on HN) analysing what writing looks like when treated like TV (consciously or subconsciously). It's a fascinating read and it made me more mindful of interiority in my own writing.

[1]: https://countercraft.substack.com/p/what-not-reading-does-to...

TofuLover··on AI companies destroy physical books – let's scan rare books before it's too late
> This is just part of a CCP-aligned moral panic, along with the water use nonsense

What was nonsense about water use?

TofuLover··on Stop Doom Scrolling, Start Doom Coding: Build via the terminal from your phone
Why?
TofuLover··on Rouille – Rust Programming, in French
The closest I've ever felt to this as a native English speaker is reading words in music scores in English. I'm a classically trained cellist, and grew up learning notation with Italian and French words for directions and expression. I've never learned either of those languages, save the words used in music notation. Seeing a score with those words in English just feels... wrong. Not in any big way, but as you said: uncanny. Definitely get the "bad psuedocode" vibe, because to me it English in music notation feels similar -- like the person who wrote it didn't know what they were doing, even though the notation makes perfect sense and the music is good. It removes some of the flair of the art of the notation itself for me.
TofuLover··on CauseNet: Towards a causality graph extracted from the web
This reminds me of an article I read that was posted on HN only a few days ago: Uncertain<T>[1]. I think that a causality graph like this necessarily needs a concept of uncertainty to preserve nuance. I don't know whether this would be practical in terms of compute, but I'd think combining traditional NLP techniques with LLM analysis may make it so?

[1] https://github.com/mattt/Uncertain

TofuLover··on An illustrated guide to OAuth
I don't think the part about front and back channels is quite correct. GET and POST requests are both encrypted in HTTPS -- including the URL (but not the domain, as DNS resolution happens separately). Front and back channel are more to do with trust boundaries, and what information is public vs private from the client's perspective.
TofuLover··on OpenTF is now OpenTofu
If your experience of tofu is only the above, I completely understand your distaste for it. But I think you owe it to yourself to try better tofu, and not as a meat alternative. Tofu on its own doesn't have much flavour, but that's the point, you need to marinate it. Google can give you some tips (squeeze the water out, then soak it in something delicious -- hell, soak it in meat juices!), but I highly recommend trying some good, low moisture smoked firm tofu. It's so good, I often just snack on it, slicing it like a sausage. But it's also great in things like burritos to add a smokey kick. Try it!