As mentioned there's also asknew: https://news.ycombinator.com/asknew
2,574 karma · joined August 9, 2023
As mentioned there's also asknew: https://news.ycombinator.com/asknew
Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.
https://github.com/KupynOrest/s3od
(BLS gets this data from surveys: https://www.bls.gov/opub/hom/cex/home.htm )
(Whether this makes you more resistant to being "fired" is still up for debate of course.)
But, back in the day FIRE used to mean something beyond just being rich - there was an anti-consumerist bent (and expectation that you'd move away from your expensive city/former job) that usually went along with lower spending.
Even that isn't strictly necessary - you can get perfectly acceptable performance by splitting a model between multiple older 12 or 16 GB cards.
If you just want a website for cheap: Bearblog, carrd.co, etc.
if you want all the bells and whistles on a platter: Squarespace, Wix, etc.
if you want to supply all the HTML/CSS yourself: Github Pages or Cloudflare Pages.
(Later, if you want to host the above (except the "bells and whistles" tier) yourself: Hetzner, Digital Ocean, etc.)
(We've received similar guidance. Even better is that our provider, GitHub Copilot, does not provide usage information to individual users if there is no per-user budget configured. So we just fire our requests into the black box and at the end of the month when IT gets the bill we maybe get a talking-to if it's excessive.)
The perhaps greater problem though is that those tests are completely trivial.
Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention and started the effort to evaluate this idea. Ashish, with Illia, designed and implemented the first Transformer models and has been crucially involved in every aspect of this work. Noam proposed scaled dot-product attention, multi-head attention and the parameter-free position representation and became the other person involved in nearly every detail. Niki designed, implemented, tuned and evaluated countless model variants in our original codebase and tensor2tensor. Llion also experimented with novel model variants, was responsible for our initial codebase, and efficient inference and visualizations. Lukasz and Aidan spent countless long days designing various parts of and implementing tensor2tensor, replacing our earlier codebase, greatly improving results and massively accelerating our research.
In any case, if the authors considered their contributions equal, that's good enough for me.Regular Qwen 3.6 benchmarks slightly better and has much wider software support though, so this is probably of interest only to organizations which disallow models trained in China.