Most of the time 3.8 works fine, but it's a bit slow if compare to 3.7 Flash. If there's already a detailed plan, 3.7 can complete the task much faster. And the best thing about agy is the usage limit was very generous.
966 karma · joined January 16, 2015
Most of the time 3.8 works fine, but it's a bit slow if compare to 3.7 Flash. If there's already a detailed plan, 3.7 can complete the task much faster. And the best thing about agy is the usage limit was very generous.
https://notes-huy-rocks.translate.goog/posts/diy-pomodoro-ti...
(google translate link because the original post was in Vietnamese)
In fact with a 64GB mac, you can run pretty much all of the latest Qwen models.
Also, anyone who has been following local LLM are well aware that the quality and performance has become way way better since Qwen3.5
They have a nice UI, support deploy any kind of backend-involved apps as long as it can be built into a docker container. While many PaaS out there seems to prioritize frontend only apps.
And they have a free plan, so people can just quickly deploy some POC before decide if it's good to move on.
Anyone know if there is any other PaaS that come with a low cost starter plan like this (a side from paying for a VPS)?
What happened after this? the factory have to replace the casting mold at their own expense or you have to pay for it?
Was this change made by a mod or OP, and why would someone making that change? I do think the original title was more descriptive, and the new title was completely out of context, or it's imply that everyone is using Starlink and know what's Roam 50GB is.
I heard in Chrome, there's a gemini nano model built-in as well, maybe this is a good example to integrate it.
Another one but turned out it was never really a big deal: some chatbots from frontier AI labs started to support those niche features (people still coming to my app for the flexibility of using multiple AI models).
I think the biggest problem was #2, life kept pulling me the other way.
It's been a good journey. Thank you so much to whoever keeps running this thread!
Great project btw!
Speaking about total time/cost, this experiment cost me just $1.01 for 2h30 on a rental GPU. But the actual successful run was less than 10 minutes for both phases. The rest of the time I was spending fixing the code, tuning the params, train, and retrain. It took me about 6 hours to build and clean the two datasets, though.
For the next step, I'm thinking of improving the model accuracy, maybe with RL, but I would not go about shrinking the model size any lower. Prior to this, I've tried a lot of different model sizes on different kinds of tasks, from 135M to 4B. I'm not sure I like the performance of these small models for code generation :D