│------------------- │ Ninfer-3090 │ Strata
│ Code generation │ 52/78 (66.7%) │ 70/78 (89.7%)
│ Code completion │ 40/50 (80.0%) │ 44/50 (88.0%)
│ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)
│ API failures------ │ 10 │ 5
- Ninfer generation: ~122 min total.
- Strata generation: ~142 min total.
So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t/s, this modified version can do 50 t/s (but it's extremely long in it's thinking, it just goes on and on.
This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).
Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t/s.
That strata has been optimized on my Oh My Pi conversations. So when I'm using it, it's probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.
I've tried to use both for online comparison shopping. QFN not only seemed less delusional but also made useful observations and problem solved ways around many different website access issues.
│------------------- │ Ninfer-3090 │ Strata
│ Code generation │ 52/78 (66.7%) │ 70/78 (89.7%)
│ Code completion │ 40/50 (80.0%) │ 44/50 (88.0%)
│ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)
│ API failures------ │ 10 │ 5I now backordered two more sparks, before the hike... hoping they will honor the agreed price and dont cancel on me. They should arrive in a month,
Its evidently clear the frontier labs won't keep on giving us cheap inference for much longer. I also watch in disbelief how people say a sub is cheaper, which probably is. But you loose on so many other things (privacy, predictability, censorship...)
Known issue with this model, I recommend setting thinking to 'medium' instead of default 'xhigh'.
I’m doubting the quality of the model, even comparing with luna 5.6, it is giving such bracingly wrong answers. It claims other tools have changed the code when pressed why it made a mistake. That code or tool didn’t come close to working on the project. It is scary.
For me, I didnt know I was on 4x and I was able to, giving it enough permissions, with a good harness like PI, to figure things out and go 4x -> 8x -> 16x.
It probably will take some opening the case and switching ssds around, but it's amazing how powerful these models have become.
It will be less of an issue for people who have already got the equipment and the practice running locally than it will be for people who have simply been relying on ChatGPT.
But on the other hand I'm giving them my best and worst ideas and the value I get in return outweighs what I'm contributing to them. I also don't have the problems of initial HW cost and maintenance.
Do you trust that a company who can't even do basic IT ops monitoring are capable of respecting a "no training checkbox"?
Is it viable to start/stop it multiple times per day?
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.