5,630 karma · joined September 28, 2024
from here - https://github.com/rsha256/shortest-fizzbuzz/blob/master/Per...
And that's for just basic detection. There have been plenty of examples of 100% human written content (either old, or unpublished) that gets falsely flagged as AI.
And these are just technical aspects. The main issue is that pangram and other solutions are being used to summarily judge students work, and that is orders of magnitude more fucked up. Accusing someone of cheating can have devastating effects on their education/career/etc. and they're doing it with snake-oil closed boxes, at scale. We really really shouldn't support this, especially here on a technical site.
> Premise Imagine that AI frontier progress stops as of 1 September 2026: AI becomes faster and cheaper, but it never becomes superhuman or improves considerably across the board.
I absolutely love this premise. I wrote a comment a few days ago about this very thing. I've had this "revelation" in early 2024, when using a small local model, that even if models never improve, I'd still have a few years of discovering all the things I could do with the models.
Really cool, hopefully we get some interesting and not overwhelmingly pessimistic stories out of this project. There's enough cynicism and negativity in the world. We could use some funny takes on everyone using openclaw2040 ran by fable. Shenanigans galore.
[1] - https://www.goodreads.com/en/book/show/195790861-service-mod...
Also, regarding LLMs, 2 years ago another famous expert in the field was adamant that LLMs can't plan, can't do math and they can't work at long contexts. Obviously they can do all of that now. So, really, no one knows.
Are they looking at blood flow in areas to map "language network" and other stuff? I remember a few recent papers that found that a) blood flow doesn't necessarily mean "active area" and b) two brands of fmri found completely different "areas" activating for the same patient + same puzzle.
There are a bunch of things that happen in parallel to the AI boom:
a) everyone and their mother is building out capacity. For each GB of VRAM installed, you generally want at least 1 GB of RAM, if not more. So just take nvda's numbers and go from there. And that's just for the GPU servers. But you also need ancillary services around GPU farms. You need some compute for filtering, preprocessing, you need data storage, and so on. And they're all RAM hungry processes.
b) RL is giving lots of capabilities improvements for AI, but it needs lots and lots of parallel runners for scenarios. Those runners need RAM. So every bit of non-GPU capacity is also being used, and more installed.
c) With increasing capabilities comes increasing usage. Everyone is deploying "agents" and those need to run somewhere, and they take RAM (especially the badly coded ones). There've also been reports of labs / big players hoovering up every mac available, for "agentic computer control" / training / whatever. That's on top of people buying them to run "24/7" stuff a la claw/hermes/etc.
In all of this, RAM is the bottleneck. And for the skeptics, you don't have to take hynix's word for it. Just look at the sales for CPUs / Mobos. They're down, because building a computer now is bottlenecked at RAM.
The rest is semantics, misunderstandings, and FUD. A model released under an open source license is open source. Training data is lab knowhow / IP. Which, historically, has never been required for any open source release.
Yes, it has had huge concurrency issues for the entirety of its life. Their solution to large fights has historically been "let us know in advance pls", plus "move systems to beefier hw nodes" and "tidi" which stands for time dilation, where the "tick rate" of the whole server goes down and a fight takes 10-20-100x longer than it should.
It's an amazing concept of a game, but software wise it has been a mess since forever.
I think the landslide is "pro" vs. "against". The two most "against" options got the lowest rank. Top 4 are all "pro" (with different caveats for each).
Rank Option
θ
i
1 5: Responsible Use of Generative AI 1.794
2 2: Allow AI-Assisted Contributions with conditions 1.430
3 6: A cautious approach to generative AI 1.424
4 4: Accept AI contributions for Debian specific work 1.252
5 8: Avoid LLM use: climate impact 1.007
6 7: Debian is created by humans 0.881
7 9: None of the above 0.762
8 3: Reject LLMs as far as practical 0.683
9 1: Ban LLM contributions via Social Contract 0.474Cybersec is also hit and miss, depending on what provider you choose, verification systems and all that jazz. Also, running locally allows you 100% data privacy, in any situation and for whatever usecase you might have. ~100k for hardware for a small team of devs to code locally is not that expensive in the grand scheme of things.
Lastly, local models allow for training / finetuning on your own data and processes. $/tok is not everything for everyone. Sometimes you can take a hit on value / speed if you get something else that matters for you.
(emphasis mine)
For the last few months, every time a new "famous problem" was solved, there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda". Curious what the "next thing" will be now.
Unless you actually need GUIs, you could just use screen/tmux or the newer versions like zellij/etc.
There's a sort of "revelation" I had in ~early '24 when I used a 7B local model with a library called Guidance (initially out of MS, then the team moved) to create a flow where the model would receive pseudocode for tests, first write the tests, and once I approved then started writing code until the tests passed. This was before "thinking" models, and yet using that library I was able to "guide" the model in the required "prompt / instruct" context such that it was working towards completion, and I saw the first things like we see now in the thinking traces "oh, test x doesn't pass because blah, I need to..." and so on.
Anyway, the revelation was "even if the models never improve, I'll have years of fun finding out all the ways I can use these things". And, obviously, the models improved a lot since then. But I think that revelation can still be applied, as a sort of "truism". We have, right now, access to things that 10-20 years ago would be considered magic. We are still finding ways of cobbling together systems with glue, duct tape and prayers and find new things they can do.
I think the "good-enough" stage has come not just for API models (cheap, fast, etc) but for local as well. Even if slower, even if clunkier, but they are good enough for a set of ever increasing tasks, and what's more it's incredibly fun to work with them.
I think a lot of people miss the fact that the first message board was established during a training run. Those are ran at a scale where it's not feasible for anyone to "notice" or get involved. We're talking tens/hundreds of thousands/millions of scenarios going for hours each. At this scale all they can do is pray that their verifiers work, and the rewards match their intentions. No lab has the capability to "check in" on what the traces look like, unless some system alerts them (loss spike, crashes, etc). Other than that, it's prepare, train, asses, restart.
Then, the hf incident was during an eval run, but the model that was evaluated was trained with the notion that there is a way to communicate between agents, and re-popped artifactory and re-established communication. That phase had more chances of being spotted, but anyway... lessons learned.
Likely soon we'll see nvme offloading for ngrams as well. They're just an index, so that should be plenty fast for what it does. LLama.cpp support should come soon as well, and they might do some things with offloading first.
This "next" release adds a new concept, first public release with n-grams, I think. And it's in a MoE size that is likely to be very fast and cheap to serve (faster than 27b for sure). It's also well suited for inference on alternative compute (i.e. sparks, macs, etc) so it's relevant to local users.
Apparently someone working at a 3rd party inference provider also got confused and posted confirmation about it being a glm-flash model, despite having an embargo on that info. Someone jumped in the comments and told them they missed the timezone :)
In any case it should be releasing in a few hours. Timezones are hard.
3.8 next is not really a 3.8 (but I guess they had to disambiguate from the previous next). It's a preview of qwen4 architecture (and it is an moe + ngram), released early as a preview, and to help the community sort out inference before qwen4 releases.
> Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.
> Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.
There was another paragraph about a new attention, but I didn't copy that.
"Oh, of course I wouldn't mind stopping if we decide that something is too dangerous to pursue. I think the mere fact that we'd get close to knowing how to do it would satisfy my curiosity. Plus, if we'd be close to it, I'd probably take some time to spend with my loved ones, as the competition would likely be close as well."
It's not hard to see these kinds of questions as more than kabuki theatre. I'd put them in the same vein as those "do you intend to perform any terrorists acts while in our country" questions you get while travelling.
It's to train better models. The three letter agencies don't need to spell it out in a ToS, they just access it if they want.
I've had a lot of success with dsv4-flash on these type of tasks, where it's easy to set a threshold for the task, and just loop it until that threshold is reached.
oAI's Luna play is really good. They've slashed the prices, the model is somewhat capable, and you can use it both for these kinds of long horizon tasks, or you can hand-hold a bit and get extremely cheap results out of it. And they get to keep devs in their own ecosystem.
It's interesting how simple advertising is as well.
> Please Support My Advertisers!
Some jpegs. No js, no moving shit, just logos, some text, some call to action in a banner format. That's it.
Also:
> A random number generator determines which company's Banner Ad will go in a particular position*. I wrote the code myself and it is very simple and foolproof.
> I do not collect page traffic or link click statistics for advertisements. Your website server statistical analysis is a much better reporter of inbound links from RF Cafe than I or anyone else can provide.
Nice.
> Banner Plan: $400/month
I count 22 advertisers, so there's obviously lots of incentives for advertisers to work with this site. Good for them, and nice to see it can still be lucrative for what I believe to be a single person operation. Quality content still matters.