312 karma · joined November 3, 2025
Issues:
• Ollama: https://github.com/ollama/ollama/issues/15931
• HuggingFace: https://github.com/huggingface/transformers/issues/29279 (labeled as feature request) https://github.com/huggingface/transformers/issues/47822 (with a bug label)
• Gemma4: https://github.com/google-deepmind/gemma/issues/768
The problem is that for tools like Ollama there is no solution except of to fix it by Ollama developers
More context
Why this is important: prompt injections are pretty dangerous and as long as user input is an instruction to the model it could be decided as code injections vulnerability. And in combination with long-memory and multi-agentic runtimes this vulnerability could stay in system for a long time. And it has not been decided as a serious security vulnerability by the global community yet
Temporal Solution
So if you're running local models just make sure to throw an error when there is a special sequence in user input. For Gemma family it is <|turn> and for tiktoken-based models it's <|im_start|>. And would be nice to see more solutions
Disclaimer
I'm not a security expert
The companies were notified 30 days ago about the issue. Only Google responded with a feedback on the issue (swiftly)
Number of parameters: 319K
(Disclaimer) The models is tiny and is not a pocket chatbot. It mistakes and is not capable to support conversation, but that's not the goal of the project
The goals of the project is to bring training to edge devices and it worked out
How can one use it? By using solar panels such device could be turned into autonomous meteorological station
• Q8: 8-bit 1.56 TB, lossless
• Q4: 4-bit, 1.51 TB
• Q2: 2-bit: 861 GB
• Q1: 1-bit, 594 GB
The smallest Q1 model keeps 78.9% accuracy, while being almost 3 times smaller than the original one. Instruction for running the model is in the model's card
Due to new learning technique the model has achieved better generalization skill without overfitting and memorization. This become possible because of new learning method which made the student model to generate samples for itself. It led to intensive reuse of existing neurons and allowed to encode information in a more dense way
While researchers are calling it a compression, I think it's a retopologization, and Microsoft had tried to do something similar in the past with their Phi model family, which they trained on reduced dictionary and simplified knowledge base first. But it seems like MS' researchers didn't explore this exact way of learning. I believe this should give even better results in the future and this is another small breakthrough moment
Paper: https://arxiv.org/html/2607.11883v1
Repository: https://github.com/shikaiqiu/requential-coding
Here the demos on HuggingFace: https://huggingface.co/collections/webml-community/transform...
It's just a surprise to see what can be done with the models in browsers today. This demos shows the abilities of the models, and this is the time for creators to bring their ideas and make solutions for real tasks
This release also adds new models to be run in browser Mistral4, Qwen2, DeepSeek-v3 and others. It has limited number of changes, what makes it pretty stable for a major version
What's about code and DX: it's not a good practice to export anything using globals, this is what JS world refused to do long ago. It turns your code into a hardly debuggable mess quickly
Why did you chose AtProto?
* Docker
* WASM
* Rustlang
* Web itself
Also it helps to start with typescript faster and easier, and to make the learning curve smoother and maintaining less complex for all developers on all platforms
YCombinator is about products, not technologies, so their example is not suitable here. Agents is a technology. They have fundamental advantage over premade software. They are already here and they are doing valuable work. They will stay for the near future. You can argue against reasons to use particular agents or particular products built around them, or particular features, but not against the whole technology.
The technology is verified, we need to find the way to make it safer and more sustainable. This would require to solve many technical tasks: build safeguards, create eco-friendly energy production, solve explainability, build physical infrastructure, develop new technical standards, etc.
You can write old-styled software for those who have nostalgic feelings about good old times, or support solutions written for old platforms. Instead of been thinking about it as of trash, think about it as it is retro
First plane didn't have cockpit, and windows, could fly only hundreds of meters, and looked more like a clothes dryer than a plane. Someone could call it inefficient (and really it wasn't efficient). It was easy to call it dangerous. But as we can see, planes evolved since then. When the first radio was invented there was no radio station to listen. When the steam train was invented it was crazy expensive and there was no government which has enough money to build all the railroads we have today
All of these technologies were modified by talented people, some of them were inventors, others later adopters. But they changed how this technologies work, and what goods they brought to us
AI is too young technology to tell if it fail or not. You can critisize it or you can change the direction of the progress. It's up to you
• Improved UI
• Added segments: filters to reuse
• Added cohorts: groups for common events
• Links and pixels are new elements for tracking
• New admin page
Breaking change:
• No more MySQL
What's so non-coop in this? This huge PR is a non-coop work and it requires to be fixed. Just imagine if someone locks the work of the whole branch just to review 9000 LOC, which they even didn't wrote. It's just no-go option
And even if there would be an AI which explains what this code is doing, people wouldn't be able to check it manually in reasonable time. It just enormous piece of work. So until there is a solution to this, such PRs should be declined and rewritten
Also it seems like the worker who brought this code doesn't know how evaluate the complexity of the task relatively to the solution, so there is a question about their qualification and level