HNHacker News
TopNewBestAskShowJobs

ofou

3,829 karma · joined September 15, 2019

AI Engineer + Music Producer https://twitter.com/omarnomad

Socials: - calendar.app.google/YhrgwoqZu3MdsaBs6 - github.com/ofou - linkedin.com/in/ofou - x.com/omarnomad Interests: AI/ML, Data Science, Digital Nomad, Education, Entrepreneurship, Open Source, Research, Science, Startups, Technology, Books, Climate Tech, Networking ---

submissionscomments
ofou··on Dutch governments builds alternative for Microsoft based on NixOS
why not just use linux/?
ofou··on Two-tier encryption in the UK
The UK is becoming 1984.
ofou··on AirPods 5
I want a simple USB-C earpods with noise active cancellation. How is that hard for Apple?
ofou··on 'VPNs are lawful technical tools,' says EU Court in landmark copyright ruling
privacy is a right, until you don't have a voice to say so.
ofou··on PyTorch: A Reference Language
this is just https://github.com/tinygrad/tinygrad
ofou··on Anna's Archive loses $322M Spotify piracy case without a fight
Demoniac move by Spotify
ofou··on A language model made in Latin America, for Latin America
The model released is called Latam-GPT, the first large-scale open-source language model developed specifically for Latin America and the Caribbean. Its size is 70 billion parameters (70B). It is a continuous pre-training (CPT) adaptation of Meta’s Llama 3.1 70B base model, trained on a regionally-curated corpus of ~300 billion tokens that emphasize Latin American languages, cultures, and contexts.

Press release: https://cenia.cl/2026/02/10/latam-gpt-la-primera-ia-regional...

Presentation (in Spanish): https://www.youtube.com/watch?v=FdLzAiQizhA

HF: https://huggingface.co/latam-gpt

Github: https://github.com/latam-gpt

ofou··on I made my own Git
btw, you can change the hashing algorithm in git easily
ofou··on Provider Variance: Introducing Exacto
This is great news for the community. We should measure a bit more closely how tool usage behaves in the wild. The gains you can get from different providers are not marginal.

Sometimes it's not just raw intelligence, but engineering as well

ofou··on UTF-8 is a brilliant design
UTF-8 should be a universal tokenizer
ofou··on Anna's Archive: An Update from the Team
the internet's own boy :( I'll be always deeply touched by him
ofou··on Anna's Archive: An Update from the Team
Shadow libraries maintainers deserve a Nobel prize for their contributions to humanity. Satoshi would be proud.
ofou··on The bitter lesson is coming for tokenization
This is stupid because a UTF8 is a tokenizer that covers all Unicode with a vocab of only 256 (yes, without a K). This is the only way of scaling the bitter lesson with tokenizers. Also, with architectures that span +1M context windows, it’s no longer an argument/issue the reduced context windows.
ofou··on Tracking Copilot vs. Codex vs. Cursor vs. Devin PR Performance
I'd submit a PR with this idea to improve coverage of agents
ofou··on Open-sourcing circuit tracing tools
Is this Garcon [1], or a new tool?

[1]: https://transformer-circuits.pub/2021/garcon/index.html

ofou··on Ask HN: Share your AI prompt that stumps every model
No luck so far with: When does the BB(6) halt?
ofou··on Ask HN: Share your AI prompt that stumps every model
considering the amount of bots in HN, not really that much
ofou··on Kilo Code: Speedrunning open source coding AI
mmm... like a whole Harry Potter series
ofou··on Winners of the $10k ISBN visualization bounty
My most sincere love to all shadow libraries out there, you're doing god's work.
ofou··on DeepSeek open source DeepEP – library for MoE training and Inference
You gotta love these guys, they're really pushing the open source frontier for all of us, thanks for sharing
ofou··on Perplexity Deep Research
https://www.emergentmind.com also offers Deep Research on ArXiv papers (experimental)
ofou··on Smuggling arbitrary data through an emoji
This is one of the reasons I've been advocating to use UTF-8 as a tokenizer for a long time. The actual problem IMHO are tokenizers themselves, which obscure the encoding/decoding process in order to gain some compression during training to fit more data in for the same budget, and arguably gaining some better understanding from the beginning. Again just a lack of computing power.

If you use UTF-8 directly as tokenizer, this problem becomes evident once you fit it into the context window. Plus, you can run multiple tests for this type of injection; no emoji should take more than up to 40 bytes (10 code points * 4 bytes per code point in the worst case). This is an attack on tokenizers, not on UTF-8.

Plus, Unicode publishes the full list of sequences valid containing the ZWJ character in emoji-zwj-sequences.txt

ofou··on Meta torrented & seeded 81.7 TB dataset containing copyrighted data
Who would have known that BitTorrent, shadow libraries, and seeders will help to train the best AI models out there, that adds a whole new meaning to a "seed".
ofou··on Visualizing all books of the world in ISBN-Space
It’s called DeepSeek. The founder just confirmed a few days ago that he got the data from Anna's to train on, I think for their latest vision model.
ofou··on Visualizing all books of the world in ISBN-Space
This is a wonderful submission to Anna's archive [1]. I really love people pushing the boundaries of shadow source initiatives that benefit all of us, especially providing great code and design. Can't emphasize enough the net plus of open source, BitTorrent, and shadow libraries that have had in the world. You can also make the case that LLMs wouldn't have been possible without shadow libraries; it's just no way of getting enough data to learn.

Just thank you.

https://software.annas-archive.li/AnnaArchivist/annas-archiv...

ofou··on OpenAI O3-Mini
I find quite interesting they're releasing three compute levels (low, medium, high), I guess now there's some way to cap the thinking tokens when using their API.

Pricing for o3-mini [1] is $1.10 / $4.40 per 1M tokens.

[1]: https://platform.openai.com/docs/pricing#:~:text=o3%2Dmini

ofou··on DeepSeek R1 Is Now Available on Azure AI Foundry and GitHub
competition is beneficial for all of us, this is great
ofou··on Searching for DeepSeek's glitch tokens
Where can I get the actual tokenizer data?

Nevermind, it's here

https://api-docs.deepseek.com/quick_start/token_usage

ofou··on AI Engineer Reading List
By the way, I think this is a fantastic reading list for creating AI products, and especially for staying updated on the latest in the AI space. However, it feels a bit scattered and might be hard for beginners to follow, IMO.

I read your book, The Coding Career Handbook, we need something similar for AI Engineering! I really enjoyed it. Thank you for creating and sharing such high-quality multimodal content :)

ofou··on AI Engineer Reading List
Dive into Deep Learning is implemented using various libraries such as PyTorch, NumPy/MXNet, JAX, and TensorFlow.

Here’s an example: https://d2l.ai/chapter_natural-language-processing-pretraini...

Page 1 of 10Next →