3,829 karma · joined September 15, 2019
Socials: - calendar.app.google/YhrgwoqZu3MdsaBs6 - github.com/ofou - linkedin.com/in/ofou - x.com/omarnomad Interests: AI/ML, Data Science, Digital Nomad, Education, Entrepreneurship, Open Source, Research, Science, Startups, Technology, Books, Climate Tech, Networking ---
Press release: https://cenia.cl/2026/02/10/latam-gpt-la-primera-ia-regional...
Presentation (in Spanish): https://www.youtube.com/watch?v=FdLzAiQizhA
HF: https://huggingface.co/latam-gpt
Github: https://github.com/latam-gpt
Sometimes it's not just raw intelligence, but engineering as well
[1]: https://transformer-circuits.pub/2021/garcon/index.html
If you use UTF-8 directly as tokenizer, this problem becomes evident once you fit it into the context window. Plus, you can run multiple tests for this type of injection; no emoji should take more than up to 40 bytes (10 code points * 4 bytes per code point in the worst case). This is an attack on tokenizers, not on UTF-8.
Plus, Unicode publishes the full list of sequences valid containing the ZWJ character in emoji-zwj-sequences.txt
Just thank you.
https://software.annas-archive.li/AnnaArchivist/annas-archiv...
Pricing for o3-mini [1] is $1.10 / $4.40 per 1M tokens.
[1]: https://platform.openai.com/docs/pricing#:~:text=o3%2Dmini
Nevermind, it's here
I read your book, The Coding Career Handbook, we need something similar for AI Engineering! I really enjoyed it. Thank you for creating and sharing such high-quality multimodal content :)
Here’s an example: https://d2l.ai/chapter_natural-language-processing-pretraini...