HNHacker News
TopNewBestAskShowJobs

julien_c

1,982 karma · joined September 10, 2010

julien a/ huggingface.co
submissionscomments
julien_c··on The local LLM ecosystem doesn’t need Ollama
> Ollama eventually added ollama run hf.co/{repo}:{quant} to pull directly from Hugging Face, which partially addresses the availability problem.

uh actually, _we_ did (generates a Docker-style manifest on the fly)

julien_c··on Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI
and a lot of traction on paid (private in particular) storage these days; sneak peek at new landing page: https://huggingface.co/storage
julien_c··on Hugging Face Removes Singing AI Models of Xi Jinping but Not of Biden
[i work at Hugging Face]

BTW the reporting in this article is sloppy/incorrect

julien_c··on HuggingFace Training Cluster as a Service
Flamingo-style, see for instance the recently released IDEFICS: https://huggingface.co/blog/idefics
julien_c··on Hugging Face raises $235M from investors including Salesforce and Nvidia
we have a secret plan to change our logo to a Face Hugger some time in the future :)
julien_c··on AWS will offer HF’s products to its customers and run its next LLM tool
that would make for a quite different branding.

Might think about it:)

julien_c··on Show HN: Deepinfra.com Serverless AI model hosting (top models from HF)
quite cool! haven't tried it yet, but what's the latency on hot-loading a model? (for instance, loading `stabilityai/stable-diffusion-2-1` for the first API call)
julien_c··on Show HN: I stripped DALL·E Mini to its bare essentials and converted it to Torch
what's the inference time on M1?
julien_c··on T0* – Series of encoder-decoder models trained on a large set of different tasks
ArXiv link to the paper: https://arxiv.org/abs/2110.08207

GitHub repo: https://github.com/bigscience-workshop/promptsource

julien_c··on PaddleOCR: Lightweight, 80 Langauge OCR
FWIW this is extractive question answering so the output of the model can only be a span of the input paragraph.

SQuAD is the prototypical example of a dataset for this task, see https://rajpurkar.github.io/SQuAD-explorer/

julien_c··on I’m Peter Roberts, immigration attorney who does work for YC and startups. AMA
For France, you can check the French Tech Visa program: https://lafrenchtech.com/en/how-france-helps-startups/french...

I've done it for one of my team members, it's pretty easy.

julien_c··on OpenAI API
Not really on unsupervised/self-supervised data though, right?

(nor on the same scale of corpora, as far as I can tell)

julien_c··on Show HN: HuggingFace – Fast tokenization library for deep-learning NLP pipelines
TL;DR: Hugging Face, the NLP research company known for its transformers library (DISCLAIMER: I work at Hugging Face), has just released a new open-source library for ultra-fast & versatile tokenization for NLP neural net models (i.e. converting strings in model input tensors).

Main features: - Encode 1GB in 20sec - Provide BPE/Byte-Level-BPE/WordPiece/SentencePiece... - Compute exhaustive set of outputs (offset mappings, attention masks, special token masks...) - Written in Rust with bindings for Python and node.js

Github repository and doc: https://github.com/huggingface/tokenizers/tree/master/tokeni...

To install: - Rust: https://crates.io/crates/tokenizers - Python: pip install tokenizers - Node: npm install tokenizers

julien_c··on Apple downloads ~45 TB of models per day from our S3 bucket
I had intended to look into Cloudflare before but didn't get a chance. Looks like it's a good time now!
julien_c··on Natural Language Processing: The Age of Transformers
Also https://transformer.huggingface.co

Disclaimer: built it.

julien_c··on MacBook Pro Keyboard Drives Me Crazy
Yes it is
julien_c··on TypeScript 3.6
Yes, Python's type system is nowhere near Typescript, and are closer to docblock annotations.
julien_c··on Microsoft is investing $1B in OpenAI
We should build one ourselves, actually. That's a terrific idea.
julien_c··on Lyft Takes over Ford GoBike
Very nicely said.
julien_c··on PyTorch adds new dev tools
The `nn.BatchNorm` CPU inference speedup is a Big Deal™ to us.

Thanks for the contributions everyone, and glad to have bet on PyTorch a few years back.

julien_c··on Single-dose propranolol tied to ‘selective erasure’ of anxiety disorders
NYT article from 2016: A Drug to Cure Fear https://www.nytimes.com/2016/01/24/opinion/sunday/a-drug-to-...
julien_c··on Single-dose propranolol tied to ‘selective erasure’ of anxiety disorders
Resting rate around 55 for someone who exercises? I don't think I would this bradycardia.
julien_c··on Single-dose propranolol tied to ‘selective erasure’ of anxiety disorders
Thanks for the pointer. I can access it here in the US.
julien_c··on A family tracking app was leaking real-time location data
Are Techcrunch writers journalists?
julien_c··on Ask HN: Is anyone still using Coffeescript?
No, but it paved the way, in many respects, to those newer approaches – and Ecmascript itself evolving.
julien_c··on Plastic Bags to Be Banned in New York
Combined with the crazy wind here it also ends up everywhere (like at the top of tall trees)
julien_c··on Plastic Bags to Be Banned in New York
You can sell reusable plastic bags so if absolutely necessary you can always buy them.
julien_c··on Facebook 'mistakenly deleted' years of Mark Zuckerberg's old Facebook posts
I don't understand why people delete their past Facebook posts. considering it's highly unlikely they are actually deleted in the backend (and not just flagged)
julien_c··on Was MongoDB Ever the Right Choice?
Very timely Show HN post: we just released Mongoku, a neat Web-based interface: https://news.ycombinator.com/item?id=19500859
julien_c··on Show HN: Mongoku, a Web-Scale GUI for MongoDB
It scales with your data (at Hugging Face we use it on a 1TB+ cluster) and is blazing fast for all operations, including sort/skip/limit. Built on TypeScript/Node.js/Angular.

[disclaimer: made it]

Page 1 of 10Next →