HNHacker News
TopNewBestAskShowJobs

fchollet

376 karma · joined November 14, 2024

submissionscomments
fchollet··on Francois Chollet is leaving Google
Hi HN, Francois here. Happy to answer any questions!

Here's a start --

"Did you get poached by Anthropic/etc": No, I am starting a new company with a friend. We will announce more about it in due time!

"Who uses Keras in production": Off the top of my head the current list includes Midjourney, YouTube, Waymo, Google across many products (even Ads started moving to Keras recently!), Netflix, Spotify, Snap, GrubHub, Square/Block, X/Twitter, and many non-tech companies like United, JPM, Orange, Walmart, etc. In total Keras has ~2M developers and powers ML at many companies big and small. This isn't all TF -- many of our users have started running Keras on JAX or PyTorch.

"Why did you decide to merge Keras into TensorFlow in 2019": I didn't! The decision was made in 2018 by the TF leads -- I was a L5 IC at the time and that was an L8 decision. The TF team was huge at the time, 50+ people, while Keras was just me and the open-source community. In retrospect I think Keras would have been better off as an independent multi-backend framework -- but that would have required me quitting Google back then. Making Keras multi-backend again in 2023 has been one of my favorite projects to work on, both from the engineering & architecture side of things but also because the product is truly great (also, I love JAX)!

fchollet··on Facebook's seized files published by MPs
If I were to divide large tech companies between "good" and "bad", the only one I would classify as "bad" is Facebook. Most are neutral-to-good. Facebook has a huge societal and human cost and has hardly any benefits to show for it (unlike, say, fossil energy companies, which are a threat to civilization but at least serve a useful role by fulfilling our energy needs).

Facebook is a democracy-threatening, attention-wasting, extractive institution, with a deeply unethical leadership. Most tech companies are neutral -- Amazon, Microsoft, Uber etc., I assume they're doing business for profit and not social good, but they do provide value and largely play by the rules. Some other companies I would classify as "good" as they provide considerable value to society as a by-product of doing business, like Apple and Google (disclaimer: I work for Google, and I am quite happy about that).

Facebook is the only large tech company I can think of that is just plain evil. There's no other like it. It's in a league of its own.

fchollet··on Artificial Intelligence Still Isn’t All That Smart
> there isn't anyone living today that has any idea how to get to AGI

Not true.

fchollet··on [dead]
When someone shows you who they are, believe them the first time.
fchollet··on User experience design for APIs
> IMO, AssertionErrors should indicate bugs in the called API. ValueErrors indicate bugs in the caller's use of the API.

Yes, I agree with this stance. `ValueError` should be used for user-provided input validation (as well as `TypeError` in some cases). But as it happens, many Python developers use `assert` statements to do input validation, and generally don't provide any error messages in their `assert` statements. I'm suggesting going with `ValueError` (and a nice message) instead.

fchollet··on TensorFlow Feature Columns
To clarify: `tf.keras` is an implementation of the entire Keras API written from the ground-up in pure TensorFlow. The first benefit of that is a greater level of blending between non-Keras-TF workflows and TF-Keras workflows: for instance, layers from `tf.layers` and `tf.keras` are interchangeable in all use cases.

Additionally, this enables us to add TensorFlow-specific features that would be difficult to add to the multi-backend version of Keras, and to do performance optimizations that would otherwise be impossible.

Such features include support for Eager mode (dynamic graphs), support for TensorFlow Estimators, which enables distributed training and training on TPUs, and more to come.

fchollet··on Understanding Hinton’s Capsule Networks, Part I: Intuition
To play with Capsule Networks in practice, you can try this simple Keras implementation: https://github.com/XifengGuo/CapsNet-Keras
fchollet··on Tensorflow sucks
But that is precisely how you should be using Keras!

* If you are implementing a standard model (that's 90% of industry use cases, and a large fraction of research use cases as well), Keras primitives considerably simplify your workflow and make you a lot more productive.

* When you need to implement something highly customized or unusual, you can revert back to writing pure TensorFlow code, which will integrate seamlessly with your Keras workflow (via custom layers, functions etc).

Basically, Keras increases your productivity for common use cases, without any flexibility cost for rare/custom use cases. It is meant to be used together with TF, not as a replacement for TF.

fchollet··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
CoreML supports Keras but not TensorFlow because Keras models form a well-structured subset of all possible TensorFlow graphs. It would be quite difficult to support completely arbitrary TensorFlow graphs, but supporting every Keras layer is relatively straightforward.

To answer your question: I had no knowledge of this ONIX project before the public announcement today. Speaking purely for myself, if I wanted to develop a universal model exchange format, the first step I would take would be to get in touch with the makers of the frameworks that sum to 80-90% of the market share -- TF, Keras, MXNet. But maybe such a strategy was thought to be superfluous in this case -- for instance, because ONIX may not actually be intended as a universal model exchange format.

fchollet··on Show HN: End-to-end deep learning experimentation platform for Tensorflow
It was always possible to train Keras models in a distributed setting (I was doing it in late 2015). And there's built-in, one-line integration with the Estimator API coming in the next version of TensorFlow.
fchollet··on The Limitations of Deep Learning
Hardly a "made-up" conclusion -- just a teaser for the next post, which deals with how we can achieve "extreme generalization" via abstraction and reasoning, and how we can concretely implement those in machine learning models.
fchollet··on Certigrad: bug-free machine learning on stochastic computation graphs
How is that the case in this specific example? It looks a lot harder to check for correctness.
fchollet··on AIGrant: Get $5,000 for your open source AI project
Great initiative! Always good to see more support for open-source and for AI tooling.
fchollet··on Introducing Keras 2
See the release notes: https://github.com/fchollet/keras/wiki/Keras-2.0-release-not...

- You can use a Keras model to compute some tensor(s), turn that into a loss, and manually add that loss to the model via `add_loss` (it just needs to only depend on the model's inputs).

- Not all of your model outputs have to have a loss associated with them. So you can do both {input, output}->loss and input->output in your workflow, as you wish. Effectively, losses and outputs are decoupled.

The VAE example hasn't yet been updated to use the `add_loss` feature, but it should be.

fchollet··on Introducing Keras 2
You can try opening an issue on Github. `val_acc` is definitely still accessible by callbacks, and the `EarlyStopping` callback, which relies on it, is fully unit-tested.
fchollet··on Introducing Keras 2
With the functional API of Keras, it would definitely make sense. In fact I do think that imperative model definition would be great to have at some point in the future. We'll see :)
fchollet··on Introducing Keras 2
To put this quote in context: this isn't specifically about PyTorch. Every couple of months since mid-2015, a new deep learning framework gets released. In the following week, someone inevitably asks "will X get added as a Keras backend?".

Supporting several backends is a strong positive. But chasing every new framework as a backend is a quick way to kill Keras, via bloat, support issues and general technical debt. We should only support a backend that is considered mature, and we should stay away from the hype surrounding the release of every new framework. There will be another hyped up framework next quarter anyway. And the one after.

It is in fact possible that Keras will eventually support PyTorch. But if it ever happens, it would be at least 1-2 years in the future. When PyTorch becomes "uncool", just like Keras :)

fchollet··on Using Keras and Deep Deterministic Policy Gradient to play TORCS
It would be easy to replace the input features with representations from a pixel-level convnet, to make this is a "real" self-driving car, going from pixels to commands.

Anyone interested in this type of research: consider cloning the repo and implementing this modification, it would make a great starter project.

fchollet··on Elon Musk on How to Build the Future
The equation should have per-subject weights to modulate the "how many people" part. For instance, if my app helps Mr. Musk be 10% more productive, that's more impactful than helping Mr. Smith be 10% more productive.
fchollet··on Show HN: Set of trained deep learning models for computer vision
Yes, you can use these models for fine-tuning (or feature extraction) on a new dataset. This tutorial would be a good place to start (esp. sections 2 and 3): https://blog.keras.io/building-powerful-image-classification...
fchollet··on Show HN: Set of trained deep learning models for computer vision
The code is under the MIT license, not the weights. The weights are under their respective licenses.

The weights are not included in the git tree and are thus not covered by the LICENSE file. They are automatically downloaded when you run the code.

EDIT: following your comment, I have added a point-by-point breakdown of licensing information in the README. This will avoid any confusion.

fchollet··on DeepMath – Deep Sequence Models for Premise Selection
Information leaks are indeed the number one challenge when evaluating this type of approach. Our evaluation procedure does not rigorously guarantee against information leaks, but it does allow for an apples-to-apples comparison with previous methods (which had the same potential issue). And showing an advantage over these methods is what we were trying to achieve.

So in summary: it's hard to prevent, but we're doing our best.

fchollet··on DeepMath – Deep Sequence Models for Premise Selection
- the most interesting thing about this approach is that it managed to prove a significantly different set of theorems than previous state-of-the-art approaches. So theorems that are hard to prove for this approach were not all hard to prove for previous approaches, and inversely this approach can prove theorems that were hard for previous approaches.

- in this specific setup, yes, all sequences are truncated for practical reasons. In theory there should be no size limit, however long sequences will of course be more difficult to meaningfully encode.

- formal software verification is definitely one of the long-terms goals of this project. I think software verification is one of the big challenges of our transition into an information society (as algorithms/AI start having more control over our lives, as we start using smart contracts, etc), and AI will help solve it. The current setup would not be of much help, however.

fchollet··on DeepMath – Deep Sequence Models for Premise Selection
One of the authors here; didn't expect to see this pop up on HN. Questions welcome!
fchollet··on Spain Runs Out of Workers with Almost 5M Unemployed
Yes. Specifically, rents are much higher in SF/SV than in NYC. SF has even higher rent than Manhattan. Most of NYC (not Manhattan or Williamsburg) can be quite reasonable.
fchollet··on Spain Runs Out of Workers with Almost 5M Unemployed
Paris is much less expensive than the bay area, but more expensive than the average US city. In fact, it's barely less expensive than NYC, one of the most expensive US cities.

Also, if you are working remotely you cannot command a bay area salary.

fchollet··on The effects of living in a poor neighborhood
> the quirky and random XD reaction GIFs and image macros that populate modern thought pieces

I'm just glad that we don't read the same "thought pieces"

fchollet··on The state of deep learning in Debian
In fact, "Brainstorm" is one of the least used packages among all deep learning libraries. For reference, here is the number of new Github issue tickets created over the past 15 days (2016-05-15 to 2016-05-30):

#1: 131 tensorflow/tensorflow

#2: 107 fchollet/keras

#3: 91 dmlc/mxnet

#4: 64 BVLC/caffe

#5: 61 Microsoft/CNTK

#6: 42 deeplearning4j/deeplearning4j

#7: 26 tflearn/tflearn

#8: 23 Theano/Theano

#9: 16 NVIDIA/DIGITS

#10: 10 pfnet/chainer

#11: 9 Lasagne/Lasagne

#12: 8 torch/torch7

#13: 8 NervanaSystems/neon

#14: 6 mila-udem/blocks

#15: 1 autumnai/leaf

#16: 0 karpathy/convnetjs

#17: 0 tensorflow/skflow

#18: 0 IDSIA/brainstorm

fchollet··on Is the online advertising bubble finally starting to pop?
Stable year-on-year growth for 15 years is not the sign of a bubble, it's just what you see with a growing industry.
fchollet··on Is the online advertising bubble finally starting to pop?
> I started calling online advertising a bubble in 2008.

Then maybe it's time to quit punditry; that was 8 years ago and the online advertising market has grown 2.5x since.

> I made “The Advertising Bubble” a chapter in The Intention Economy in 2012.

The online advertising market has grown 30% since.

Since no market can grow forever, at some point there will be a dip in the industry. Maybe in less than five years, maybe in ten years. And when it finally happens, this kind of pundit will be telling us "they had called it correctly, before anyone else".

← PreviousPage 2 of 13Next →