HNHacker News
TopNewBestAskShowJobs

niclane7

101 karma · joined April 30, 2021

submissionscomments
niclane7··on Red queen hypothesis – A new way forward for self-improving AI
Yes, this is key
niclane7··on Red queen hypothesis – A new way forward for self-improving AI
true. And all of us are finding RSI is a great lever to explore.
niclane7··on Ask HN: Who is hiring? (January 2024)
Flower - https://flower.dev/ | Multiple Positions | Remote-first, also Germany and UK | Full-time | REMOTE, VISA

Flower is an open-source framework, ecosystem and community for training and using AI on distributed data with federated learning, and related decentralized technologies. Companies like Banking Circle, Nokia, Samsung, Capgemini, Porsche, and Brave use Flower to easily improve their AI models on sensitive data that is distributed across organizational silos or user devices. Almost all AI today is based on centralized public data — a small fraction of the data we have; we believe that training on orders of magnitude more data will unlock the next leaps in AI.

We are backed by Y-combinator, and prominent venture capital firms and angel investors including, First Spark Ventures, Factorial Capital, Hugging Face CEO Clem Delangue, Betaworks, and Pioneer Fund.

Open Positions (see https://careers.flower.dev/ for most recent postings and information)

- Senior Backend Python Engineer https://careers.flower.dev/o/python-engineer

- Senior Android Engineer https://careers.flower.dev/o/android-engineer

- Research Scientist (all seniority levels) https://careers.flower.dev/o/research-scientist

- Solutions Engineer (all seniority levels) https://careers.flower.dev/o/solutions-engineer

- Head of Community and Content (senior ICs also considered) https://careers.flower.dev/o/head-community-content

- Head of Business Operations https://careers.flower.dev/o/head-biz-ops

niclane7··on Launch HN: Flower (YC W23) – Train AI models on distributed or sensitive data
It is reasonable to think of it that way. Certainly high-level information from the data is extracted and embedded within a model, but only the information necessary for the model being trained. Whereas if the data itself was being sent, then all of the information is available. Additionally, through added protections (differential privacy being one) it is possible to engineer the federated system such that the data itself can not be reconstructed from model itself.
niclane7··on Launch HN: Flower (YC W23) – Train AI models on distributed or sensitive data
Flower has an agreement to develop interoperable components with OpenFL. This is part of the broader plan by Intel to work with a consortium of players (that includes Flower Labs) and have the output code sit with the Linux Foundation. Enabling TEE support within OpenFL for SA assessible to Flower users is precisely the type of opportunities we seek to make possible by working with Intel on this.

This is the official press release for those who are interesed: https://www.intel.com/content/www/us/en/newsroom/news/transi...

More broadly, in regards too your comment -- our current SA support does not require hardware support, which is what we targeted first, so that can be broadly adopted in many potential hosts of FL aggregation servers. It is suitable for most applications in need of privacy, although still requires certain assumptions to be met such as the number of nodes within a round, and other factors.

niclane7··on Launch HN: Flower (YC W23) – Train AI models on distributed or sensitive data
Thanks for the question, very natural to ask. We are also fans of PySft. It offers support for a very wide range of privacy enhancing machine learning tools. But where Flower and PySft differ is in focus. Federated learning is difficult and requires many technical moving parts all working together (e.g., secure aggregation, differential privacy, scalable simulation, device deployments, integration with conventional ML frameworks etc.). All of these need to tightly integrated, and in a manner that performs federated learning efficiently. This is where Flower currently excels. It offers comprehensive, extensible and, most important, easy to use construction of federations that need these different parts together. We believe it offers the best user experience for federated learning currently out there. We hope in the future many tool suites that offer private machine learning (like PySft and others) will actually adopt Flower components so we can all work better together.
niclane7··on Launch HN: Flower (YC W23) – Train AI models on distributed or sensitive data
Yes, this would be exciting to see. One approach wouldn't require federated learning however. If you had direct access to the data then you could build a conventionally trained large language model (i.e., collect all the data together placed in a data center). However, given the context of this discussion -- you are probably asking about if we could use Flower to train in a federated manner. I believe so. Although again, we'd probably be training a LLM which brings added complications due to its size (and other factors). Internally at Flower we have been testing methods to overcome this and are confident we can pull this off. One could imagine someone hosting a pre-trained LLM and contributing institutions acting as nodes in the network, each performing some small part of the training based on the fraction of the literature they have access to. We plan to release LLM based federated technology in the coming months.

For those that are interested: The best work currently I've seen on training very large models under federated learning, that also makes very realistic assumptions about the likely underlying participating hardware, is this: https://arxiv.org/abs/2206.11239 -- although I expect more in this direction to come soon.

niclane7··on Launch HN: Flower (YC W23) – Train AI models on distributed or sensitive data
Regarding federated labeling, you might be interested in some recent prototypes built on Flower that use forms of self supervised learning. By combining SSL with federated learning we can start to leverage unlabeled data and this will be a big deal once it becomes common place. I'd suggest looking at these two research papers that build on Flower and include members of the Flower team as authors:

https://arxiv.org/abs/2207.01975

https://arxiv.org/abs/2204.02804

niclane7··on Launch HN: Flower (YC W23) – Train AI models on distributed or sensitive data
Yes, we have developed modular and efficient secure aggregation and differential privacy solutions that can help people dial in the amount of protection they need. We have documented an early version of the secure aggregation here: https://flower.dev/docs/secagg.html Documentation and updates on both methods will be released soon.
niclane7··on Flower – A Friendly Federated Learning Framework
This is cool. Nice graphic. I like this event they are talking about on the site. I plan to go.