323 karma · joined July 23, 2013
Still, I find it interesting. If you can't synthetically alter someone's performance to be "worse", is it OK that the NFL synthetically altered Alicia Key's performance to be "better"?
For a more consequential example, imagine Biden's marketing team "cleaning up" his speech after he has mumbled or trailed off a word, misleading the US public during an election year. Should that be disclosed?
This seems oddly specific to the inverse of what happened recently with Alicia Keys from the recent Superbowl. As Robert Komaniecki pointed out on X [1], Alicia Keys hit a "sour note" which was silently edited by the NFL to fix it.
[1] https://twitter.com/Komaniecki_R/status/1757074365102084464
In my opinion, these kinds of apps are the future of data science / data analyst work. Forget no-code, just enable these professionals to work in a single programming language that they're familar with and give them visualization superpowers. The Python ecosystem has https://www.streamlit.io/ and https://gradio.app/ now. R has https://shiny.rstudio.com/. I think we'll see more.
[1] https://ai.facebook.com/research/publications/bart-denoising...
Actually, now that I think about it... this might be a good source for bootstrapping further training data.
[1] https://en.wikipedia.org/wiki/Flesch%E2%80%93Kincaid_readabi...
I think the most interesting part of this problem is the ambiguity in desired result. Emails are inherently unstructured text data, which results in conflict about how people want them to look like.
I'd love to test it on more marketing emails, or content marketing -- these seem to be the two applications with the most direct ROI.
I do have a generic popup that you can to summarize any passage, so you could copy-paste some long passage (email, article) to piggy back on the functionality. But it isn't tightly integrated into the Gmail UI as of now.
For the design of this tool, it was important for me to "not get in the way" if you're already writing a short email. That's why I designed the widget so that it will only appear in your gmail window if you go over 100 words. It can serve solely as a "reminder" for the rare cases you might go over... almost like a training to help you form short email habits over time.
If you're talking about https://www.gkogan.co/blog/increase-reply-rates/, this was a primary inspiration for this project -- I even stole the name ;)
Good to meet you -- I'm a big fan of DVC. In our implementation, we've taken the approach of conforming to the standards (typically open-source) set by the frameworks for serialization (for example https://www.tensorflow.org/guide/saved_model). In our API design, it was important to integrate at the framework level (e.g. a tf.keras.models.Model object) for our client libraries. If you're using one of the widely available frameworks that we support, this results in a simple API where the serialization / deserialization is more of an implementation detail. If you're using a custom or rarer ML framework with an unstandardized serialization format, an open-source approach might work better.
Hope that was helpful!
We do support deploying to a private AWS cloud in our private beta that we can help get you set up with. Azure is not yet supported for private deployments.
1) GPU acceleration and AWS VPC deployments are only available in the private beta (free tier is hosted on our private infrastructure). Apply here and we can set up a meeting with you asap: https://modelzoo.typeform.com/to/Y8U9Lw.
2) We've started with TensorFlow and Hugging Face Transformers. We're currently working on scikit-learn and PyTorch support (via https://github.com/PyTorchLightning/pytorch-lightning). What kind of frameworks do you use?
I've worked as a software engineer in machine learning for a few years, bringing models to production in a variety of areas. Although there exists an open-source ecosystem for deploying ML models, most of these tools are targeted towards infrastructure engineers -- Kubernetes, Docker, and web server frameworks. As a result, there exists a gap today between some of the data scientists and machine learning engineers that develop these models and the skills required to deploy them.
I built Model Zoo to address that gap. Deploy your model to an HTTP endpoint with a single line of code, from any Python environment. Plus, you get all the features you'll need from a production ML system for free (monitoring features / predictions, autoscaling, web interface for documentation).
Test it out with one of our quickstarts here for free. You can experiment with it in-browser via Google Colaboratory or in your own Python environment:
https://docs.modelzoo.dev/quickstart/tensorflow.html or https://docs.modelzoo.dev/quickstart/transformers.html
[1] https://seattle.curbed.com/2017/8/15/16153622/ofo-bike-share...
Definitely agree with you that the result of these experiments and others similar to it so far have produced unstructured and unenjoyable music, when compared to human-level compositions. I also agree that you could achieve similar results by some rule-based system involving random walks through a musical scale, common chord progressions, etc.
However, to me what's exciting about this and similar projects is that the network is learning these rules on it's own. The whole point of machine learning is that we don't have to explicitly state or even understand the underlying structure or music theory, because the algorithm figures it out on it's own. And sure, right now, it is only learning structure that we can codify manually in rules. But who's to say that as neural network techniques become more advanced, they won't be able to learn more abstract concepts such as the overarching dramatic structure of a musical piece?