511 karma · joined October 13, 2011
In large user populations like at Databricks I think the ultimate answer will come from experimentation instead of offline evals. We are already doing this in small groups, exposing them to new candidate models and then measuring per-developer cost and perceived quality changes.
It is true that this problem can be mostly managed by the techniques we mention here. Those are actually pretty difficult to set up at scale, so many companies (including us) we only really did this in earnest once we started to see those large cost oscillations.
The main reason we shared this here is to maybe help other companies get infrastructure in place before massive cost swings rather than after.
We'll definitely keep iterating on Dolly and releasing everything openly.
Working on a model without this issue. Certainly our goal is totally open models anyone can use for anything.
This uses a fully open source (liberally licensed) model and we also open sourced (liberally licensed) our own training code. However, the uptraining dataset of ~50,000 samples was generated with OpenAI's text-davinci-003 model, and depending on how one interprets their terms, commercial use of the resulting model may violate the OpenAI terms of use. For that reason we are advising only noncommercial use of this model for now.
The next step here is to create a set of uptraining samples that is 100% open. Stay tuned.
https://github.com/databrickslabs/dolly
Sorry it took us a day to get the external repo setup.
DataFrames impose just a bit more structure: we assume that you have a tabular schema, named fields with types, etc. Given this assumption, Spark can optimize a lot of internal execution details, and also provide slicker API's to users. It turns out that a huge fraction of Spark workloads fall into this model, especially since we support complex types and nested structures.
Is the core RDD API going anywhere? Nope - not any time soon. Sometimes it really is necessary to drop into that lower level API. But I do anticipate that within a year or two most Spark applications will let DataFrames do the heavy lifting.
In fact, DataFrames and RDDs are completely inter-operable, either can be converted to the other. This means that even if you don't want to use DataFrames you can benefit from all of the cool input/output capabilities they have, even just to create regular old RDDs.
Online graph algorithms aren't there yet (probably what you mean). We just started adding online MLlib algorithms, so this is the main focus for now.
Re: databricks cloud - shoot me an e-mail and I'll see if I can help. Right now demand exceeds supply for us on accounts, but I can try!
Here's an excerpt: "Come out and suppoert Chickenshed, an inclusive theatre company based in London that brings people of all ages, backgrounds and abilities together to create groundbreaking and exciting new theatre."
Time will tell how the world will change now that this sensitive information is out in the open.
To say that every missed opportunity is a failure is like saying that every CA lottery I didn't buy a ticket for was a mistake. That strategy risks over-fitting to successful fluke's, to people who made bad bets that happened to work out. There is enough loose money floating around today that some people out there are going to make bad bets and succeed.
I don't see why you/YC wouldn't focus on bets where you have an unfair advantage over the market and feel fine passing those up where you don't, even if they have some non-zero probability of success.
At the end of the day you are taking a calculated risk. The interesting question isn't whether you wished you had funded them, because that is asking you to make a risk-free (in hindsight) decision. The real questions is how you decide whether a missed opportunity indicates a lapse in the way you calculated risk/reward, or simply a bet you ended up on the wrong side of despite it being the right bet at the time.
That's the question I'd really be interested in hearing the answer to. How do you decide whether a missed opportunity represents an error in your process?