271 karma · joined January 7, 2013
Open source always depended on a viable business model (of one or many companies) that can sustain not just the release, but also an ongoing maintenance of the standard.
There is plenty of AI extensions, but the experience matters. The depth of integration matters. When you execute queries against production warehouses and you make decisions based on the results of AI-generated code, accuracy matters. We had our first demo of an AI agent running in 2 days, it took us another 2 years to build the infrastructure to test it, monitor it, and integrate it into the existing data source.
You'd be surprised how many people collaborate together. Software engineering is solitary, collaboration happens in GitHub. But data analysis is collaborative. We frequently have 300+ people looking at the same notebook at the same time.
.py never worked for data exploration. You need to mix code, text, charts, interactive elements. And then you need to add metadata: comments, references to integrations, auth secrets. There are notebooks that are several pages long with 0 code. We are building a computational medium of the future and that goes beyond a plaintext file, no matter how much we love the simplicity of a plaintext file.
Didn't expect to see this trending here! We worked hard to execute on our vision of a data notebook and I'm glad we finally got a chance to open source it. We stand on the shoulders of giants. AMA!
But keep in mind that US market is unique, as it's extremely overpriced with very poor quality/coverage. Looks like average price in the US is $6.00/GB. Compare this to countries like Israel ($0.02/GB) or Colombia ($0.20/GB). Whenever I travel, I usually have a better coverage and faster connection in a jungle than downtown SF.
Looking at the data, you need to scroll past 9 US cities (all big cities) until you find the first European city (relatively small city you never have a reason to go to) in this crime ranking: https://www.numbeo.com/crime/rankings.jsp
Europe is definitely safer than the US.
When ES6 came out but wasn't yet widely supported it became pretty common to add a build step to transpile the backend code to. This brought a whole bunch of issues that transpilation brought with it but was seen as a temporary evil worth the tradeoff because the new language features brought so many productivity improvements, but stuck for so long it became pretty standard. And when TypeScript became popular, it felt like we'll just never get rid of this complexity explosion that build systems bring. And then Deno came.
Just imagine the amount of config files you'll be able to delete once you no longer need to build your backend codebase.
Is there some hidden complexity or is it just a consequence of engineers building a product for other engineers? Also, any tips what worked for you?
Notebooks are a direct evolution of REPLs and at the time of writing (Knuth published his paper on literate programming in 1984), REPLs have been well known (they had been around since 1970s). Yet there is not a single mention of them in the original paper, nor any prediction how literate-programming-capable notebooks could look like. At the time, these were two different concepts each serving its own need (documentation vs exploratory programming).
We used both ELK and Datadog before and over time I somehow stopped questioning why the logs are so slow. I thought we are simply hitting the theoretical limits of how fast searching in logs can be and unless there is a significant leap in hardware this is what we have to deal with. Then we tried Logtail and now we are migrating everything we can.
Huge part of it is simply the UX. There's a wide range of what kind of work a data scientist does. Some train models that go into production, some analyze the datasets and build reports. Probably best to try both products with your workload and see what works better.
We built Deepnote so that the work you do as a data scientist can be shared with both engineers and non-technical folks. We're not really an mlops platform. We make a really good notebook that integrates with other platforms.
But I'd like to improve on this experience. There are many ways how to do it (great job btw), but we want to explore how a versioning system native to notebooks would look like. We're still iterating on that.