HNHacker News
TopNewBestAskShowJobs

ishcheklein

200 karma · joined March 7, 2018

Building DVC.org, CML.dev, and other ML workflow tools.
submissionscomments
ishcheklein··on Is Google an Evil Corporation?
Can we apply evil/good terminology to corporations at all?

Btw, today on the HN front page - https://news.ycombinator.com/item?id=24105465

"There is no question that Google has a monopoly on online search. Any reasonable person knows this is a fact. The three vital questions I hope to help you answer today are: 1) whether Google has used its overwhelming market power as a monopoly to benefit itself while crushing competition, 2) whether these actions have had negative impacts on the open internet as a public good and 3) whether these actions have created harm for everyday internet consumers. The answer to all three questions is an emphatic YES. "

ishcheklein··on Launch HN: Datafold (YC S20) – Diff Tool for SQL Databases
Thanks! And how and where does setup happens which database to use to run the query for the specific SQL file? Also if it's part of some pipeline will it have to run the whole pipeline from the very beginning?
ishcheklein··on Launch HN: Datafold (YC S20) – Diff Tool for SQL Databases
Hey! Looks great! Is there an example of the Github integration - how does it looks like?

I'm one of the developers and maintainer of the DVC project and we recently released CML.dev- which integrates with Github and can be used to run some checks on data as well. But in our case it's about analyzing files more or less. I'm curious how does that integration look like in your case.

ishcheklein··on We Failed with ProductHunt Launch
"First of all no matter what you launch, if you hope that you simply put your submission out there, and it will end up in the top 5 of the day, most certainly this is not going to happen. If you wanna get there, you have to come up with a plan to promote your listing and get the upvotes." - so, it looks like that was (at least partially) wrong.

Why would you think that it's not possible to launch it more or less organically? It's definitely possible here on HN, why would PH should be very different?

At the very end - don't focus on promoting that much, especially artificially, build the product people love and you get your front page.

ishcheklein··on Show HN: We're building a virtual meetup platform
Congrats! Looks like a pretty crowded space already (can be a good sign - problem is not solved yet). Curious how do you differentiate yourself from other players in this space - tulula, Airmeet, run the world?

We've been trying different tools recently for our needs (we are building opens-source tools for ML - DVC.org, CML.dev and naturally need a solution for online meetups). I liked tulula and I liked a regular Zoom for smaller events where anyone can stream a video, talk if needed, etc.

In your mind what would be a killer feature to compete with Zoom long term?

ishcheklein··on One year of automatic DB migrations from Git
Looks interesting - GitOps for databases. Can I rollback to a specific version?

I think Rails has a combination - migration files + schema file that serves a source of truth https://edgeguides.rubyonrails.org/active_record_migrations..... It solves a problem of a cold start, but I'm not sure it can generate a migration from the changes to it.

ishcheklein··on CLI Design (2013)
I'm one of the maintainers of the DVC.org tool and was even trying to find a person CLI UI/UX engineer to get UI to the next level. Have someone heard about people specializing in this? It should someone who just feel pain when UI is not good, someone with great empathy to end users.

Any other great resources on this topic?

ishcheklein··on Show HN: Painless tab-completion script generator for Python applications
We've made a painless tab-completion script generator for Python applications! It's called shtab and it currently works with argparse, docopt, and argopt to produce bash and zsh completion scripts. This tool was originally created to help dvc, but we realised it could be made more generic and valuable to the world's entire ecosystem of Python CLI applications. Find out how to take advantage of it in this blog post.
ishcheklein··on Open-Source Version Control System for Machine Learning Projects
Thanks, documentation part warms my heart - it's been quite important part day zero.
ishcheklein··on Open-Source Version Control System for Machine Learning Projects
Hi! Great question. It started as a pet project, but now is being supported and belongs to a VC funded entity - Iterative. + it has independent contributors, it's an Apache 2.0 project after all. In terms of business long term goals, I would you can compare us to Hashicorp. There are no plans to monetize or even build premium features in DVC or CML.dev project, rather we'll be building enterprise layer on top of them. I really hope that clears out your concerns. If not - I would be happy to answer any question.
ishcheklein··on Open-Source Version Control System for Machine Learning Projects
Hey! DVC maintainer here. I would say they overlap to some extent, but mostly complement each other for now. And that's what we see across our user base. MlFlow is used mostly as a ML logger - you insert some Python code into your scripts and log to MlFlow everything related to an experiment. While DVC is strong is managing and versioning data and models with a Git-like experience.
ishcheklein··on Reflecting on a year of making machine learning useful
This post resonates (and mentions) with the Software 2.0 concept introduced by Andrej Karpathy, I guess key point is that data (slicing and dicing, and creating good tools to mange that process) can be more important than the modeling itself.

Also, a few points on interpretability, importance of reproducibility and a few thoughts on how ML tools should be built.

ishcheklein··on Post-Commit Reviews
It feels to me that it's pretty much the same. You still have to do review at some point, means engineer will be distracted, and probably it will require immediate attention since it's already merged, other things can depend on it, reverting is an high velocity team can be a challenge on its own. Not saying that it's bad per se, it just feels that you get the same amount of work and "distraction" anyway.
ishcheklein··on Team Improvement Techniques
I've seen different teams and sometimes for some reason it's quite challenging to get folks participate in retrospectives in a meaningful way. E.g. it's hard to come up with a few things that went bad for engineers. And it's not a matter of trust, it's something else. Does it come with practice? What are the techniques to get people involved?
ishcheklein··on Launch HN: Openbase (YC S20) – reviews and insights for open-source packages
Do you have a plan to expand to other projects, not related to JS?
ishcheklein··on Show HN: Deploy Scikit and Keras Models with a Simple Drag and Drop
Looks interesting! How about models that require dictionaries - e.g. tf-idf to convert text into a feature vector? Does it allow for some preprocessing?
ishcheklein··on Robinhood raises $320M more, bringing latest round to $600M at $8.6B valuation
Meanwhile - “They make it so easy for people that don’t know anything about stocks,” he said. “Then you go there and you start to lose money.”

https://www.nytimes.com/2020/07/08/technology/robinhood-risk...

I wonder at what stage regulators come and start destroying the simplicity? Like Zoom initially unsecure but user-friendly has recently become an enterprise-grade confusing-to-use tool.

ishcheklein··on Launch HN: Aquarium (YC S20) – Improve Your ML Dataset Quality
Hey! DVC maintainer and co-founder here. First of all, congrats and let me know if we can help you or you have some collaboration in mind! A few questions - how does workflow look like - do you expect users to upload all data to your service? How can data then be consumed from the platform?
ishcheklein··on Show HN: Tiny password manager with all data stored encrypted on your machine
Nice! I've been using https://pwsafe.org/ + Dropbox for quite a while. Is it similar, what are the new features?
ishcheklein··on Software 2.0 (2017)
I think it's orthogonal question, right? Even if there were some mistakes, how many more death there would have been if they were not thinking about rigorous process behind ML development?
ishcheklein··on Software 2.0 (2017)
Ideas discussed in his post might seem too far from the reality and controversial, but first don't forget they have been doing Tesla's self-driving ML models for a few years now - so, he definitely has some material to generalize. It's one of the most advanced ML models in production, and without reflecting on the process it would be hard to develop it, would be hard to maintain it, etc.

Also, from my own experience building DVC - when you do any ML project you do have code indeed, but not doubt that data can be considered as important element as code, we need to take it seriously - track, review, etc, etc.

ishcheklein··on Show HN: Continuous Machine Learning – CI/CD for Machine Learning Projects
I think DVC (+CML) is a good solution for this. It "wraps" artifacts that you store into Git. And Git repo abstracts access to the cloud. In your case of the mounted storage it will look like `git pull` + `dvc checkout` after model is "merged" in the production/master branch.

CML can automate and make the process of preparing the model to be merged into that branch reliable, visible, robust, etc.

I'm happy to help with this flow, ping me on Twitter - @shcheklein in DM or ivan on DVC Discord.

ishcheklein··on Show HN: Continuous Machine Learning – CI/CD for Machine Learning Projects
Thanks! we'll check it out and be in touch.
ishcheklein··on Show HN: Continuous Machine Learning – CI/CD for Machine Learning Projects
Agreed! I think, it's similar to software engineering. It's not that often that large websites are being deployed to prod without someone approving/QA-ing it manually first, right?

CI/CD systems in this case help automating this as much as possible, but do not completely replace decision making process, I would say.

ishcheklein··on Show HN: Continuous Machine Learning – CI/CD for Machine Learning Projects
Hey Sytse! Thanks for the kind words. We're very interested in a deep integration with Gitlab :) Can you share some examples of these security scan reports, please? Right now, we return CML reports as comments in Merge Requests like this: https://gitlab.com/iterative.ai/cml-cloud-case/-/merge_reque.... We'd appreciate any tips or suggestions.
ishcheklein··on Show HN: Continuous Machine Learning – CI/CD for Machine Learning Projects
Hey! Disclaimer - I'm one of the DVC maintainers :) Super excited for the team on this release!

For the last two years we have seen over and over again how our users take DVC and use it inside Gitlab, Github, etc. This product was born partially as a result of these discussions, partially as an initial visions for the ML tools ecosystem - Hashicorp-like.

Having A software engineering background I really hope that integrating ML workflow into engineering tools will be the future of this space. And with CML and other tools (e.g. https://github.blog/2020-06-17-using-github-actions-for-mlop...) we see this happening.

ishcheklein··on Why Can't I Reproduce Their Results?
Hey, DVC maintainer here. For those who interested in this topic, I like this one about the same problem - https://petewarden.com/2018/03/19/the-machine-learning-repro... (industry focused) and an excellent talk from Patrick Ball - https://www.youtube.com/watch?v=ZSunU9GQdcI&t=1s how they structure data projects.
ishcheklein··on Pkg.jl telemetry should be opt-in
It helps developing and prioritizing features faster. What is so harmful about it? Assuming it's anonymized properly, if no one resells it, if it's explicit (doesn't matter opt-in or opt-out).
ishcheklein··on Python Utility ISort 5.0 Release
Congrats! We've been happy using it for our python projects.
ishcheklein··on Rich Tables in the Terminal
This looks awesome! Gonna try to use it for DVC where we need to show tables of experiments, metrics, etc.
← PreviousPage 2 of 3Next →