HNHacker News
TopNewBestAskShowJobs

idomi

19 karma · joined April 23, 2020

submissionscomments
idomi··on Observability for LLM apps with structlog and DuckDB
I think it really depends on your need, if you're working on your local and trying to get something to work, caching, scale etc, might be an overkill.
idomi··on Show HN: Shiny Express – Reactive web framework for data science in Python
Pretty cool, it's about time we got some facelift for Shiny.
idomi··on Ploomber (YC W22) Is Hiring Senior Software Engineers in NYC
Come join us!
idomi··on Evidence.dev – Business Intelligence as Code
Feel free to open an issue if you're looking for new features/something's missing. Also we'd love to chat more about your Azure Data Studio!
idomi··on [dead]
When data scientists are training locally on small datasets, they have to deal with dependency installation, parameterization, and provisioning infrastructure. Once they want to train the model for production on a full load dataset, the complexity increases, and a new configuration must be considered.

Today we're launching a GUI feature in ploomber to solve this issue, by allowing our users to drop notebooks and execute them on the cloud without spinning up clusters or worrying about any infrastructure.

The service is based on our open-source software and have a free-tier that allows to scale multiple models before depleting the quota.

Would love to hear thoughts and impressions of it! You can also reach out to me directly at ido@ploomber.io

idomi··on Who needs MLflow when you have SQLite?
+1 I also think it's faster that way on both environment setup and ad hoc rapid experiments, from my experience using the library in a team doesn't scale well, it becomes pretty slow.
idomi··on Who needs MLflow when you have SQLite?
Pretty interesting. I think this is part of this notion to release half baked products, like some of the stuff in there are really cool, just enough to get you in but it doesn't scale and usually is complex to deploy/use.
idomi··on Who needs MLflow when you have SQLite?
How many data scientists that use Databricks for modeling do you know?
idomi··on Show HN: Ploomber Convert – Convert Jupyter notebooks to PDF, no setup required
Hi, we’re Ido & Eduardo, the founders of Ploomber. We’re launching Ploomber Convert today, a web-based application that allows data scientists to convert notebooks to PDF, no setup required.

As data scientists, we have to share our work with non-technical colleagues to communicate results. To allow them to read our findings, we use nbconvert, which enables us to export notebooks to PDF or HTML. Unfortunately, nbconvert requires Pandoc, TeX/XeLaTeX, Pyppeteer, Chromium, and other packages, which is complicated. Ploomber Convert provides the nbconvert functionality without installing a single package.

Ploomber Convert is built on top of AWS and runs all the necessary packages in a docker container. Since notebooks often contain sensitive information, we do not store any notebooks or PDF files.

Ploomber Convert is free to use. Go to https://convert.ploomber.io, drop your Jupyter Notebook to convert, hit ‘Convert to PDF’, and save it.

We want to make Ploomber Convert the go-to tool for data scientists to turn their notebooks into shareable reports. We’re working on adding support for Quarto, custom CSS templates, export to HTML, and other features. Let us know what else you need!

We’re thrilled to share Ploomber Convert with you! So, if you had difficulties exporting your Jupyter Notebooks or if you have any feedback, please share your thoughts! We love discussing these problems since exchanging ideas sparks exciting discussions and brings our attention to use cases we haven’t considered before!

idomi··on TikTok’s Poison Pill
I doubt it, they got too much traction at this stage.
idomi··on Ask HN: Who do you talk to about system architecture and design?
First you can always research on your own, there's tons of resources online and in git. If that doesn't work I find a teammate or a friend with the right domain expertise. There are also some useful slack communities for instance on MLOps etc.
idomi··on Things I've learned building a modern TUI framework
There are so much stuff you could do with it right?
idomi··on Show HN: Calculator for US individual income tax, from 1970-present
How is this different than turbotax for instance?
idomi··on Show HN: Ploomber Cloud (YC W22) – run notebooks at scale without infrastructure
Thanks for sharing! Yeah definitely a problem, we've been trying to build it mainly on the feedbacks we're getting from the open-source solution.
idomi··on Show HN: Ploomber Cloud (YC W22) – run notebooks at scale without infrastructure
You're right to some extent - at least on some of the concepts. We don't focus on a specific ML use case, like computer vision. In addition, we're oriented towards notebooks and this notion allows us to break it into smaller tasks, cache the results and execute in parallel (locally or via this new cloud service). BTW, we tried talking to the founders but couldn't get a hold of them, if you or anyone know them - we'd love to chat!
idomi··on Show HN: Ploomber Cloud (YC W22) – run notebooks at scale without infrastructure
Yes, you should follow best practices and isolate each job to the smallest task possible (and then reuse components). We have this functionality in 2 flavors, you can define hooks as part of your pipelines(https://docs.ploomber.io/en/latest/api/spec.html#id2), in addition you can define this dependency as part of your jobs DAG (https://docs.ploomber.io/en/latest/get-started/basic-concept...), e.g get the data, clean it, train the model and test it.
idomi··on Lessons learned from running Apache Airflow at scale
Make sure to checkout ploomber, our support is seamless tons of docs (https://docs.ploomber.io/) and we take our users seriously. P.S. We integrate with airflow and other orchestrators if you still need to tackle those.
idomi··on Lessons learned from running Apache Airflow at scale
When it comes to scale and DS work I'd use the ploomber open-source (https://github.com/ploomber/ploomber). It allows an easy transition between dev and production, incrementally building the DAG so you avoid expensive compute time and costs. It's easier to maintain and integrates seamlessly with Airflow, generating the DAGs for you.
idomi··on Show HN: Kestra - Open-Source Airflow Alternative
Ploomber actually can work with shell script so you could export you js code into a pipeline. But besides that not that I know of.
idomi··on Show HN: Kestra - Open-Source Airflow Alternative
Or Ploomber?
idomi··on Launch HN: Elementary (YC W22) – Open-source data observability
Pretty strong accusation, are you sure re-data isn't "inspired" from Monte Carlo? :)
idomi··on Launch HN: Elementary (YC W22) – Open-source data observability
Great job Elementary team! Does this in essence similar to the aws deeque project but fancier and more inclusive of edge cases, common scenarios? (https://github.com/awslabs/deequ)
idomi··on On the Myths and Problems of Jupyter Notebooks
Great piece! I've seen tons of data scientists that move from notebooks into .py scripts, usually it's an org decision. I don't see the point in writing the same code twice.
idomi··on Launch HN: MutableAI (YC W22) – Automatically clean Jupyter notebooks using AI
Not really, data scientists love notebooks, it's pretty inefficient to move out of the notebooks.
idomi··on The Unbundling of Airflow
Well written. I think that airflow is being enforced in organizations as the main orchestrator even though it's not always the right too for the job. In addition, organizations has to enforce a micro-services approach to have modular components. Besides that managing those frameworks is a nightmare. We built Ploomber (https://github.com/ploomber/ploomber) specifically for this reason, modular components and easy deployments. It standardize your pipelines and allows you to deploy seamlessly on Airflow, Argo (Kubernetes), Kubeflow and cloud providers.
idomi··on Launch HN: Ploomber (YC W22) – Quickly deploy data pipelines from Jupyter/VSCode
Thanks for letting us know! So the audio is ok on laptops and mobiles. On the video you're right, it's a bug we need to fix since it's within an iframe you can't expand it into full screen mode.
idomi··on Launch HN: Ploomber (YC W22) – Quickly deploy data pipelines from Jupyter/VSCode
We took an approach of keeping the notebook interface/IDE, and behind the scenes Ploomber converts it into .py files so you can collaborate with teammates through Git. The users can still open those as a notebook and interact with the files regularly.
idomi··on Kedro – Creating reproducible, maintainable and modular data science code
I believe essentially all of the tools mentioned here are focusing on the engineering persona and not the Data Scientist. Writing classes and functions isn't the jargon of a Data Scientist. At Ploomber we tried to put the Data Scientists in the center of everything, helping them to work together with OPS. Check it out! https://github.com/ploomber/ploomber
idomi··on Kedro – Creating reproducible, maintainable and modular data science code
We've built the Ploomber to work with R as well comparing to others who believe only python is for ML. Ploomber is an open source tool that can integrate with R and R studio out of the box. Check it out! https://github.com/ploomber/ploomber
idomi··on Kedro – Creating reproducible, maintainable and modular data science code
We've built the Ploomber open source tool for that exact reason - true open source! We've been trying to focus on the data scientists, not taking them out of jupyter and definitely making sure they can execute what they want without a dedicated infra/ops person. Check it out! https://github.com/ploomber/ploomber
Page 1 of 2Next →