HNHacker News
TopNewBestAskShowJobs

dbecker

1,340 karma · joined December 22, 2011

I help companies use LLM's.

Formerly: Data Scientist at Google, VP Product at DataRobot, ran consulting engagements for 6 companies in Fortune 500. Created `Kaggle Learn`

submissionscomments
dbecker··on Advancements in machine learning for machine learning
> These ML-compilers are being overhyped. It's all the same trade-off as a traditional compiler

Funny you should say that. Because traditional compilers have been incredibly useful.

dbecker··on MGM Seeks Contractors to Repair Infra in 3 Weeks
The people who build this all leave on October 15?

What could possibly go wrong with that?

dbecker··on Falcon 180B
Were you using the base model or the conversational model?

The post says:

The base model has no prompt format. Remember that it’s not a conversational model or trained with instructions, so don’t expect it to generate conversational responses—the pretrained model is a great platform for further finetuning, but you probably shouldn’t driectly use it out of the box.

dbecker··on Agent GPT – Assemble, configure, and deploy autonomous AI Agents in the browser
I've tried AutoGPT agents about 10 times for tasks that seem like they'd be a good fit. That meaning they need to scrape current data, but otherwise they'd be possible for a person willing to put in some time to combine data sources.

An example is "create a table showing the current ratio of rental/sales prices (on a per square foot basis) for residential properties in the 10 most populated counties in Colorado."

I'm still yet to get a reasonable result from trying this.

dbecker··on ChatPDF – Chat with Any PDF
I tried this a few times about a month ago. It sounded cool, and the experience of using it was even better than I expected.

I'm glad it was reposted so I get another chance at developing a habit of using it.

dbecker··on Darts: A Python library for easy manipulation and forecasting of time series
> That’s actually exactly how vegetarian buffets work.

I'm going to go out on a limb and guess you don't visit many restaurants that advertise vegetarian buffets

dbecker··on Streamlit's $35M Series B
Streamlit is awesome.

It's nicer than sending someone static results, and it isn't much more effort.

And vastly better than sending a notebook to someone unless you expect them to modify the notebook a lot.

And learning time to make Streamlit useful for a small internal apps is probably ~15 minutes for most people.

dbecker··on New Explainable AI Algorithms
This is really nice work.

If you had a mechanism to subscribe for future updates, I'd do it.

Either way, this is cool stuff.

dbecker··on On moving from statistics to machine learning, the final stage of grief (2019)
The author starts with

The data science world may reject me and my lack of both experience and a credential above a bachelors degree

More likely the data science world will reject him because he is so confident a field he has so little experience or knowledge of.

Data scientist is a profession rather than the name of an academic field. So data scientists' job is to solve practical problems. That involves a lot more than class assignments, and in some cases involves using machine learning to maximize predictive accuracy (because common ML models like gradient boosting capture interactions and non-linearities in a richer way than the GLM models the author is familiar with).

Their argument "that's a garbage model because we can't reasonably interpret underlying parameters," is replacing their personal criteria above what is needed to solve some problems.

They can blame it on only having bachelor's degree. But the real problem is the belief that a bachelor's degree taught them everything there is to know, and those in the DS field are ~ idiots who got lucky enough to be paid more.

dbecker··on Show HN: A tool to convert Jupyter notebooks to beautiful blogs
If you send me an email (dan@kaggle.com) I'd love to set up a time to show you some mockups that may be the solution you are looking for.
dbecker··on Vim.dev is redirected to Emacs website
spaces.dev is still available, in case someone wants to buy it and redirect to something like tabs.com
dbecker··on Loneliness on the Job: Why No Employee Is an Island
I haven't worked for an employer that's held that view since high school (when I bagged groceries).

I'm skeptical any successful business treats high-skilled employees that way.

dbecker··on How to Learn Piano with Technology
But it's much cheaper than in-person lessons.
dbecker··on Kaggle Learn review: there is a deep learning track and it is worth your time
I'm the lead on the Kaggle Learn project, and the author of the deep learning track.

I'm happy to answer questions here.

I agree with the commenter saying you need to do your own projects to understand these topics.

Our deep learning track is meant to be the fastest path to knowing enough to do your own projects. You can do the entire track, including the hands-on exercises, in a single sitting.

We won't make you an expert in an afternoon, but you'll know enough to start doing your own projects. For most people that's also the point where Deep Learning becomes fun enough that you'll find time to keep learning.

Kaggle Learn is still in a very early stage. We'll add more lessons soon. But we'll stay committed to the goal of getting you up-to-speed quickly, so you can take on your own projects.

dbecker··on MIT 6.S094: Deep Learning
People make them (and even race) self-driving cars. Check out donkeycar.com
dbecker··on Bayesian analysis of ego-depletion studies favours null hypothesis
It's disappointing to see links to the abstracts on the front page of HN (given the paper itself is behind a paywall).

In this case, the find some non-zero effect, but call it zero because the difference was not statistically significant.

That's likely a reflection of their small sample size rather than evidence for the null hypothesis as the title suggests.

But it's hard to know, or even have an intelligent discussion about it, since the paper itself is behind a paywall.

dbecker··on The Growing Peril of Index Funds: Too Much Tech
This article is talking about diversification to reduce risk, which is unrelated to "timing the market" and which can be consistent with passive investing.
dbecker··on Genetic Algorithms for Training Deep Neural Networks for Reinforcement Learning
Yes and maybe.
dbecker··on TensorFlow r1.4
From tensorflow import keras

Trivially simple example here: https://www.kaggle.com/dansbecker/transfer-learning-scratch-...

dbecker··on Comparison of Neural Network Simulators
This list seemed to miss almost all modern deep learning tools... with the exception of theano, for which the version info here is years out of date. But this small detail at the bottom of the page explains it:

This page was last modified on 10 November 2014

dbecker··on 6 Reasons Why I Am Done with AirBnB as a Renter
The host guarantee on AirBnb isn't as great as most people assume.

We ate $1500 in damages on one of our first visitors, and we're off the platform now. The whole experience was a real bummer.

dbecker··on How economics became a religion
Economic research can be roughly divided into three subfields: Macroeconomics, Micro theory, and Empirical Micro.

The first 2/3 of an undergrad curriculum at most universities is all macro and micro theory. So, most of what an undergrad sees is prove this or solve that.

But far more econ faculty focus on empirical work for their personal research.

It's possible that we focus on theory because it spreads economic orthodoxy faster than empirical work. But economists spend much less time doing theory when they are outside undergraduate classrooms.

dbecker··on Waymo filing says Travis Kalanick knew engineer had Google info
The myth that you can build a billion-dollar business without some shady dealings needs to die

Do you have an explanation for why Uber has been outed with so much more shadiness than most (if not all) other billion dollar tech companies?

Have they done more? Are they worse at PR? Something else? Do you dispute that they have been exposed for more shadiness?

dbecker··on Amazon Prime Wardrobe lets you try on and return clothes free
Wow, borrowing a Holocaust quote.

You just compared good customer service to the largest genocide in human history.

dbecker··on Tensorflow v1.2 released
I have an MBP and an iMac, both of which came with nvidia GPU's. The MBP is older, but I doubt either of these are atypical machines out there today.
dbecker··on You can probably use deep learning even if you don't have a lot of data
We tried to mirror the original analysis as closely as possible - we did 5-fold cross validation but used the standard MNIST test set for evaluation (about 2,000 validation samples for 0s and 1s). We split the test set into 2 pieces. The first half was used to assess convergence of the training procedure while the second half was used to measure out of sample predictive accuracy.

Predictive accuracy is measured on 1000 samples, not 20.

dbecker··on You can probably use deep learning even if you don't have a lot of data
The model we are comparing against makes 10X as many errors.

I hadn't imagined someone would argue that's not a meaningful difference.

Though the difference is statistically significant too.

dbecker··on You can probably use deep learning even if you don't have a lot of data
if you try really hard and get lucky enough, you can probably do as well as a simple regression in this case.

Maybe you aren't familiar with deep learning: but this isn't "trying really hard." This is doing basic stuff that anyone using deep learning probably knows.

And deep learning doesn't just "do as well" as the simpler model. It does meaningfully better at all sample sizes.

dbecker··on Don't use deep learning when your data isn't that big
If you are Google, Amazon, or Facebook and have near infinite data it makes sense to deep learn. But if you have a more modest sample size you may not be buying any accuracy

The author explores sample sizes up to 85, and then suggests this is the relevant range except at Google, Amazon, Facebook, etc.

But the VAST majority of people considering deep learning have sample sizes between those extremes. Results on small samples are interesting, but it's disingenuous to market this as typical of the world outside Google.

dbecker··on Solar Roof
The $80k roof is especially scary if you live in an area with any type of natural disaster risk.

The area I live in just had a huge hail storm. Most roofs in our neighborhood will require repairs. That's a bummer with asphalt shingles, but it would be a lot worse with an $80k roof.

Most people won't think about those risks, but over the time period you need to consider for a roof, it's a big risk in most places.

Page 1 of 17Next →