HNHacker News
TopNewBestAskShowJobs

StonyRhetoric

83 karma · joined August 6, 2019

submissionscomments
StonyRhetoric··on Ask HN: I’m 41 and still unmarried – what should I do?
A lot of useful things have already been said, so I'll attempt to add value by asking a few questions that might be useful for you to ponder. I'll be as constructive and encouraging as I can.

1. You mentioned you were raised by a single mom - have you seen a healthy marriage before? Perhaps the marriage of close friend or relative? Right now, you have some image in your mind about what "falling in love" and "having a family" means. Where did this image come from? All of us are affected by TV and movies, of course, but, as an adult, you need to see real flesh-and-blood marriages, up-close. Find one or two couples you trust and respect and just ask, and they should be happy to let you in and help. Figure out how to answer this question - what does a good marriage look like to you?

2. More generally, it seems that you are in introspective person - lots of books, podcasts, therapy, exercise, apps. You've dated a lot. But it seems like it's solitary endeavor for you. Do you have close friends/family? Do you have someone you absolutely trust, who understands your life in enough detail that they can see through the "everything's great, how about you?" surface layer? Basically, it's tough to keep your own counsel, to be your own dating coach. You need someone else who can be a loving critic. Ideally, someone who is mature and relationally-successful.

3. We, as ego-protecting humans, have an incredible capacity for self-deception. Something isn't quite adding up. Given your qualities, your evident desire, and your extensive efforts - you should have been more successful they you currently are. Are you sure you're telling the whole story here? For example, some people deeply desire emotional intimacy, but become fearful when actual intimacy is within reach, because they can't bring themselves to show true vulnerability, and risk true rejection. Some people don't want to lose optionality, as, after all, marrying one person means losing out on the option of marrying anyone else. There's something missing here. You don't have to tell me, but it is important that you tell yourself.

4. While often unpleasant to think about, dating has a comparative, competitive aspect. At the very least, your competition will include 30-year-olds who resemble what you were like at 30. In what ways have your grown and become better than the 30-year-old you? What is your "competitive advantage"? And what kind of man will value and appreciate these qualities? Who is your "target audience"? Bluntly, what type of guy is going to pick you over the 30-year-old version of you?

5. Not a question, but I want to agree with comments encouraging you to look at "second-hand men". Basically, you missed the first bus, where all the conventional/normal people paired up between 25-35 years old. What's left are those who didn't get paired-up, or were paired-up but no longer are (divorced, widowed). You might have to go outside the apps, and maybe outside your usual circles in Austin to find them. Almost by definition, they will not be on the stereotypical life trajectory. In my mind, you're looking for someone who has been knocked down but has gotten back up, someone that life has already sanded away the rough edges. My aunt got married pretty late (40), and she found a chain-smoking, obese businessman who was working himself into an early grave. He quit smoking for her, started losing weight, and their two kids are now in college. He's a really cool guy, super funny and generous. How she saw that, back then, I don't know. Somehow I don't think the apps would have matched them.

6. Also not a question - but, if you haven't noticed, people in Texas tend to get married early, and Austin is a college town - there will be an endless supply of marriage-minded 20-year-old women there. Food for thought.

Good luck. And in case it matters, I am a happily-married Texan man with two kids, and, if you ever decide to convert to Christianity, I'd be happy to introduce you to an eligible (and slim and tall) doctor with a somewhat controlling mother.

StonyRhetoric··on Evidence-based software engineering: book released
I'm having a hard time deciding if this is the work of a well-meaning but misguided obsessive, or GPT-3 generated pseudo-gibberish.

In any case, this book could use a great deal of editing. If there is wisdom in here, I don't want to dig through hundreds of pages to find it.

StonyRhetoric··on Show HN: Obsidian – A knowledge base that works on local Markdown files
This is amazing - I have a homebrew version of this - markdown only, with some scripts to wire-up the links, relying on unique file names.

This is miles better that what I have, and I look forward to using it.

A feature request - mermaid JS support please! This would allow me to import my ~3k existing markdown files without modification.

StonyRhetoric··on How to build your own feature store for ML
This is a good idea - every ML operation should have something like this, to store, organize, version data, check for drift, do time-travel, backups/replication et cetera.

But to borrow from Steve Jobs, I think this is a feature, not a product. If you've already done the hard work of setting up a data lake or data warehouse in a cloud provider, the cloud provider can give you backups and replication, and even some time-travel. Using something like Delta Lake or even just the standard Kimball DW audit columns will get point-in-time queries. Feature versioning is just query versioning in source control, and if you have schema, you can schema version with views if you need to. If you don't have a data lake, data warehouse ... well, you'll still need to gather and clean all your data before you put it into a feature store, and that's where 90% of the work is.

I'd love to learn more, I'm sure I'm missing something, but it seems that they're re-solving the solved part - data storage and versioning. Checking for drift and data integrity is a nice bonus, but again, lots of libraries for that. I guess I could see it being beneficial for ML shops that don't have modern development practices, but if you don't have that, you have bigger problems anyways.

StonyRhetoric··on Why is AI so useless for business?
This is clickbait (1) to promote his startup, Proda.

ML is used in business workflows all the time - to date, I have built several solutions that are being used for 53 clients, internal and external.

Here is what makes B2B ML hard: People have to trust it.

This isn't some movie-recommendation engine, which spams you with more bank heist movies after you watch one. B2C ML systems can get it wrong, and customers are generally forgiving, because it's a low stakes game. B2B applications are generally higher-stakes, because they impact business workflows, and if someone has decided to automate it, it's probably a high-volume, critical workflow. It has to be extremely accurate, and demonstrably better than the equivalent human system.

The problem has to be well-defined enough that an ML system can act with high-accuracy, but not well-defined enough that a rule-system could replace it. Don't use ML if a rule-system will do a better job. (For those scenarios, you can still put an ML anomaly-detection system to make sure the rule-system is still valid, and to guard against data input changes.) As just mentioned, the problem also has to be important enough and high-volume enough to warrant an ML solution. The percentage of problems that fulfill these criteria is not very large.

Now to actual ML development and deployment - the model is the tip of the iceberg. The rest of the iceberg is data acquisition, feature selection, data/feature versioning, automated training, CI/CD, model performance monitoring, et cetera. If ML is being developed inside a software development organization, this isn't a problem, most people will understand this. If it is being developed within an embedded BI team inside a business unit - they will generally not have support/runway needed to build the full system. The ML model might make it to production, but it will probably run naked, be brittle, and hard to retrain. A dramatic failure with business impact is just a matter of time.

There are a lot of low-code, no-code ML solutions that have been developed, or are being developed, and some of the supporting infrastructure as well, but, at the risk of sounding parochial/protectionist, you need a rock-solid, end-to-end, integrated, data management system that is fully understood by whomever needs to pick up the phone at 2AM. It's the interfaces that are hard, and chaining together a bunch of third-party black-box systems just means more interfaces and behavior you don't control. Choose and use these systems wisely.

So yeah, B2B ML is hard. But it's generally not due to lack of data, and transfer learning is generally not necessary. Understanding business processes is important, I agree, but that's comparatively easy. It's what consultants have been doing for decades. The hard part is choosing a problem where ML can add value, and then executing on it with enough integrity that people will actually trust it.

(1) Ok, clickbait might be harsh. But it is self-promotion, and the article itself is a collection of generic banalities. I feel it falls on the wrong side of the line.

StonyRhetoric··on Data Science: Reality Doesn't Meet Expectations
As the lead data scientist at a small-ish fintech, I can confirm many of the frustrations and disappointments in the OP. But my trajectory was slightly different - from being the only "data science guy" in 2016, to now leading an autonomous team of four, with quarterly meetings with the CEO, and monthly meetings with our tech leadership. I decide tech stack, workflow, and hiring. Execs decide priorities. Sure, some of it was dumb luck, some of it was actually having a CEO that cares about data strategy, but I like to think at least some of it was me.

So here's what I think I did right:

1. Provide indisputable, obvious business value every month. You should consider yourself an in-house consultant to whichever cost center your salary is drawn from. If you're product development, prove value to them. If you're operations, or sales, or marketing, prove value to them. After about two months, you should be able to justify your existence in two sentences. Just remember, most of your company probably thinks of you as a optional add-on.

Your first few projects should attack high-impact pain points with the simplest solutions possible. My first projects were basically ETL into some basic regression into a dashboard. No machine learning required. But it was better then what they had (which was often nothing), and it was STABLE and RELIABLE. And that leads to the next point...

2. Build trust. With my dead-simple models, nothing ever blew up, there were no nonsensical answers, and there wasn't much brittleness when new categorical features or more cardinality was added. It mostly just worked. And that built my reputation for me. They didn't have to understand what was going on in the model, but they knew, from experience, that they could trust the result. Once I had the credibility, I could start building more complex, more elaborate models, and asked them to trust those as well. If they don't trust your models, then no business value has been created, and your job is worthless.

3. Recognize that data science is being done everywhere in the organization, and respect it. Every department has someone who has built a monster spreadsheet that contains more embedded domain knowledge then you could hope to learn in a month. As data scientists, we like to think that we're helping the organization by building critical metrics to improve performance. But here's the catch. If the metric was truly critical, someone has built it already. It might be ad-hoc, use poor-methodology, and be somewhat wrong, but it works and is good enough. You have to find that person, learn from them, and improve on it.

4. Be as self-contained as possible. Ideally, your critical path should not depend on other teams doing things for you (except for IT setting up data access). You should be able to do it all. From front-end dashboards, to ETL, to DevOps. Remember, you're an in-house consultancy. You should be able to take problems and just handle them, rather then be a perpetual bother and distraction to other teams.

There's more, but if you do these four things, I think you can build the reputation in your company for creating useful, accurate data tools that help other people do their jobs better. After that's achieved, people will breaking down your door to get your help. That's where my team is now - we've got a backlog for at least 18 months, with our work priorities often being set directly by the CEO.

StonyRhetoric··on Show HN: Aim and Shoot – A game where your opponents are neural networks
Played three times, made it to gen 23 on the third time. Fun game.

Without knowing much about game mechanics, a "domestication" strategy seems to work well.

1. Pick a corner. Bottom-right was what I chose.

2. Move there without getting shot.

3. Shoot first at the robots shooting at your direction, then those with guns pointed in your direction. Then shoot the robots that are shooting. Save the robots that are pointed the other way, not shooting for last.

4. After a few generations, all the robots will be pointing the other way, not shooting. Kill the ones that twitch first.

5. There seem to be randomization events, and some of your domestication will be lost. Try to survive those and re-domesticate.

6. Eventually you run out of non-replenishable HP and die.

StonyRhetoric··on Ask HN: How to Make Money from Machine Learning?
I manage a team that directly uses ML to create software services that our customers pay money to use. This is relatively rare. Most ML applications are internal - your customer is your own company. In this case, you should think of yourself as an in-house ML consultancy for your company. You need to reliably create value for the company, and be able to measure it.

In my opinion, there are three business "tiers" of ML.

1. Process automation. You're turning a defined business process currently done by a human into something automated, with some custom rule logic, of which some of it may be via ML-trained models. The easiest, because the criteria are well-defined and everyone knows what success looks like.

2. Data Mining/Analysis/Insight. Your company sits on some unexploited set of data, and you want to generate useful business insight from it. ML models can help make sense of it. This takes the traditional business intelligence function to the next level. Harder, because you may need to educate the company on what ML offers. They may not even realize what types of new questions ML can answer.

3. Customer-facing automated-decision services. This is the most demanding application from a business perspective, but not necessarily from a technical perspective. The standards for quality, stability, accuracy, should be much, much higher. If it's customer-facing, it can't mess up, or people will stop trusting it. The customer may be internal or external.

StonyRhetoric··on It’s worth spending weeks on research before wasting years on a hopeless project
Speaking from the other side, my advice is to over-communicate, especially if you haven't built up a deep trust and reputation with your team and your manager.

Here's what your manager is worried about:

  1. Will the job be done well?
  2. How long will this take?
  3. Are you on-task, or are you stuck, overwhelmed, going rogue?
You can help your manager answer these questions by communicating these things:

  1. I have a plan, and here is the plan.
  2. Here is my progress in executing this plan.
  3. I have made these findings so far. I anticipate 
  these difficulties and challenges.
How this is communicated to the team, and your manager, depends on your team culture and workflow. But something typical might be:

  1. Two sentences in standup. Think through what you are 
  going to say, read a prepared statement from a sticky 
  note if necessary.
  2. A short email (<100 words) answering any questions 
  people have regarding your research. Use your own 
  judgment as to audience size and email frequency.
  3. A document or wiki page that is the final work 
  product of your research.