HNHacker News
TopNewBestAskShowJobs

staticautomatic

3,651 karma · joined October 19, 2015

if there is a bear at fastmail
submissionscomments
staticautomatic··on Web Scraping in Python – The Complete Guide
Best web scraping guy I ever met (the type you hire when no one else can figure out how) was a Perl expert. I don’t know Perl so I don’t know why, but this is very real.
staticautomatic··on Web Scraping in Python – The Complete Guide
Yeah it’s often gonna be a per site, lots of xpath queries, email me when it breaks kind of endeavor.
staticautomatic··on Show HN: PRQL in PostgreSQL
It’s a tough call. I run a small analytics team and am starting to train some analysts to code. Just the other day I basically told one of my reports to focus on learning Python and let ChatGPT teach him SQL by example because I think it’ll be easier to grok the explanations. Now I’m looking at PRQL and Malloy and asking myself if it’s really a path I should send them down, and I’m not sure it’s a good idea.
staticautomatic··on Diseconomies of scale in fraud, spam, support, and moderation
I live in SF and my pre Uber experience was about as nonexistent as the cabs and the people at the cab company who were supposed to answer the phone.
staticautomatic··on Think Python, 3rd Edition
Oh sure, if I'm understanding you correctly. It's been a few years since I've wrenched on such code, but IIRC, you can pass any seralizable object you want to the function being executed. We used to pass around an in-memory class object, no problem. The tricky part was control flow, but that wasn't celery's problem.
staticautomatic··on Think Python, 3rd Edition
Could you give an example of state internal to the worker?
staticautomatic··on Think Python, 3rd Edition
Rather than handling the multiprocessing and message passing yourself it would be much easier to use celery + gevent and let celery do the work of spinning up processes for executing the tasks.

It may feel limiting but my advice would be to keep all the celery job queue stuff isolated from your server, especially if you are using an async web framework. Have your web server just put all the jobs in the celery queue and let it handle executing them, regardless of whether they’re cpu or io-bound. If you try to optimize too much by doing something like leaning on celery for cpu-bound tasks but letting your web server handle the io-bound ones you’re going to be in for a world of hurt when it comes to both debugging and enforcing the order of execution. Celery has its warts but you’ll at least know where in the system your problem is and have reasonably good control over the pipeline.

staticautomatic··on Think Python, 3rd Edition
That is correct. Honestly I’m not sure this is a good use case for Python but it’s definitely possible. I’ve used the RabbitMQ + gevent + multiprocessing pattern in the past and it works but I find the code extremely hard to reason about. If I were doing it again from scratch I’d probably choose another language with better concurrency primitives.
staticautomatic··on Think Python, 3rd Edition
Well RabbitMQ is built in Erlang so there you go :)
staticautomatic··on Think Python, 3rd Edition
Get garbage collected by what? Doesn’t Python use reference counting?
staticautomatic··on Sora: Creating video from text
I wonder if they could theoretically race multiple people at once like chess masters.
staticautomatic··on Is the "modern data stack" still a useful idea?
Yes, we spent money and engineering time setting up a "data layer" for Google Analytics, though it was mainly to ensure continuity with existing reporting on core site metrics.

That said, most routine data collection in our org isn't from custom instrumentation requiring engineering lift. It happens when someone sets up their own event in Google Analytics or adds a new field in Salesforce or whatever. This kind of bloat is inevitable, and at my scale (which is pretty big but not like F500 big), with my stack, it's always going to be cheaper to just ingest everything than to spend labor on trying to manage and minimize it, especially when I can ingest a lot of stuff for cheap or free and then discard what I don't need.

staticautomatic··on Is the "modern data stack" still a useful idea?
I don't doubt the bugs. In fact, I expect them because so many of the connectors are community contributed and so I try to think of it as more like GitHub than a curated set of expertly-built connectors. That said, for the time being it's more important that I not break my other rules. The promise of a perfectly working connector (if such a thing exists) in exchange for unknowable pricing or a SDK I'm gonna have a hard time training people on is a tradeoff I feel I can't reasonably make right now, but I'm very open to the idea that my calculus may change.
staticautomatic··on Is the "modern data stack" still a useful idea?
I think this is fair and I don’t doubt that AirByte jumped on the MDS marketing bandwagon, but the fact that their pricing isn’t inscrutable sets them apart in my view.
staticautomatic··on Is the "modern data stack" still a useful idea?
I'm a newly minted head of analytics who transitioned from a different domain, so I never had to muck my way through the MDS but attentively watched others from the sidelines over the last few years. Best I can tell, "the modern data stack" is just a marketing phrase invented by a cadre of vampire vendors. The lessons I learned watching others translated into a few simple requirements for our nascent "stack" that most importantly include transparent pricing I can reason about and divvy up, as many integrations as possible so I can minimize rolling my own, and a straightforward framework for ETL code. These three requirements plainly disqualify most of the MDS universe.

With the benefit of starting basically from scratch and not having to mess around with real-time analytics, it's pretty easy to ignore the MDS vendors. So far I've landed on BigQuery, AirByte, GitHub, BI Engine, Looker Studio, and Pandas 2.x or DuckDB for local stuff. I send as many things as possible straight to BQ, lock junior analysts out of gigantic tables, archive periodically to partitioned parquet files in cold storage, use mostly turnkey integrations, and ruthlessly prioritize custom ETL jobs. Putting GitHub in the mix isn't super ergonomic and we may be in the market for new tools once we cross the "big data" frontier, but that'll be a while from now. I'll probably never know or care what the MDS vendors think I'm missing.

staticautomatic··on Is the "modern data stack" still a useful idea?
With respect to the example I gave, no, there actually isn’t. In fact I don’t even understand what you think the work and costs involved are.
staticautomatic··on Is the "modern data stack" still a useful idea?
Sorry but I can't tell if you're being sarcastic or not. Google Analytics and BigQuery are both nominally free, the service expense is $1K/year, and the implementation cost literally a few minutes of my time.
staticautomatic··on Is the "modern data stack" still a useful idea?
It doesn't always make sense to "collect only the data you need." For example, someone in our org will ostensibly have a good reason for setting up an event in Google Analytics. As the head of analytics I have no use for their event but I'm going to end up ingesting it into BigQuery anyway because the built-in GA4->BQ transfer sends ALL the data, and I'm fine with that because it's damn near free for me to ingest and store. I could instead "collect only what I need" by running the transfer through an ELT tool and filtering out that event, but why would I bother doing that work and paying the ELT vendor money when Google will do it for free and charge me like $1K/year to store 100M rows of everything collected in GA4?
staticautomatic··on Woman Got Cremation Ads in the Mail After Getting Chemotherapy
They certainly can under one of the permissible purposes enumerated in the DPPA.
staticautomatic··on Goodbye non-KISS appliances
In CA you can use the lemon law for other stuff. I got the cost of a laptop back that way in small claims court.
staticautomatic··on X blocks Taylor Swift searches after fake AI videos go viral
We already have “false light” laws to address this, and they’re rooted in the common law right of privacy rather than the constitutional right.
staticautomatic··on Generalized K-Means Clustering
What are people using k-means for? I can count on one hand the number of times I’ve had a good a priori rationale for the value of k.
staticautomatic··on Ibogaine banishes PTSD, small study finds
You mean like the administration of magnesium in this study?
staticautomatic··on Pg_rman: Backup/Restore Tool for PostgreSQL
Perhaps someone in this thread will know…if I have a set of pg_dump files, is there any way to read them in memory without an instance of PG running somewhere? It’s an actual problem of mine atm.
staticautomatic··on Scotia, California is owned by a New York investment firm
I’d add Sinclair, WY to that list. I stumbled upon it while driving cross country. It’s kinda surreal.
staticautomatic··on How We Handle Cap Table Information
You mean someone chooses not to invest because your cap table is on Carta?
staticautomatic··on How We Handle Cap Table Information
And what will the damages be? Tort claims require a showing of harm.
staticautomatic··on NHS to investigate Palantir influencer campaign as possible contract breach
B2B companies do this via LinkedIn ads on the regular.
staticautomatic··on Westworld: The first film with CGI (and its source code)
That doesn’t make sense. Nobody would have been having sex with cold robots.
staticautomatic··on Whistleblower Aid
Retaliating against a whistleblower employee is bad news bears in the US
← PreviousPage 9 of 34Next →