HNHacker News
TopNewBestAskShowJobs

scott_s

34,207 karma · joined March 2, 2008

Computer science research and systems software development.

http://www.scott-a-s.com

submissionscomments
scott_s··on The Dataflow Model Revisited
It is indeed interesting, and it's from Nathan Marz, who was the creator of Storm. That was the first major open source streaming platform. It predated what I consider the default open source streaming platform, Flink.

Glancing through the docs, my main three reactions are: 1) It's sophisticated system which, as you say, effectively opens up the components of a database engine to be used as needed. 2) Folks will still want to eventually land their data in a "normal" data-at-rest storage format like Hive tables or Parquet. 3) Folks will still want SQL.

scott_s··on The Dataflow Model Revisited
That work is still relevant! It's just that end-users don't need to be aware of it. The position of the paper I submitted, which I basically agree with, is that "streaming" shouldn't need to be something end-users care about. It's something the system does based on needs.

Databases already have a dataflow style architecture: that's how they implement queries. Because SQL is relational, SQL queries become dataflow execution plans.

One way to think about the programming model I worked on is that it was like exposing a query plan API directly to users, instead of giving them SQL.

scott_s··on The Dataflow Model Revisited
Agreed agreed.

During my time in the space, my pithy saying about the system I worked on [1] was that we could scale up or down. If you wanted to do streaming packet filtering with microsecond latency, we could do that. If you wanted to do complex analytics on structured data, we could do that. We did have deployments that "scaled down" and were more stream processing rather than streaming analytics. But analytics is by far the dominant use case, and SQL and relational databases are the better abstraction there. And for the stream processing cases, folks tend to stick to their existing lower-level stacks.

[1] I worked on IBM Streams, https://www.ibm.com/docs/en/streams/4.3.0?topic=welcome-intr..., which had its own language, compiler and runtime system. IBM sold this technology in 2023: https://21cs.com/en/resources/articles/2023/10/10/21cs-acqui....

scott_s··on The Dataflow Model Revisited
I worked in the streaming area for a decade, doing research and development (see: https://scholar.google.com/citations?user=Rdf5OIYAAAAJ&hl=en). After moving on from streaming specifically and moving into the general problems in large data warehouses, I also concluded: just default to SQL for all analytics and the database lens is the best way to think about streaming for analytics.

I still do think that stream programming models are extremely interesting and powerful. But I used to think they would eventually become more mainstream as a way to elegantly program for high throughput, low latency massively parallel systems. That has not been the case, and I no longer think that it will be. People get by with the existing programming languages and models, that seems to be fine.

scott_s··on C++26: Standard Library Hardening Experiments
I have not been following, but some searching lead me to the conclusion that contracts are currently in the draft, but Stroustrup and some others are loudly saying that is a mistake. See: https://wrocpp.github.io/posts/contracts-dispute/
scott_s··on People who grew up with high economic connectedness earn more
A good place to start: https://news.ycombinator.com/item?id=46195226. (The submitter should be familiar.)
scott_s··on Why Large Language Models Fail at Tabular Prediction
I find that an odd take. The paper claims to establish what causes the problem: dimensionality. They are clear in that they don't understand why. But this sort of work is what needs to be done to eventually solve the problem.
scott_s··on TorchCodec 0.14: HDR Video Decoding for CPU and CUDA, and Fast Wav Decoder
We dynamically load the shared object files. See: https://github.com/meta-pytorch/torchcodec/blob/8bbce656797c...

Solutions that shell out to the `ffmpeg` binary are not going to perform well.

scott_s··on TorchCodec 0.14: HDR Video Decoding for CPU and CUDA, and Fast Wav Decoder
I was copying URLs fast-and-furious. The PyTorch Stable ABI reference is: https://docs.pytorch.org/docs/2.12/notes/libtorch_stable_abi...
scott_s··on TorchCodec 0.14: HDR Video Decoding for CPU and CUDA, and Fast Wav Decoder
1. A higher-level API that better integrates into the PyTorch ecosystem.

2. Ease of going back-and-forth between CPU and GPU; in our experience, there's still a lot of scenarios where CPU decoding makes sense.

3. Audio decoding support.

Please take a look at our tutorials to get a feel for what TorchCodec can do: https://meta-pytorch.org/torchcodec/stable/generated_example...

scott_s··on TorchCodec 0.14: HDR Video Decoding for CPU and CUDA, and Fast Wav Decoder
You can get older version of TorchCodec that work with older version of PyTorch, but it unfortunately will not have the new features (HDR video decoding; fast Wav decoding) in the latest release. See the compatbility matrix: https://github.com/meta-pytorch/torchcodec#compatibility-wit...

Up until recently, TorchCodec releases worked with one-and-only-one version of PyTorch. This is because up until recently, PyTorch did not have a stable ABI, and we needed to pin TorchCodec releases to PyTorch releases. But! PyTorch now has an excellent Stable ABI (https://github.com/meta-pytorch/torchcodec#compatibility-wit..., https://www.youtube.com/watch?v=HNdEmnvMvGE&t=1s) and TorchCodec is taking advantage of that since version 0.12.

scott_s··on TorchCodec 0.14: HDR Video Decoding for CPU and CUDA, and Fast Wav Decoder
The one you have installed. :) We don't distribute FFmpeg and instead find your installed version at runtime. We support versions 4 through 8.
scott_s··on TorchCodec 0.14: HDR Video Decoding for CPU and CUDA, and Fast Wav Decoder
For disclosure, I've worked on TorchCodec. I'm happy to answer any questions!
scott_s··on We might all be AI engineers now
You are correct, but this is not a new role. AI effectively makes all of us tech leads.
scott_s··on We might all be AI engineers now
That's not what the author means. Multiple times a day, I have conversations with LLMs about specific code or general technologies. It is very similar to having the same conversation with a colleague. Yes, the LLM may be wrong. Which is why I'm constantly looking at the code myself to see if the explanation makes sense, or finding external docs to see if the concepts check out.

Importantly, the LLM is not writing code for me. It's explaining things, and I'm coming away with verifiable facts and conceptual frameworks I can apply to my work.

scott_s··on What's up with all those equals signs anyway?
I think of, and look up, this drunken rant at least once a year.
scott_s··on ACM Is Now Open Access
Yes, and that peer review happens through the ACM. It serves an organizing function. The conferences themselves are also in-person events, and most of the important research papers come out of those conferences.
scott_s··on ACM Is Now Open Access
It doesn't. arXiv is exclusively a pre-print service. The ACM digital library is for peer-reviewed, published papers. All of the peer-review happens through the ACM, as well as the physical conferences where people present and publish their papers.
scott_s··on ACM Is Now Open Access
IEEE may do it, as it's a professional organization. That is, they're a non-profit dedicated to the furtherance of the field. Being open access fits their mission, and the costs can be handled by dues and fees. Springer and Elsevier are for-profit publishers. I don't know how if they can have an open-access business model.
scott_s··on ACM Is Now Open Access
Great news. They temporarily opened it in 2020 during the pandemic. I argued it should remain so in a post: https://www.scott-a-s.com/acm-digital-library-should-remain-.... I'm glad it's finally happened.
scott_s··on I ignore the spotlight as a staff engineer
Gather metrics and regularly report them.
scott_s··on What Killed Perl?
Agreed. In grad school, I used Perl to script running my benchmarks, post-process my data and generate pretty graphs for papers. It was all Perl 5 and gnuplot. Once I saw someone do the same thing with Python and matplotlib, I never looked back. I later actually started using Python professionally, as I believe lots of other people had similar epiphanies. And not just from Perl, but from different languages and domains.

I think the article's author is implicitly not considering that people who were around when Perl was popular, who were perfectly capable of "understanding" it, actively decided against it.

scott_s··on Scientist exposes anti-wind groups as oil-funded, now they want to silence him
That's true of all renewable energy sources. So we should take advantage of all of them, as much as is feasible.
scott_s··on Claude Sonnet 4 now supports 1M tokens of context
You train on data. Context is also data. If you want a model to have certain data, you can bake it into the model during training, or provide it as context during inference. But if the "context" you want the model to have is big enough, you're going to want to train (or fine-tune) on it.

Consider that you're coding a Linux device driver. If you ask for help from an LLM that has never seen the Linux kernel code, has never seen a Linux device driver and has never seen all of the documentation from the Linux kernel, you're going to need to provide all of this as context. And that's both going to be onerous on you, and it might not be feasible. But if the LLM has already seen all of that during training, you don't need to provide it as context. Your context may be as simple as "I am coding a Linux device driver" and show it some of your code.

scott_s··on Claude Sonnet 4 now supports 1M tokens of context
Because training one family of models with very large context windows can be offered to the entire world as an online service. That is a very different business model from training or fine-tuning individual models specifically for individual customers. Someone will figure out how to do that at scale, eventually. It might require the cost of training to reduce significantly. But large companies with the resources to do this for themselves will do it, and many are doing it.
scott_s··on Claude Sonnet 4 now supports 1M tokens of context
> Of course, because I am not new to the problem, whereas an LLM is new to it every new prompt.

That is true for the LLMs you have access to now. Now imagine if the LLM had been trained on your entire code base. And not just the code, but the entire commit history, commit messages and also all of your external design docs. And code and docs from all relevant projects. That LLM would not be new to the problem every prompt. Basically, imagine that you fine-tuned an LLM for your specific project. You will eventually have access to such an LLM.

scott_s··on My AI skeptic friends are all nuts
The tools are at the point now that ignoring them is akin to ignoring Stack Overflow posts. Basically any time you'd google for the answer to something, you might as well ask an AI assistant. It has a good chance of giving you a good answer. And given how programming works, it's usually easy to verify the information. Just like, say, you would do with a Stack Overflow post.
scott_s··on Exploiting Undefined Behavior in C/C++ Programs: The Performance Impact [pdf]
It's not as obvious a win as you may think. Keep in mind that for every binary that gets deployed and executed, it will be compiled many more times before and after for testing. For some binaries, this number could easily reach the hundreds of thousands of times. Why? In a monorepo, a lot of changes come in every day, and testing those changes involves traversing a reachability graph of potentially affected code and running their tests.
scott_s··on Why is Warner Bros. Discovery putting old movies on YouTube?
Murder in the First is one of them, and it is a long favorite of mine: https://www.youtube.com/watch?v=X42yOL5Ah4E&list=PL7Eup7JXSc...

It has the best performance I've ever seen by Kevin Bacon, and a solid performance from Christian Slater. Gary Oldman is a solid villian. R. L. Emery does his usual thing, but he's really good at that usual thing. I think about lines and ideas from it frequently. Granted, this is partly because the movie came out when I was 15 and I watched it a formative age with friends. But I've also watched it recently, and I think it holds up.

scott_s··on C: Simple Defer, Ready to Use
Look at the Linux kernel. It uses gotos for exactly this purpose, and it’s some of the cleanest C code you’ll ever read.

C++ destructors are great for this, but are not possible in C. Destructors require an object model that C does not have.

Page 1 of 34Next →