HNHacker News
TopNewBestAskShowJobs

mikehollinger

1,155 karma · joined August 5, 2015

I lead teams that design and build high performance hardware and software, most-recently in applied AI for Sustainability, and before that, in the deep learning and hardware accelerator spaces.

Find me at https://www.mikehollinger.com

submissionscomments
mikehollinger··on Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
> Is anthropomorphizing a real problem

Yes. There's a difference between scrapping a session and starting over, or going back and branching something, or using sub-agents to see five outcomes, vs arguing with a system in a long drawn out chat.

Like - I know that if a model starts doing something silly, instead of correcting it - I can probably go back and edit two steps prior to add an extra guardrail, or extra data, or whatever.

mikehollinger··on Book review: There Is No Antimemetics Division
Oh. Nice.

I'd add the Zones of Thought series by Verner Vinge. [1]

[1] https://en.wikipedia.org/wiki/A_Fire_Upon_the_Deep

mikehollinger··on Robots that learn
(needs a tag to be 2017)
mikehollinger··on Helix: A vision-language-action model for generalist humanoid control
From a different robot (Boston Dynamics' new Atlas) - the system moves at a "reasonable" speed. But watch at 1m20s in this video[1]. You can see it bump and then move VERY quickly -- with speed that would certainly damage something, or hurt someone.

[1] https://www.youtube.com/watch?v=F_7IPm7f1vI

mikehollinger··on Time-Series Anomaly Detection: A Decade Review
> what were you thinking then before your aha moment? :D

My naive view was that there was some sort of “normalization” or “pattern matching” that was happening. Like - you can look at a trend line that generally has some shape, and notice when something changes or there’s a discontinuity. That’s a very simplistic view - but - I assumed that stuff was trying to do regressions and notice when something was out of a statistical norm like k-means analysis. Which works, sort of, but is difficult to generalize.

mikehollinger··on Time-Series Anomaly Detection: A Decade Review
This doesn’t capture work that’s happened in the last year or so.

For example some former colleagues timeseries foundation model (Granite TS) which was doing pretty well when we were experimenting with it. [1]

An aha moment for me was realizing that the way you can think of anomaly models working is that they’re effectively forecasting the next N steps, and then noticing when the actual measured values are “different enough” from the expected. This is simple to draw on a whiteboard for one signal but when it’s multi variate, pretty neat that it works.

[1] https://huggingface.co/ibm-granite/granite-timeseries-ttm-r1

mikehollinger··on Things we learned about LLMs in 2024
> About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy....

and

> a key thing with LLMs is that their ability to help, as a tool, changes vastly based on your communication ability.

I still hold that the innovations we've seen as an industry with text transfer to the data from other domains. And there's an odd misbehavior with people that I've now seen play out twice -- back in 2017 with vision models (please don't shove a picture of a spectrogram into an object detector), and today. People are trying to coerce text models to do stuff with data series, or (again!) pictures of charts, rather than paying attention to timeseries foundation models which directly can work on the data.[1]

Further, the tricks we're seeing with encoder / decoder pipelines should work for other domains. And we're not yet recognizing that as an industry. For example, whisper or the emerging video models are getting there, but think about multi-spectral satellite data, fraud detection (a type graph problem).

There's lots of value to unlock from coding models. They're just text models. So what if you were to shove an abstract syntax tree in as the data representation, or the intermediate code from LLVM or a JVM or whatever runtime and interact with that?

[1] https://huggingface.co/ibm-granite/granite-timeseries-ttm-r1 - shout-out to some former colleagues!

mikehollinger··on Turning the Crank: Design as a Mechanical Process
https://c4model.com/ is very useful for this. :-)

I've told it before, but when we were doing some clean sheet work a while ago I decided to use the C4 model and drew out the obligatory "Context" diagram with "user" "phone" "laptop" "app" sort of stuff.

I found them silly and (honestly) I still find that if I see one "in the wild" with no further elaboration I become suspect.

However two hours later, because of that silly context diagram, I realized that we had both an online and a semi-disconnected mobile app that could be offline for hours, and that certain things -had- to use a queue and expect an arbitrary amount of time for a task to run, and it completely changed how we thought about the core of how we implemented something pretty important.

Sold. :-)

mikehollinger··on What do you visualize while programming?
Relationships and sequences.

But if you want to talk about REAL complex systems talk to a microprocessor logic owner or architect trying to shoot a bug.

A while ago we found a bug that could crash a system (fixed in a new RIT of the chip) if we did X then Y in state … we didn’t know.

Listening to the various leads for the sub-units on a phone call trying to reason about what was happening I found myself visualizing this increasingly complicated steam powered machine, with parts sprawling, tiny gears whirring, and bits zipping about whenever X happened.

It was humbling.

mikehollinger··on Meta Movie Gen
> compression mostly makes imperfections go away

The ultimate compression is to reduce the video clip to a latent space vector representation to be rendered on device. :)

Just give us a few more revs of Moore’s law for that to be reasonable.

edit: found a patent… https://patents.google.com/patent/US11388416B2/en

mikehollinger··on Xkcd 1425 (Tasks) turns ten years old today
Eh. A better analogy - the output would decide that there needs to be conduit between floors for chilled water, hot water, sewage, dutifully make several 4” pipes, and then from floor to floor forget which is which.
mikehollinger··on Pivotal Tracker will shut down
I am fascinated by how complex JIRA is. We evaluated it in 2008. It seemed fine enough.

Looking at it 16 years later, and… what is this nonsense? It’s so customizable that it’s loaded with footguns.

mikehollinger··on 0xCAFEBABE & 0xFEEDFACE (2003)
Semi-related story with some insider baseball:

There are quite a few memorable words you can spell using 32 or 64 bits—like BA5EBA11. This is the story of me -not- choosing one of those.

These bit-pattern words are handy because they’re easy to recognize, especially in a random memory dump.

On my first “real” assignment, I was writing real-time embedded C code for a 16-bit processor that communicated with a host microprocessor on a server. We needed to run periodic assurance tests across a bus to ensure reliable communication with the host since we weren't constantly using the bus.*

We were given an unused register address on the host processor and told to write whatever we wanted to it. The idea was to periodically write a value, read it back, and if we encountered any write errors, incorrect reads, or failures, we’d declare a comm error and degrade the system in a controlled manner.

Instead of using zeros or something like 0xDEADBEEF, I decided to write 0x4D494B45 - "MIKE" in ASCII. It was unique, unlikely to be tampered with, it worked, and no one argued with me. The code shipped, the product shipped, and all was well. We even detected legitimate hardware errors, which I thought was pretty cool.

Fast forward two generations of systems, and long after I’d moved on from that team, the code had been ported around but that assurance test remained unchanged. Everything was fine until they brought up a new generation of systems, flipped on the firmware for that device, and 10 seconds later, my assurance test clobbered an important register. The entire system promptly checkstopped and crashed. It took the team days to figure out what was wrong, and I had to explain myself when they found "MIKE" staring back at them from the memory dump.

That was a fun project. ;-)

* Note: It would've been bad if our device went out to lunch because we were responsible for energy management of the server. If the power budget was exceeded and we couldn't downclock and downvolt the processor, something might have crashed or been damaged.

mikehollinger··on IBM Audible Random Timer
Even highly precise and expensive timers will drift.

Lesson learned: we built a 42-node several petabyte storage cluster. Worked fine as we were starting small, and got bigger, and bigger. When we “had it” and disconnected from our internal network and moved to an isolated network we discovered quickly that the timebase on all the systems drifted very quickly to be seconds and then minutes out of sync after several days. This was because they couldn’t access NTP.

Set up an ntp server and “resolved” it, except now we had an entire cluster of hardware that was (in lock step) drifting out of sync with the rest of the world. Fixed that. Moved on.

Another example: we bought a bunch of really cheap LED ropes to make our LGBT+ pride float’s backlighting for a pride parade several years ago. I wired it. The ropes were all independent and we planned to have them rotate colors. The little controller boxes had a mode that would smoothly transition through presets. It turned out that simply turning on the power would cause the lights to start in sync, and then quickly drift out of sync in a really mesmerizing way. They’d occasionally come back together but frankly it was better than anything we could’ve programmed.

We also learned we could control the rate of the “smear” by over or under volting the LEDs a bit.

Even if it’s digital, everything is eventually analog. And analog is weird.

mikehollinger··on Alexa is in millions of households and Amazon is losing billions
>If they want to salvage Alexa, they need to forget shopping and start doubling down on the smart home and assistant experience.

Agree.

Go look at the Alexa Skills for any random category and sort by "best sellers," then sort again by "average review." There isn't an ecosystem.

For Lifestyle, the "4th best selling" [1] skill is "North American Roofing," which is for a company in Tampa.

There should be more there. Given the devices with a touch-sensitive screen, some form of presence detection, location awareness, and other things, there's a lot of missed potential there.

[1] https://www.amazon.com/s?i=alexa-skills&bbn=13727922011&rh=n...

mikehollinger··on Frances Hesselbein's leadership story (2022)
The Peter Drucker compliment is pretty amazing.

As an aside for anyone who’s technical and wants to understand how most corporations work, read “Effective Executive.” It’s from the 60’s but is still very relevant.

It (more or less) is “how to be a knowledge worker.”

mikehollinger··on Google AI recommends adding Elmer's glue to pizza cheese after scanning Reddit
A counter example: I like making hasselback potatoes for special dinners. It’s incredibly tedious to peel, slice, and stack 4 pounds of potatoes.

I dumped in my recipe and notes, suggested the 1 lb bag of shredded raw potatoes instead of the sliced potato and asked for an adjusted recipe I could cook on a cooktop.

These are the most delicious latkes I’ve ever eaten.

mikehollinger··on Chameleon: Meta’s New Multi-Modal LLM
> Numbers like these really don't bode well for the long-term prospects of open source models, I doubt the current strategy of waiting expectantly for a corporation to spoonfeed us yet another $100,000 model for free is going to work forever.

I would add “in their current form” and agree. There’s three things that can change here: 1. Moore’s law: The worldwide economy is built around the steady progression of cheaper compute. Give it 36 months and your problem becomes a $25,000 problem. 2. Quantization and smaller models: There’ll likely become specializations of the various models (is this the beginning of the “Monolith vs Microservices” debate? 3. E2E Training isn’t for everyone: Finetunes and Alignment are more important than an end to end training run, IF we can coerce the behaviors we want into the models by finetuning them. That along with quantized models (imho) unlocked vision models which are now in the “plateau of producivity” in the gartner hype cycle compared to a few years ago.

So as an example today, I can grab a backbone and pretrained weights for an object detector, and with relatively little data (from a few lines to a few 10’s of lines of code, and 50 to 500 images) and relatively little wall clock time and energy (say 5 to 15 minutes) on a PC, I can create a customized object detector that can detect -my- specific objects pretty well. I might need to revise it a few times, but it’ll work pretty well.

Why would we not see the same sort of progression with transformer architectures? It hinges on someone creating the model weights for the “greater good,” or us figuring out how to do distributed training for open source in a “seti@home” style (long live the blockchain, anyone?).

mikehollinger··on Using ARG in a Dockerfile – beware the gotcha
> I don't see the gotcha, that's how it is supposed to work. It's just their purpose

The issue here is that docker evolved rather rapidly and in a “let a thousand flowers bloom” sort of manner. And because of that you have these subtle but confusing differences between behaviors that aren’t really all that consistent.

A good example of this is how the shell is handled from layer to layer(sorta this) or even how CMD and ENTRYPOINT behave (or don’t).

If the spec has allowances for behaviors like this generating warnings would be the best possible outcome (eg referencing a variable that theoretically isn’t set). Maybe certain runtime / runc / build envs complain but the author didn’t see the complaint.

mikehollinger··on Passkeys: A shattered dream
I use 1Password to manage passkeys. It’s pretty nice. They sync across my devices, I can erase one if I need to, and generally with the exception of something odd with Firefox and one website they just work.
mikehollinger··on The design philosophy of Great Tables
> According to everyone that has ever told me how to do powerpoint presentations, they need an intro slide that tells the audience what they are about to be told in the presentation they are about to be told, and a conclusion slide that tells the audience what they were told in the presentation they were told.

Sorta? I guess that might be like driving. Hands on the wheel, gas on the right, etc. But what about driving in the rain (presenting to a cold prospect)? What about driving uphill (addressing a conflict)? What about driving offroad (explaining to an irate customer what happened?

There's storytelling and communication in each of those, but they're different.

See here[1]. It does have an outline, and it does have takeaways, but there's a a structure to the story in the presentation.

It opens with "3 parts" but immediately grabs your attention with item 0. I do that because I'm presenting to engineering students, and there's some CS/ECE folks in the audience and we count from zero. Plus its fun to shake people out of "here we go, another presentation" mode.

Each part has some interesting stories and plot points. The first explains moore's law, and advances in ML / DL. You can't tell from the charts, but there's a narrative that says 2011 to 2016, we saw this advance. There's a relentless march of technology. Isn't that interesting. Part two talks through real projects, and real outcomes, and the shift in value and complexity of delivery at scale vs "we got it working on a laptop." Part three looks at the future - and calls back to pt 1, and proposes a question - what rate do you think stuff'll advance? Here's my guess. Here's some resources. Keep the convo going. Thank you. etc.

It's an incredibly fun presentation to give, and people enjoy it as well. This doesn't have the basic structure you shared, but -- it kind of does. We come back to the "agenda" pages when transitioning between the 3 sections. It has a clear punch line. And it contextualizes what we're going to talk about.

It does this, though, without having a boatload of bullet points and outlines, and, looking at it with fresh eyes, probably makes no sense without a talk track, but people seem to like it and find it useful.

[1] https://ibm.box.com/v/ou-ai-session-2024

mikehollinger··on The design philosophy of Great Tables
You might want to read Edward Tufte's Beautiful Evidence.[1] He discusses stuff like what you brought up about readability and distracting from the message / point of the data.

If you've seen sparklines, [2] Tufte coined the term.

Whenever I do a UI review I end up paging through it just to see if there's something we're not thinking about, and its an interesting book to just open to a random page and read.

Plus he has an entire treatise on why PowerPoint is terrible.

[1] https://www.edwardtufte.com/tufte/books_be

[2] https://en.wikipedia.org/wiki/Sparkline

mikehollinger··on Amazon ditches 'just walk out' checkouts at its grocery stores
> Though it seemed completely automated, Just Walk Out relied on more than 1,000 people in India watching and labeling videos to ensure accurate checkouts.

Yup.

I got the chance to go shopping at one in 2018 (so it's been a while). You could tell there was some "reconciliation" happening, because your receipt didn't show up until 20-30 minutes after you left. My guess is that this was some person via mechanical turk crawling thru the data and indexing what you -really- bought.

Of course, my colleague and I (who were working on computer vision at the time) did stuff like take our bags off, put them in the middle of the floor, and roll a can of soda into the bag, just to see what would happen.

mikehollinger··on 10 Years of Hacker News "Ask HN: Who Is Hiring" Trends
Oh absolutely. It's pretty amazing tbh. I wonder how many people were turned away from positions because they didn't have "SLQ" experience?
mikehollinger··on 10 Years of Hacker News "Ask HN: Who Is Hiring" Trends
Semi-related - I continue to be amused by the number of people that blindly-copy things without understanding what's what. I've seen / heard people refer to LLMs for Vision, which are certainly a thing, but not always what people mean or intend.

And my favorite still is the job posting and experts that refer to "R, Python, SLQ, etc" elsewhere. [1]

This seems to be a copy-paste off of GlassDoor which has -still- not been corrected. [2]

[1] https://www.google.com/search?q=%22R%2C+Python%2C+SLQ%2C+etc...

[2] https://www.glassdoor.com/employers/Job-Descriptions/Data-Sc...

mikehollinger··on Following the Lean Startup Method
Y - the concept of frequent validation differentiated the stuff I’ve worked on that’s done well from the stuff that didn’t. Put a different way: the items that we didn’t get good validation on had to landed with a thud in market, or needed to make sizable changes before they were accepted. The items we validated effectively had to make the same changes, but we got there in smaller steps with less pain.

Also: don’t forget business model validation. Having great tech that sellers won’t sell, or customers can’t buy is a tremendous waste.

mikehollinger··on How photos were transmitted by wire in the 1930s
This is remarkable. Lasers were invented in 1960 as a point of reference.

The painted rope on two spools was also remarkable because of its simplicity - and it holds up today on explaining what a "download" is. ;-)

And even cooler - if we wanted to, I would wager a high school student could implement something along these lines using lego today. There's optical sensors, and you could rig up something to hold a pen like this [1] to render the image.

[1] https://www.youtube.com/watch?app=desktop&v=dHmgaLgFRGM

mikehollinger··on A former Gizmodo writer changed name to 'Slackbot', stayed undetected for months
I’ve worked with someone named True, who, when I went to go add her to some event or another, something along the way helpfully changed it to “TRUE.”

I also worked with a guy whose last name was Null. His email was null@ for a period of time.

mikehollinger··on X.com is Twitter, but what are [a-z].com?
I can email mike@somedomain.com but I can't "call" or "text" mike@somedomain.com . You can kinda sorta see this now with imessage / facetime, but that's not consistent and implemented in a standard protocol.
mikehollinger··on X.com is Twitter, but what are [a-z].com?
> People just google “bob builder” anyway.

Today they do. When a domain was $200 to register in the 90’s, people treated URLs like phone numbers were also treated at the time - to be written down, memorized and then typed precisely in (with slashes!) to find whatever Bob the builder was offering.

It’s odd to me tbh that phone numbers were solved with contact lists and address books, along with the occasional “new phone, who dis?”

Page 1 of 10Next →