NSA releases 1982 Grace Hopper lecture
nsa.gov
nsa.gov
> "Now back in the early days of this country, when they moved heavy objects around, they didn't have any Caterpillar tractors, they didn't have any big cranes. They used oxen. And when they got a great big log on the ground, and one ox couldn't budge the darn thing, they did not try to grow a bigger ox. They used two oxen! And I think they're trying to tell us something. When we need greater computer power, the answer is not "get a bigger computer", it's "get another computer". Which of course, is what common sense would have told us to begin with."
In the Pacific Northwest, US, early loggers would leave the huge ones - to the point where pioneers could complain about a lack of available timber in an old-growth forest.
When the initial University of Washington was built, land-clearing costs were a huge portion of the overall capital spend. The largest trees on the site weren't used for anything productive; rather, they were climbed, chained together, and domino felled at the same time. By attaching the trees together, they only needed to fell one tree which brought the whole mess down into a pile and they burned it.
I think there's a lesson here about choosing which logs you want to move.
<https://www.etymonline.com/word/log> (definition 2)
Chip log: <https://en.wikipedia.org/wiki/Chip_log>
Dead reckoning: <https://en.wikipedia.org/wiki/Dead_reckoning>
Scilly naval disaster of 1707, in which four warships and 1,400 to 2,000 sailors of the British fleet were lost in a single navigational failure: <https://en.wikipedia.org/wiki/Scilly_naval_disaster_of_1707>
This directly inspired the Longitude Rewards prizes for developing accurate navigational methods (<https://en.wikipedia.org/wiki/Longitude_rewards>) and John Harrison's invention of a marine clock (<https://en.wikipedia.org/wiki/John_Harrison#Longitude_proble...>).
More recently, the Honda Point disaster in which the US Navy, 1923 saw the loss of 7 destroyers at flank speed of 20 knots off the Santa Barbara coast: <https://en.wikipedia.org/wiki/Honda_Point_disaster>.
It's interesting to note that advances in timekeeping typically translate to improvements in location determination.
Actually, humans have been doing exactly that through breeding over the millennia. They were just limited in their means.
This analogy has some "you wouldn't download a car!" vibes — sure I would, if it were practical (: And vertical scaling of computers is practical (up to some limits).
You spoke of one log, and the time scales involved. But suppose you have an entire forest of logs. Then it may indeed be worth breeding bigger oxen (or rather, inventing tractors).
I don't mean to accuse Hopper of shortsightedness, but when quotes by famous people, like the above, are thrown around without context, they encourage that dogmatic thinking.
So, I was more replying to that quote as it appeared here, rather than as it appeared in her talk.
I don’t think there’s anything about the original post, with quote about oxen, that reads as dogmatic, or invites such perspective.
Also, I think we can all agree most innovation happens as an extension of “making the best of what’s available” rather than independent of it, on a fully separate track.
Using two oxen can lead to realizing a bigger ox would be beneficial.
... they're trying to tell us something. When we need greater
computer power, the answer is not "get a bigger computer", it's
"get another computer".
that does read as dogmatic advice to me, taken in isolation. It boils down to "the answer is X."
Not "consider these factors" or "weigh these different options," but just "this is the answer, full stop."That's dogma, no?
(that aside, I do slightly regret the snarkiness of my initial comment :)
The only exception is really where you have a bounded task that will never grow in compute time.
Was that all just a bunch of wasted effort, and what they should have been doing was build more and more 50MHz chips?
Of course not. There are lots of advantages to scaling up rather than out.
Even today, there are clear advantages to using an "xlarge" instance on AWS rather than a whole bunch of "nano" ones working together.
But all this seems so straightforward that I suspect I really don't understand your point...
If you waited for chips to catch up to your workload, you got smoked by any competitors who parallelized. Waiting even a year to double speed when you could just use two computers was still an eternity.
> Was that all just a bunch of wasted effort, and what they should have been doing was build more and more 50MHz chips?
No, that’s a stupid question and you know it. You set it up as a strawman to attack.
Hardware improvements are amazing and have let us do tons for much cheaper.
However, the ~4ghz CPUs we have now are not meaningfully faster in single thread performance compared to what you could buy literally a decade ago. If you’re sitting around waiting for 32ghz that should only be “3 years away”, you’re dead in the water. All modern improvements are power savings and density of parallel cores, which require you to face what Grace presented all those years ago.
Faster CPUs aren’t coming.
xlarge on AWS is a ton of parallel cores. Not faster.
There is risk in reinforcing a narrow-minded approach that "all we need is more oxen." It limits one's imagination. That's the essence of what I've been advocating against in this thread, though perhaps my attempts and examples have merely chummed your waters. Ironically, I'd say Grace Hopper rather agrees, elsewhere in the linked talk[1].
> Faster CPUs aren’t coming.
Not with that attitude, ya dingus (:
[1] "https://www.youtube.com/watch?v=si9iqF5uTFk&t=1420s
I think the saddest phrase I ever hear in a computer installation
is that horrible one "but we've always done it that way." That's a
forbidden phrase in my office.I liked grace hopper's comments as a rebuttal against "only vertical! No horizontal!" but I'd agree that reading that rebuttal dogmatically would be just as bad of a decision.
Bigger is better in terms of height and girth when it comes to capabilities. At any given time, figure out the most cost efficient number of oxen of varying breeds for your workload and redundancy needs and have at it. In another year if you're still travelling the Oregon trail you can reconsider doing the math again and trading in the last batch's oxen for some new ones, repeat as infinitum or as long as you're in business.
Your point has been made and I’m telling you very explicitly that it’s bad. The years of waiting for faster processors have been gone for basically a generation of humans. When you hit the limit of a core, you don’t wait a year for a faster core, you parallelize. The entire GPU boom is exemplary of this.
Despite the very plain language her talk has a lot of depth to it and I do think how interesting how on the money she was with her thoughts all the way back then.
The scale-up/scale-out tradeoff applies to many things, both in computing and elsewhere. I was trying to make a larger point.
I guess it's appropriate, in this discussion about logging, that we got into some mixup between the forest and the trees (:
That’s idiotic unless you have other constraints. The parallelism allows you to also break apart the oxen to do multiple smaller logs at the same time when their combined force isn’t needed.
One potentially more likely solution back in the day was to just accept the job was going to take a while. This would be analogous to using a block and tackle. The ox can do the job but they're going to pull for twice as long to get it done. Imagine pulleys cost $10, but a second ox costs $1000 and a yoke costs $5000, and getting the job done in less time is not worth $5,990 to you.
Admiral Hopper's lecture wasn't delivered too long after 1976, which saw the release of both the CRAY-1 (single CPU) and the ILLIAC IV (parallel). ILLIAC IV, being more expensive, harder to use, and slower than the CRAY-1, was a promising hint at future possibility, but not particularly successful. Cray's quip on this subject was (paraphrasing) that he'd rather plow a field with one strong ox than $bignum chickens. Admiral Hopper was presumably responding to that.
What they both seem to miss is that the best tool for the job depends on both the job and the available tools. And they both seem to be completely missing that, if you know what you're doing, scale up and scale out are complementary: first you scale up the individual nodes as much as is practical, and then you start to scale out once scale up loses steam.
One normal-size, versus 10 miniature ones?
Needs research (:
In your attempt to take down the analogy you just reinforced it. They quickly hit the limits of large oxen and had to scale up far faster than any selective breeding could help.
The exact same thing happened in computing even during the absolute hay day of Moore’s law. Workloads would very quickly hit the ceiling of a single server and the way to unblock yourself was not to wait for next gen chips but to parallelize.
This technique famously remodeled the iconic Tiffany building in NyC. https://www.mgmclaren.com/projects/crane-lift-at-tiffanys/
Though I have to say the part about the cost of not implementing standards, the cost of not doing something, felt scarily relevant right now.
It's still kind of hard even now. To date in my career I've had more successes with improving existing systems' throughput by removing parallelism than I have by adding it. Amdahl's Law plus the memory hierarchy is one heck of a one-two punch.
https://en.wikipedia.org/wiki/Cray_X-MP
because you could still make bipolar electronics that beat out mass-produced consumer electronics. By the mid 1990s even IBM abandoned bipolar mainframes and had to introduce parallelism so a cluster of (still slower) CMOS mainframes could replace a bipolar mainframe. This great book was written by someone who worked on this project
https://campi.cab.cnea.gov.ar/tocs/17291.pdf
and of course for large scale scientific computing it was clear that "clusters of rather ordinary nodes" like the
https://www.cscamm.umd.edu/facilities/computing/sp2/index.ht...
we had at Cornell were going to win (ours was way bigger) because they were scalable. (e.g. the way Cray himself saw it, a conventional supercomputer had to live within a small enough space that the cycle time was not unduly limited by the speed of light so that kind of supercomputer had to become physically smaller, not larger, to get faster)
Now for very specialized tasks like codebreaking, ASICs are a good answer and you'd probably stuff a large number of them into expansion cards into rather ordinary computers and clusters today possibly also have some ASICs for glue and communications such as
https://blogs.nvidia.com/blog/whats-a-dpu-data-processing-un...
----
The problem I see with people who attempt parallelism for the first time is that the task size has to be smaller than the overhead to transfer tasks between cores or nodes. That is, if you are processing most CSV files you can't round-robin assign rows to threads but 10,000 row chunks are probably fine. You usually get good results over a large range of chunk size but chunking is essential if you want most parallel jobs to really get a speedup. I find it frustrating as hell to see so many blog posts pushing the idea that some programming scheme like Actors is going to solve your problems and meeting people that treat chunking as a mere optimization you'll apply after the fact. My inclination is you can get the project done faster (human time) if you build in chunking right away but I've learned you just have to let people learn that lesson for themselves.
My big sticking point is that for some key classes of tasks, it's not clear that this is even possible. I've seen no credible reason to think that throwing more processors at the problem will ever build that one tool-generated template-heavy C++ file (IYKYK) in under a minute, or accurately simulate an old game console with a useful "fast forward" button, or fit an FPGA design before I decide to take a long coffee-and-HN break.
To be fair, some things that do parallelize well (e.g. large-scale finite element analysis, web servers) are extremely important. It's not as though these techniques and architectures and research projects are simply a waste of time. It's just that, like so many others before it, parallelism has been hyped for the past decade as "the" new computing paradigm that we've got to shove absolutely everything into, and I don't believe it.
What you can do is run g and h currently in something that looks like f(g(), h()). And you can vectorize.
A lot of early multiprocessor computers only gave you that last option. They had a special mode where you'd send exactly the same instructions to all of the CPUs, and the CPUs would be mapped to different memory. So in many respects it was more like a primitive version of SSE instructions than it is to what modern multiprocessor computers do.
Originally, the whole point of the Hadoop architecture was that the data were pre-chunked and already sitting on the local storage of your compute nodes, so that the overhead to transfer at least that first map task was effectively zero, and your big data transfer cost was collecting all the (hopefully much smaller than your input data) results of that into one place in the reduce step.
Now we're in the cloud and the original data's all sitting in object storage. So shoving all your raw data through a tiny small slow network interface is an essential first step of any job, and it's not nearly so easy to get speedups that were as impressive as what people were doing 15 years ago.
That said I wouldn't want to go back. HDFS clusters were such a PITA to work with and I'm not the one paying the monthly AWS bill.
All of those were duds. Other than the PS3 Cell, also a dud, none of those architectures were built in quantity. They're really hard to program. You have to organize your program around the data transfer between neighbor units. It really works only for programs that have a spatial structure, such as finite element analysis, weather prediction, or fluid dynamics calculations for nuclear weapons design. Those were a big part of government computing when Hopper was active. They aren't a big part of computing today.
It's interesting that GPUs became generally useful beyond graphics. But that's another story.
Some TPUs are also structured around fixed dataflow (systolic arrays for matrix multiplication).
It's not really comparable to the other examples you cite - the ncube/transputer/connection machine, in that it was programmed conventionally, not requiring a special parallel language, but Tandem's NonStop was this, starting in ~1976 or 1977. Loosely coupled, shared-nothing processors, communicating with messages over a pair of high-speed inter-processor busses. It was certainly a niche product, but not a dud. It's still around having been ported from a proprietary stack-machine to MIPS to Itanium to X86.
EDIT: I suppose it can be compared to a Single System Image cluster.
Dual socket Epyc is pretty big these days.
If you can fit your job on one box (+ spares, as needed), you can save a whole lot of complexity vs spreading it over several.
It's always worth considering what you can fit on one box with 192-256 cores, 12TB of ram, and whatever storage you can attach to 256 lanes of PCIe 5.0 (minus however many lanes you need for network I/O).
You can probably go bigger with exotic computers, but if you have bottlenecks with the biggest off the shelf computer you can get, you might be better of scaling horizontally, but assuming you aren't growing 4x a year, you should have plenty of notice that you're coming to the end of easy vertical scaling. And sometimes you get lucky and AMD or Intel makes a nicely timed release to get you some more room.
Amusingly, in the horse-powered era, once railroads started working, but trucks didn't work yet, there was a "last mile" problem - getting stuff from the railroad station or dock to the final destination. The 19th century solution was to develop a bigger breed of horse - the Shire Horse.[1]
(I've owned a Percheron, and have known some Shires.)
(relatives have owned Friesian's and Clydesdale's and Norwegian Fjording Horses, but it's neither here nor there)
To advance; to further; to prefect; to make to increase; to promote the growth of. “We must develop our own resources to the utmost. Jowett (Thucyd).”
Rather than develop as in the contemporary software engineering sense synonymous with create.
In this sense, the breed was further developed to serve railway terminals from the original breed created to service maritime ports.
So I think “developed” is fairly appropriate here, though “adapted” might have been more clear.
In engineering we are accustomed to getting involved early in the creation process, and our usage of "develop" reflects this bias.
Outside of that bubble, "develop" is very explicit that the thing already exists. For example, developing a musical theme, a muscle, or a country.
I hadn't twigged on the origin of "station wagon". TIL.
That article was re-posted here on HN and elsewhere but didn't seem to get much attention and I feared the worst, since 1-inch magnetic video tape degrades with time. Very frustrating since such vintage VTRs do exist in working order in the hands of museums, video preservationists and collectors. Now six weeks later we get the best possible news! Hopefully, that article and the re-postings helped spread the word and someone in control of access to the tape got connected to someone with the gear.
And what an amazing piece of history to have preserved. I'm only ten minutes into the first tape but she's obviously a treasure - clear thinking, great communication and a sharp wit. Even captured here later in life you can clearly see why she was so successful and highly regarded by her peers (including some the most notable people in early computing history).
>While NSA did not possess the equipment required to access the footage from the media format in which it was preserved, NSA deemed the footage to be of significant public interest and requested assistance from the National Archives and Records Administration (NARA) to retrieve the footage. NARA’s Special Media Department was able to retrieve the footage contained on two 1’ APEX tapes and transferred the footage to NSA to be reviewed for public release.
[0] https://www.nsa.gov/Press-Room/Press-Releases-Statements/Pre...
1: https://www.schneier.com/blog/archives/2004/10/the_legacy_of...
Further, the truncated version of DES that got standardized far outlasted its expected lifetime --- the National Bureau of Standards expected DES to have a useful lifetime of about 5 years. And even at the time it was understood that you could expand the keysize by tripling up the DES core.
I think there's a really big difference between publicly weakening a standard, in effect telling the world "we want a standard that is adequate for commercial purposes but inadequate for military purposes, so as to retain our national edge", and doing what they did with Dual-EC, where it was impossible (apparently) for people to reason about what NSA was up to.
Schneier was clearly able to reason about what NSA was up to, and told everyone in 2007 not to use Dual-EC, 6 years before the Snowden revelations.
I believe you have admitted that you thought that “Dual-EC has a backdoor” was a wild conspiracy theory until the Snowden revelations? Which makes the “impossible (apparently)” part a classic case of projection.
(I thought nobody should use Dual EC! But that was my reason for thinking it wasn't an NSA backdoor, because it was too dumb to be one. I underestimated the industry's capacity for "dumb". Also: I was dumb! I am dumb a lot.)
Thankfully we found a better way that ensures cryptographic security, which is to get former NSA interns to write the PQC standards, instead of proper NSA employees.
A better question: why do you think so many of your cryptographic feline friendz were so excited about isogenies for the past decade? Where do you think they all obtained that identical enthusiasm from? Why do you think SIKE made it so far in the contest and only got eliminated through luck?
I'm guessing this isn't a conversation that's going to take us into Richelot isogenies.
Is Dual-EC-DRBG fine because we never saw the FVEY Python exploit that breaks it?
I think my theory here is that NSA coordinated an action whereby they figured no one was reading obscure algebraic geometry papers from 1997. In our low-attention-span world, it’s not the worst plan.
(Hell, folks didn’t realize TAOSSA contained 0day for a long time. Simply putting something in front of the public doesn’t mean they’ll read or comprehend it.)
Dual EC isn't broken by an exploit script. It's broken with a secret key.
No, it leaves every SIKE-protected system in the world exposed to _everybody who reads obscure algebraic geometry papers from 1997._ We got really lucky that the two dorks who do read those papers decided to share their insights.
For all you know, there’s a paper sitting at the Institute For Advanced Study that would let you write a marvelous pq-crystals-shattering Python script, but they’ll never tell you the combination to the safe.
(Again: TAOSSA contained 0day exploits, and few noticed for a decade.)
However in this context the debate is just over the PQ scheme (not the overall system). Also, NSA are not planning to mandate a hybrid system for government use. Others may do the same.
I supposed they did (allegedly) pay RSA Security to make this the default choice in BSAFE but that seems like an awful lot of work to hack one product.
Another thing I was very certain (and certainly wrong) about was that no competent team was using BSAFE in 2010. The more I've learned about cryptography the less confidence I've held onto in industry cryptography practices outside of Google, Apple, and Microsoft. I would have assumed the major networking vendors were playing at roughly the same level. Yikes, no.
Only if everything you know about the NSA comes from the evil, cackling, mustache-twirling caricatures of it promulgated by angry people on the internet.
Once you look beyond the politics, propaganda, and axe-grinding that is endemic to the online world you find out all sorts of fascinating things about the U.S. government.
If you consume news primarily from, say, Hacker News, then sure.
If the NSA, and other intelligence agencies, had any influence on the election, why wouldn't they do exactly what it would appear they are doing now and get a milquetoast liberal elected to office who will easily capitulate to their demands?
Did they inject Biden with a dementia drug to force a withdrawal and engineer the timing such that the current Vice President was pretty much the only viable option for the US Democrats to rally behind?
Seems like a tightrope feat of Rube Goldberg Heath Robinson needle threading.
It's not a one-way relation to power. Intelligence agencies are nothing if not opportunistic, they can influence elections but if one of the candidates is clearly incompetent there isn't much they can do unless he drops out. You're forgetting that Jill Stein would've never been endorsed by Biden; what appears to be chaotic and contingent actually has a strong set of boundary conditions of possibility that all the contingency is contained within, and intelligence agencies, including even the state department for foreign affairs, try to control that. Not individual actions, but the ability to perform them, the rationality of it. The fact that you can't even imagine a candidate besides Donald Trump who poses a serious threat to the state intelligence apparatus shows you that they've already won, or at least nearly so.
?
How'd you get this incorrect insight into what I think ... and what makes you think that Trump is a serious threat to the US state intelligence apparatus?
I think you claimed that Harris was somehow not an ideal candidate for the current hegemonic forces in the US, or at least those forces of power wouldn't do what they could to make sure she gets elected. One of Chomsky's points was precisely this, they goad you with progressive political candidates who don't actually threaten power. The two main forces of power in the US are capitalist industry and the state, but the truth is that both have an interest in maintaining power relations such as they are, and so what we are witnessing in most elections is just a sort of balancing act between direct and indirect means of control. With Trump you have someone who is so insanely narcissistic that he is completely unreliable and there is essentially no way of using him to maintain state control as such.
Trump is a threat to democracy, not to TLA's.
> I think you claimed that Harris was somehow not an ideal candidate for the current hegemonic forces in the US,
I made no such claim. Perhaps you might like to scroll back and identify where I did, I suspect you've confused me for another.
As if America is a democracy
What you're suggesting doesn't exist - and is being skirted around by the news - is in fact widely available. Google's right there.
> If the NSA, and other intelligence agencies, had any influence on the election, why wouldn't they do exactly what it would appear they are doing now and get a milquetoast liberal elected to office who will easily capitulate to their demands?
This strikes me as working backwards from a conclusion. If in your view the intelligence community would operate in that way, how would you ever know one way or the other?
One thing we can certainly agree on is that Trump is the real threat. It is pretty damning of our age that "not having a platform"(to your satisfaction) is supposed to be met as a serious criticism, but her opponent's openly unhinged behavior is just "how it is".
She doesn't need one: the fact that she's not Trump, and she's not old enough to be senile or on death's door, is all she needs for most voters. It's not like the Democratic Party had a bunch of other viable candidates in a position to mount a presidential campaign this close to the election.
If you want to criticize the US for having a crappy FPTP election system that basically guarantees only two viable parties on the national stage, that's fair, but that's not the fault of journalism outlets, it's baked into the Constitution and other legislation.
<The real threat, the known threat to state security is Trump, because he and his followers are crazy. If the NSA, and other intelligence agencies, had any influence on the election...
Also, those news outlets may very well have their own agenda they're pushing, without any help from the intelligence agencies or anyone else: back in 2015, the media did help to make Hillary look bad. Perhaps they're blaming themselves partially for Trump getting elected, so this time around they want to make sure they don't turn off voters to the non-crazy candidate just because she isn't perfect. (And granted, Kamala doesn't have nearly as much baggage as Hillary did, which helps a lot.)
Meanwhile, Harris has made specific policy promises on a wide range of issues.[3]
[1] - https://www.piie.com/research/piie-charts/2024/trumps-bigger...
[2] - https://www.donaldjtrump.com/platform
[3] -https://www.cnn.com/interactive/2024/08/politics/kamala-harr...
For 99 44/100 percent of the online outrage bait, I'm like "you're not that interesting, and they almost certainly don't care about you anyway."
"The National Archives and Records Administration has standard procedures and approved vendors for this.[1] One of their approved vendors, Colorlab, has 1" type C equipment. Colorlab is conveniently located just outside the Capitol Beltway, about 20 miles west of NSA HQ at Fort Meade. Colorlab does preservation and conversion work for the Library of Congress, Warner Bros., Universal, NBC, The New York Public Library, Paramount, HBO, etc. NARA has a standard form for government agencies requesting this service.[3] It looks like it's not even charged against the sending agency - Archives picks up the bill."[1]
Maybe somebody got the message.
The zeitgeist of the time was shifting emphasis onto management (MBA type stuff) but the army had a saying; you can't manage a soldier into war, you lead them.
You manage things, you lead people.
Toxic waste was a highly relevant cultural phenomena at the time. I believe she was referencing the "Valley of the Drums" toxic waste site which was proposed as a superfund site in 12/82. Love Canal made the subject popular 5 years earlier.
For some reason I'm extremely interested in toxic waste. Anyways bit of reference.
Toxic waste is a real life monster that make for the best horror stories. Fictional monsters as in creatures aint got shit on real life willful poisoning of entire communities causing the suffering and deaths of millions. And its all done intentionally because someone wants more money - greed. The real monsters take on a dual form - the head of the beast being the people responsible and the body being the invisible poisons carelessly tossed onto the earth. Makes lovecraft and others look like mickey mouse.
For anyone interested, this is far and away the best book I've read on the subject of toxic waste. It became rare over the past 5 years. Used to be available on Open Library but I think they received a DMCA. Even the NYC public library only has one copy located at the main branch. Library wouldn't let me loan it out.
https://books.google.com/books/about/The_Road_to_Love_Canal....
This needs Torrent-seeded to death, in the public interest.-
The things that actually happened to real people, the things real people have done... they are just way more incredible and terrible than any one author is capable of dreaming up. The fact that they actually happened also gives them more heft...
"Have you checked to see if this area is sitting on EPA Superfund designated land, or down stream?" The response is usually a blank look.
So I ask them why is this area called Silicon Valley? Then I ask if they realize how incredibly toxic the solvents used in chip manufacturing are? And then I ask how much 1950s and 60s companies cared about environmental concerns? Most people connect the dots pretty quickly. "Holy shit." Is the usual response.
It really wouldn't surprise me if Fairchild, Intel and the rest just took barrels of used chemicals out back and dumped them into holes in the ground back when.
Google got hit by this a few years ago when they built an office building on top of toxic waste and now have to have 24/7 basement ventilation to make sure workers there don't get sick.
There are whole neighborhoods built on that same polluted land. I'll get my orange from Safeway, thanks.
It's definitely not just the Valley. I live in rural western Ohio. It was really eye-opening to see how much contamination there is even here, in a relatively sparsely-populated area. Once I knew the extent my feelings about local real estate changed dramatically. Everybody should research toxic sites in their area.
No doubt the manufacturing center in Dayton, OH, helped drive local contamination. I'm in a suburb 30 miles away in another county, however. We've got fun Superfund sites like the old county incinerator (PCB), two contaminated aquifers (tetrachloroethene and trichloroethylene) under the largest town in the County (from three sources, too!), and lead from a battery "recycler".
I simply can't understand the mentality earlier generations had re: environmental contamination. I hear it in my father (71) re: anthropogenic climate change ("I can't believe the activities of humans could change such a large system...") and I imagine similar sentiments were in the minds of people dumping PCB or lead into the ground. It's chilling to me.
Now, it's a fairly densely populated urban neighborhood. It was declared a superfund site in the 80s, and they're still monitoring the site and working on nearby properties for remediation, decades later. https://www.health.state.mn.us/communities/environment/hazar...
> I simply can't understand the mentality earlier generations had re: environmental contamination. I hear it in my father (71) re: anthropogenic climate change ("I can't believe the activities of humans could change such a large system...") and I imagine similar sentiments were in the minds of people dumping PCB or lead into the ground. It's chilling to me.
It is bonkers, but you still see it all the time. "Oh, my choice to drive a 10 MPG SUV to work every day doesn't matter. I'm only one person."
Some additional reading here: https://semspub.epa.gov/work/09/100018492.pdf
It's possible to build semiconductor devices without polluting the soil and water table, but it means every factory needs to build at least a small chemical waste processing plant onsite, or (better) design new closed-loop manufacturing processes that minimize or eliminate waste.
https://www.sourcengine.com/blog/growing-sustainability-effo...
The landfill was created in 1982 by the State of North Carolina as a place to dump contaminated soil as result of an illegal PCB dumping incident.
The "illegal PCB dumping incident" refers to dumping of PCB contaminated oil from the Ward Transformer Factory along the sides of highways in several (14) NC counties, back in 1978.
[1]: https://en.wikipedia.org/wiki/Warren_County_PCB_Landfill
She also mentions they were using computers to enhance satellite photos, it took 3 days to process but they could determine the height of waves in the middle of the pacific and the temperature 20 feet below the surface.
[1] The Bug in the Computer Bug Story
"I think it's rather nice that the Navy is keeping a few of the early artifacts like the first bug and me and a few other things."
:)
lovely, reminds me of an argentinian tv presenter that we make jokes about regarding her age (97 currently and going strong)
I wrote a little more about this: https://jkaptur.com/bugs/
She appeared on the David Letterman show a few years later in 1986, https://hackcur.io/grace-hopper-on-letterman/.
Some numbers from a few minutes of searching:
- Then: Intel 8021: 1 kB ROM, 64 B RAM, 11MHz, about $.40-.50 in 2024 dollars
- Now: ATtiny25: 1kB ROM, 128 B of RAM, 20 MHz, maybe $.70-$.80 each for a huge order
Not sure if this is the right comparison, and I'm sure there are lots of other differences that the topline numbers don't capture and that I don't know about (e.g, power consumption, instruction set, package size, etc. etc.)
The CH32V003 is $.10-$.20, quite usable in my experience and features a 16 kB ROM, 2 kB RAM and a 32-bit 48 MHz RISC-V.
The PMS150 is available for <$.05 with ~1.5 kB ROM, 60 B RAM and an 8-bit 8 MHz CPU.
If you're excluding chinese manufacturers, the STM32G030 is sub-$.80 in quantity for 32 kBs of ROM, 8 kBs of RAM and a 32-bit 64 MHz ARM CPU.
mmus are a performance hack; they make your memory-protected code run faster than if you use a jit compiler that inserts memory bounds checks. but suppose running code that way costs you a factor of 10× in performance. so maybe your 30-dhrystone-mips processor (https://www.lcsc.com/product-detail/Microcontrollers-MCU-MPU..., say, or https://www.lcsc.com/product-detail/Microcontrollers-MCU-MPU... for 2×) performs roughly like a 3 dhrystone mips processor. if we believe https://netlib.org/performance/html/dhrystone.data.col0.html that's roughly the performance of a sun 3/160, an amiga 2000, or a 40 megahertz amd clone 80386 pc (though much slower than an intel 386)
that's much faster than many multiuser machines i've used
the promise of java (and oberon) was that such large runtime overhead would be unnecessary with better static checking, and j2me and oberon seem to have largely borne that out
the bigger issue is i think that cheap microcontrollers don't have much ram or off-chip bandwidth. if you want a megabyte on-chip, you end up with things like https://www.digikey.com/en/products/detail/stmicroelectronic... (a 480-megahertz cortex-m7 with 128 kibibytes of flash and a mebibyte of ram for usd11.27), https://www.digikey.com/en/products/detail/infineon-technolo... (a dual-core 150-megahertz cortex-m0+ with 2 mebibytes of flash and a mebibyte of ram for usd12.92), https://www.digikey.com/en/products/detail/nxp-usa-inc/MIMXR... (a 600-megahertz cortex-m7 using external program memory and a mebibyte of ram for usd14.48), or https://www.digikey.com/en/products/detail/renesas-electroni... (a 240-megahertz renesas rx72n with 4 mebibytes of flash and a mebibyte of ram for usd20.85)
i'm pretty sure none of these have mmus either but i forget the rx architecture
a whole esp32 module like https://www.lcsc.com/product-detail/Development-Boards-Kits_... is cheaper and has more ram though. that one has 8 megs of psram and costs usd4.93
https://www.lcsc.com/product-detail/Microcontrollers-MCU-MPU... 1.5¢, 32 kibibytes in-application-programmable flash, 4 kibibytes sram, 48 megahertz, nearly 1 32-bit arm instruction per clock
https://jlcpcb.com/partdetail/NyquestTech-NY8A051H/C5143390 1.58¢, 1 kibiword otp prom, 48 bytes of ram, 20 megahertz, nearly 1 8-bit pic16-like instruction per clock, english datasheet https://www.nyquest.com.tw/upload/2024_02_293/NY8A051H_v1.6....
https://www.lcsc.com/product-detail/Microcontroller-Units-MC... 10.5¢, 1 kibiword otp prom, 60 bytes ram, two hardware threads ('fppa') context-switching every cycle so you can get better real-time response, 16 megahertz, nearly 1 8-bit instruction per clock. english datasheet https://www.padauk.com.tw/upload/doc/PMC251%20datasheet%20V0...
https://www.lcsc.com/product-detail/Microcontrollers-MCU-MPU... 8.7¢, 20 kibibytes flash, 3 kibibytes ram, 24 megahertz, nearly 1 32-bit arm instruction per clock. english datasheet https://download.py32.org/Datasheet/en/PY32F002A%C2%A0datash...
https://www.lcsc.com/product-detail/Microcontrollers-MCU-MPU... 12.45¢, 16 kibibytes of flash, 2 kibibytes sram, 24 megahertz, nearly 1 32-bit risc-v (rv32ec) instruction per clock
these are generally much lower power than the 8021, but really the place to look for power consumption is ambiq; these are all conventional cmos rather than the subthreshold logic ambiq uses
they also incorporate a lot more peripherals
https://en.m.wikipedia.org/wiki/Intel_MCS-48 says it's a cut-down 8048. the 8048 itself has a max clock speed 11 megahertz, 15 clocks per machine cycle, with about 70% of instructions taking one machine cycle, 30% taking two, for about half a mip. not vax mips, tho, an 8-bit mip. there's a manual for the chip family at https://manualsdump.com/en/download/manuals/intel-mcs-48/253...
the biggest omission is that they cut it down from 28 to 21 i/o pins and eliminated interrupts, which would be a big loss in a modern microcontroller, but it was nmos rather than cmos, so you're looking at unholy power consumption anyway; 40 milliamps (typ., p. 192/478 (6-49)) at 5 volts is 200 milliwatts. but they also cut the clock speed, the minimal machine cycle time on the 8021 is listed as 10 μs (with a 3 megahertz crystal) rather than the 8048's 2.5μs or the 8049's 1.36μs. so you get 0.07 8-bit mips. it uses dynamic logic and dynamic ram to save space so you can't clock it at less than 20% of that, so you can't do low-power sleep, ever
they also omitted the subtract instruction, which the 8048 doesn't have either. i guess you can use cpl, inc, add. (i see that what the manual suggests is cpl, add, cpl.) there's enough space for multiply and divide subroutines but they ain't gonna be fast
also, you can forget about programming an 8021's program memory. it's mask-programmable only; you have to do your test programming on an 8748 before you place your order with intel for a batch of 8021s with a custom silicon mask encoding your 1024 bytes of already tested and debugged firmware. so those 42-cent∆ prices were necessarily in rather large batches. the 8051 and 8048/8049 had an 'ea' pin you can pull high to get it to execute code from external memory instead, but i don't think the 8021 did; the manual says, "no external rom expansion capability is provided."
0.07 8-bit mips is about 0.005 dhrystone mips, although i think the 8021 is too small to run dhrystone. the cypress chip i linked above is about 60 dhrystone mips at 48 megahertz, so about 12000 times faster, for about a 25× lower price. it also has 4096 bytes of ram instead of 64 (16×), 32 kibibytes of nonvolatile program memory instead of 1 (32×), plus 8 kibibytes of rom, and you can program the flash in-application (i.e., under the control of the program it's running). it has a hardware multiplier, which is about a 10× additional speedup for dsp type stuff. in deep sleep it uses 2.5μamps. at full speed it's a bit of a power hog by current standards, slurping a hefty 13 milliamps (at 1.8 volts if you like, so 23 milliwatts, almost ⅛ the 8021, but if you're in deep sleep most of the time you can go another 5000× lower). and it's 1.6mm×2mm. and you can program it in c. it has only 9 gpio pins, though!
despite nominally being a psoc the cypress chip has no analog peripherals, not even a comparator. it does have internal oscillators (with no external components), pwm generation, i2c, spi, uart, and quadrature input, and its gpio pins have seven drive strength modes
so depending on whether cpu speed or memory space is the bigger bottleneck for your application, price/performance has improved between 400× and 300000× since the 8021. for things that were constrained by battery life or reprogrammability, the difference isn't quantitative, it's just that the 8021 couldn't do the job at all
in the metric of interest to hopper, though, which was computers per buck rather than mips per buck, it's only about 25× better than then
______
∆ https://data.bls.gov/cgi-bin/cpicalc.pl?cost1=.13&year1=1982... says 42¢
It wasn't until persistent internet connections became the norm, that this would have been shown to be an illusion. An illusion we continue to suffer for to this day.
https://news.ycombinator.com/item?id=40958494
Related:
The NSA Is Defeated by a 1950s Tape Recorder. Can You Help Them? -https://news.ycombinator.com/item?id=40957026 - July 2024 (24 comments)
Admiral Grace Hopper's landmark lecture is found, but the NSA won't release it -https://news.ycombinator.com/item?id=40926428 - July 2024 (9 comments)
> able to retrieve the footage contained on two 1’ APEX tapes
I'm no expert, but I think they meant 1-inch AMPEX tape.
Also, perhaps they should record it in doubly. #ThisIsSpinalTap #Stonehenge #InchsToFeet
p.s. I'm imagining ECHELON flagging your comment and someone from the NSA quickly modifying the document on the sly...
Hi, anonymous NSA employee!
Wow.
11.8" is a nanosecond.
So much wisdon and understanding:
"Get a molocule set, red balls can be computers, blue balls can be databases..."
"Get out of the domain of the paper, you cant draw in paper any more, they've got to be in three diminsions."
--
EDIT: @-probably_wrong
You have to think that she also has to put things into the minds of others who dont have here prescient forethought of systems.
Listen to all she says about "systems of computers" where "you need a computer to run all these other computers" and that you have ~160KB over head to run a 32KB program... andn how she breaks down all the constituents in a cluster, down to security auth...
And she tells you to buy a Scientific Molocule Model set to be able to design computer systems in 3d.
She is fn the foundation of cloud.
I wholeheartedly wish certain folks - us all, undoubtedly, but particularly certain people, such as Hopper - were inmortal.-
IMO the fundamental problem is that visualizing complex systems fails because the result is either too cumbersome to be useful or too simplified. UML tried to solve that issue (in 2D) by allowing you to go into more/less detail as you need it, and yet its adoption in modern software development is uneven at best. And there's a reason why we use flowcharts mostly for beginner's problems.
The actual solution, I believe, was getting away from visualizations by making robust software (well...) with clear interfaces to abstract the complexity away. Reaching this conclusion took a lot of work by plenty of brilliant minds, so I'm not faulting her for not being that accurate in that particular prediction.
I mean, at the time that she was saying this, you couldnt fill a Trump Rally with as many people on the planet at the time knew the future of compute the way she did.
The closest we have nowadays to her "3D flowcharts" idea would be UML in general [2] and Orthogonal State Machines [3] in particular, but I think that what her problem really needed was better encapsulation and interfaces between systems.
[1] https://en.wikipedia.org/wiki/Flowchart
[2] https://en.wikipedia.org/wiki/Unified_Modeling_Language
[3] https://en.wikipedia.org/wiki/UML_state_machine#Orthogonal_r...
> The seven-year-olds, of course, are learning BASIC, running the computers. I know one man that bought a computer and took it home, his son is teaching him BASIC. His son is seven. Of course I know another guy that took a computer home, now he has to apply to his three children for computer time."
> They're tremendously bright and they are out there, the brightest youngsters we have ever had."
As a GenXer who was a 10yo computer nerd programming BASIC in 1982, I'm proud to know that she was talking about myself and my peers!
But then she goes on to predict that the brightest kids will come from rural areas because they have "good schools". No idea where she got that idea from - I moved from a city to a rural area around that time and the education wasn't any better and my access to computers was totally gone. In my experience, most rural areas, especially in flyover states, neither had the money for a computer lab, nor the teachers that knew how to use them.
She is speaking from the perspective of someone leading a team of people that volunteered for the Navy and were then selected to do technical work for her team.
I think that given the years in which she lived, she would have certainly be seeing the side effect effects of the growing divide in opportunity between the urban and rural areas. A smart, hardworking kid in the city or suburbs is going to have a fulfilling job or go off to college at a higher rate than the equivalently talented and hardworking kid from farm country, where joining the military is probably the best chance at advancement that they're going to get.
That rural school system was a marked change from the one near Columbus AFB I had previously attended --- most students were from the base and the school received a generous amount of DoD funding to offset that, so all of the teachers had Masters degrees, and a number of them were accredited as faculty at a nearby college --- classes were strongly divided between social such as homeroom, P.E., social studies, &c. (attended at one's grade level) and academic (attended at one's grade level with a cap on 4 years ahead if in grade 8 or lower --- said cap was removed at 8th grade and students could begin taking college courses --- many graduated high school and were simultaneously awarded a 4 year college degree).
>attended at one's grade level with a cap on 4 years ahead if in grade 8 or lower
should be:
>attended at one's _ability_ level with a cap on 4 years ahead if in grade 8 or lower
And again, the wealthy area where I live now, if you are a bright kid that shows an interest in the military they are going to try to get you to do ROTC at the top engineering school possible, or setting up interviews with the Congressman or Senator to try to get a nomination into one of the academies.
So by percentages, the best enlisted men and women are likely to be from more rural areas.
What a peach. She inspired this jarhead.
And I mean this - amazing - she is at a level that few CEOs and public figures ever got. She’s got tremendous charm, intelligence and wit
Back to Grace Hopper:
What I lived most about the first presentation was her questions about “Valuing Information” and getting nothing but blank stares whenever she asked this question.
This is also a huge missed research domain in biomedical research. For me the vale of information (data utility) is a function of data connectivity potential. If I have one vector of data that connects to 1000 other vectors of data then that vector has much greater utility (all else equal) than onr vector that only connects on one other vector. The first vector is “smart” data that leads to valuable information, the second vector is “mute” data—just barely information at all.
Link on this general topic hat I hope GH would have enjoyed as much as I did her luminous and funny talk: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2751652/
Leland W. Slater, "Everything You Ever Wanted to Know About Microcomputers," Computop i cs . 3/82, pp. 38ff.
Slater, L.W. (1982). Everything you ever wanted to know about microcomputers (but didn’t know WHO to ask). Navy Regional Data Automation Center Publ., Norfolk, Virginia, 19 pp.
Leland W. Slater, "Everything You Ever Wanted to Know About Microcomputers," Computopics, 3/82, pp. 38ff.
(Computopics appears to have been the magazine of the Washington DC chapter of the ACM.)
Update: I have heard back and should have a scan of it soon.
In 2024 I can say this has not changed one bit. Even though we’ve gone through the Information Age, we still make decisions like we live in the Gilded Age.
I mention two things outside of social media, which is what most people think is the internet, about what I can do with a computer and people stare at me like I'm speaking alien languages. I come to hacker news and realize I'm not even 1% as smart as most of you.
A good video to watch. She's really funny. Really smart.
That aged pretty well.
Jump to a random timestamp. Listen for 3-4 seconds. Repeat it 10 times.
Quite impressive.
(As a part of my English language exam when I was a student, I had to write a text about a subject of my choice. I wrote about the history of programming languages: that's where I discovered Grace Hopper and mentioned her work in my essay).