Building the First GUIs
computer.rip
computer.rip
Peak "augmentation" seems to have been (believe it or not) the Microsoft era, complete with its "productivity suite" (except productivity has gone missing) and killer apps like the spreadsheet that is still a computing interface out of this world.
The demise of the MS monopoly came about by the exploitation of opportunities offered by interconnected computers with zero regard to lofty empowerment ideals. The role of the computing device has become to offer just enough attraction so as to extract value from the individual using it. Achieving that at scale has required the dumbification of the UI, removing any cognitive stress. Except, no pain - no gain.
If the hypothesis is correct a next generation of important (G)UI's may come about if and when the much delayed harnessing of the networked universe to augment individuals starts taking place. It might be driven by the challenges of the new stage we are in. Think of concepts to address information overload or managing online personas. Or it might facilitate emerging possibilities, eg interacting with local algorithmic agents etc.
I wish the expressiveness of tabular data representations was more obvious to more developers. All you have to do is sprinkle in a little bit of relational calculus and you can model literally anything as an xlsx document. Then, unlike virtually every other programming environment on earth, you can hand that document to any other non-technical human in the same business and they will immediately understand what you are trying to show.
9/10 times the spreadsheet will give the business what it wants, assuming you invest enough time in developing a decent one. The only reason you should add code on top of that is if you identify additional value-add with persistence, validations, systems integration, concurrent access, fancy UIs, etc.
On the other hand, I wish spreadsheets exposed a more developer-friendly feature set. It wouldn't take much: a sane language for formulas, some form of procedure/subroutine functionality with document or sheet scope, separation of semantic content and graphics.
Part of what's infuriating about Excel is how it's a local optimum not too far away from an even better local optimum. I could see myself actually using spreadsheets more, even professionally, but those poor things are so badly misused it's nauseating.
And even with just VBA, I've seen some impressive spreadsheet applications (in fact, some even were versioned and had official SDLC process in some companies).
Assume you had a few worksheets and you want to produce a new one based on some projection of those. You could press some hotkey "New Worksheet from Query..." and then type in the SELECT you want to use for the projection. The column headings could be optionally specified with the appropriate SQL syntax. The schema for the internal SQL dialect could be dynamically generated by inferring table name from sheet name, and column name from the first row in each. You could even have this re-evaluate in real-time, so any changes to base sheets would instantly update the projected sheets.
VBA, COM and things like that rest upon the assumption that the worksheet is just a complicated data structure like the DOM.
Want to take Excel seriously? Let's commit to that. Take sheets as the computational model and build from there. A natural operational semantics for them is a graph rewriting system, same as ML and Miranda.
I'd argue that GUI evolution has all but stagnated: The shift from specialized tools for individuals to mass adoption brought us much more new ways of interacting with our devices. The adoption of personal computing devices in the last several years has been unprecedented. We'd still be reading paper newspapers, watching TV and using the landline network for communication if it were not for the recent GUI evolution.
The main problem I see is the trend to remove power user features altogether instead of having them be hidden in the settings. Ease of use and customizability can get along just well, but are rarely to be found. So your statement about augmentation is right anyway.
Edit: I think that different OSes, different devices for different target audiences is not an actual problem. The uses of a computer can be so vast. The trend to integrate smart devices has only begun and requires different usability approaches. We should not compare Excel UX to Tinder UX
I'd say peak augmentation came with having Wikipedia available in my pocket or from a smart speaker. There's still so much more that could be done. If anything, word processors and spreadsheets are a local maximum and I think we are probably stuck on another one now.
No, that would be Atari STs GEM OS, or the Amiga’s AmigaOS. Or the Acorn Electron. Windows was laughing joke amongst non-PC users for the entirety of the 80s and the first half of the 90s. Even Tandy produced a better frontend to DOS than the early versions of Windows.
> It was more that early releases of Windows failed because they were inferior to other DOS GUIs.
I think the author agrees with you, they’re just noting that this is how many people today (including myself, prior to having read the article) perceive the evolution of GUIs.
That said, you are absolutely correct that there were multiple competing microcomputer platforms at the time and most of them had GUIs on offer.
The PC world would have been a far more interesting scene if IBM had consulted with or purchased technology from someone like Atari or Commodore.
Don't get me started on audio.
IBM really blew it. Arguably, we're lucky they did, given the market opportunities that they created as a result.
Is it only because I myself am building an alternate OS interface or has this become a popular subject of conversation?
And then the designers came and threw away so much of that knowledge because they wanted to make something that looks good in screenshots. Visible controls are ugly they said, so all buttons and controls get hidden away under some nondescript icon that doesn't even look like an interactive element (maybe 3 lines in the corner or something). Scrollbars are uncool so now they disappear immediately and you just have to know where to look for them. White space is visually pleasing so just fill the screen with it. Old rules about making buttons look like buttons are completely ignored so the screenshot is pretty. Sure you have to hunt around like a blind person in an alley trying to make things actually work, but it sure looks nice doesn't it?
Ad-tech and similarly myopic business pressures are frequently at odds with usability, and those business pressures ultimately determine what solution makes it to (or remains in) production. Less usable websites and tools are frequently "better performing" when measured by a metric leadership is beholden to, and usability is easily sacrificed if faith in its value doesn't extend to the very top of a company.
I know this is HN and it's fun to pick on incompetent designers, but there are a lot of parallels with the lack of care for security that frequently tops the news on here. Poor usability often reflects company values, not designer capability.
>there are a lot of parallels with the lack of care for security
Absolutely, 100%. Same idea.
We still have much further to go. My main display doesn't have a single UI element on it. It's a terminal in full screen mode. Except devoting 100% of the display to content wasn't enough for me. I wanted a third dimension of content. So I modified the terminal software so that I could control transparency by shift + scroll wheeling which means I can play movies in VLC running in full screen mode underneath the full screen terminal. That's my design. Terminal and television. The most important use cases of the most important devices of the twentieth century.
Not everyone knows that, I don't even think to try it unless there's a corresponding control or menu item. It's also not universal, outlook is one of the most heavily used GUIs in the world and ctrl-f means forward in that. Then you've got the ubiquity of touch screens where ctrl-f isn't even an option.
> They were probably formally studying people who had never seen or used a GUI before
Everyone's seen a GUI, but that doesn't mean they've seen the particular GUI they're trying to use at any given time.
I think it is true that people learn and we can assume other things now. But still there must be discoverability of such features.
I long for the day when more elaborate and luxurious art-deco-based interfaces become fashionable in UI design.
Building shitty UI because it looks cool in your marketing materials is doing a criminal disservice to your users and ultimately your business.
This was made in 1990, sponsored by the ACM CHI 1990 conference, to tell the history of widgets up until then. Previously published as: Brad A. Myers. All the Widgets. 2 hour, 15 min videotape. Technical Video Program of the SIGCHI'90 conference, Seattle, WA. April 1-4, 1990. SIGGRAPH Video Review, Issue 57. ISBN 0-89791-930-0.
https://www.youtube.com/watch?v=9qtd8Hc90Hw
Also by Brad Myers:
Taxonomies of Visual Programming (1990) [pdf] (cmu.edu)
https://news.ycombinator.com/item?id=26057530
https://www.cs.cmu.edu/~bam/papers/VLtax2-jvlc-1990.pdf
Updated version:
And they are all ultimately derived from IBM CUA, so that's a good start: https://en.wikipedia.org/wiki/IBM_Common_User_Access
Yes. It was called "web UI".
This brought a level of casualness and lack of rigor, that infected native UIs as well.
In the 90s there was a lot of GUI research, including by major companies like Microsoft and Apple, and lots of innovations.
But even more importantly, the standard practices of the main OS makers, favored a uniform look, with clear affordances (e.g. button bevels), and so on. The apex of with would be something like Windows 2000, Mac OS 8, BeOS, NeXT, and the like.
Stuff like "mystery meat navigation", "hamburger menus", "invisible scrollbars", huge padding, buttons that look like text, and so on, including "let's make the whole desktop app interface in the DOM" where post-2000 additions, and not for the better.
In some ways xfce4 reminds me a lot of windows 2000.
Most of the apps would be an assortment of (non Xfce) Gnome-inspired GTK, Electron (e.g. Slack, VSCode), custom (Firefox, Chrome, Sublime Text), and/or KDE stuff, each with its own warts and "modern" approaches...
People say that CLIs are composable and GUIs aren't, and so CLIs are better for power users. They're right, for today's GUIs, but yesterday's GUIs were composable. We've lost that.
Imagine where we could be today if this was our starting point in the 70s.
I myself am working on a web-based note-taking system which works with every browser since Mosaic. :)
Well, there are two conflated problems here:
1) The GUI and interacting with humans, itself.
I don't think we've even scratched the surface here. I'm not a big fan of the Instagram/TikTok "HAH! Everything is hidden, Boomer! If you aren't spending 5 hours a day with our application talking to your friends how to use our application you'll never figure it out." However, at least some stuff is different.
2) The software implementation of the GUI underneath
I do think we've gotten lost in a local extremum here. The whole "The GUI must run on the main thread" when we've got gajillions of cores spread across the CPU and GPU is a gigantic problem. We really need a multithreaded UI implementation.
Younger generations don't feel like these features are hidden at all, and wouldn't feel like it's any different from most new apps. What Gen Z are very good at is discoverability inside apps using things like swiping. A younger person will interact with the app by swiping up and down or left or right, where older people might be more trained for button pressing. Having spoken to some younger people, they find some apps like Facebook harder to use, since there are too many buttons, and not enough contextual motion.
From watching the younger generation, I disagree.
There is no app discoverability or consistency at all. The first thing they do is go watch the video about how to do X in the app.
Sure, being younger, they can triple tap and multi-finger rotate better than older people. That doesn't make the GUI better.
I would love to see circular menus in GUI interaction on tablets. Sadly, only games seem to think they're a good idea.
And now that I think about it some more, I have a VR headset, and the UIs are often more complicated than swiping on your phone. It's more like pointing at buttons, menus and links with your controller. That's typically the case for gaming consoles as well.
As such, I'm not sure phone apps are still the latest in UI design, it's more that the smaller screen touch interface and apps that typically do one thing well better fit the form factor of smart devices. Granted, the younger generation use their phones for everything, until they put on a VR headset or have to work in a desktop-like program.
Sketchpad by Ivan Sutherland was released in 1963, see e.g. https://en.wikipedia.org/wiki/Sketchpad. The NORAD/SAGE display systems which had light pens as pointing devices were even older.
The earliest “overlapping windows” system was, iirc, the smalltalk environment.
I say “environment” because that back then you booted the machines (Alto and later the D machines) into a completely hermetic environment; Smalltalk, Interlisp-D, and Cedar/Mesa ran on the bare hardware (had their own networking stacks and drivers and even had their own microcode. There was no meaningful distinction between OS and application.
My memory is a little fuzzy as this was quite a while ago and there was a lot of experimentation going on all over the place. I left PARC to go back to MIT in 1985.
I assume you know these interesting presentations: https://www.youtube.com/watch?v=2Z43y94Dfzk, https://www.youtube.com/watch?v=uknEhXyZgsg
EDIT: see also e.g. https://computerhistory.org/blog/introducing-the-smalltalk-z...
[1] http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.656...
https://en.wikipedia.org/wiki/History_of_the_graphical_user_...
Here's his presentation of his prototype, in 1982, quite a while before the Mac or Lisa were released:
This has, in practice, proven to be basically wrong, yes.
But I giggled reading it because I'm absolutely addicted to trackpoint based keyboards specifically to avoid needing to move my hands (for both my convenience and the sake of my wrists).
Edit: ... and now I note that there's a footnote saying the author shares said preference.
One of the reasons is that the first Intel processor to support multi-tasking was 386. When DOS was created, there was no hardware support for multiprocessing and virtual memory.
For designs that wanted an MMU, you could use an external one. Sun had external MMUs in their early workstations. Even Apple had their own MMU in the Macintosh II line before they switched over to the 68030.
You could add a 68851 to a Macintosh II, I didn't think Apple developed their own MMU for it. The Lisa had a custom MMU.
"The machine shipped with a socket for an MMU, but the 'Apple HMMU Chip' (VLSI VI475 chip) was installed that did not implement virtual memory (instead, it translated 24-bit addresses to 32-bit addresses for the Mac OS, which would not be 32-bit clean until System 7)."
Not quite. The 80286 was not a good CPU, but it could do multitasking:
https://www.cpu-world.com/CPUs/80286/index.html
> The second generation of x86 16-bit processors, Intel 80286, was released in 1982. The major new feature of the 80286 microprocessor was protected mode. When switched to this mode, the CPU could address up to 16 MB of operating memory (previous generation of 8086/8088 microprocessors was limited to 1 MB). In the protected mode it was possible to protect memory and other system resources from user programs - this feature was necessary for real program multitasking. There were many operating systems that utilized the 80286 protected mode: OS/2 1.x, Venix, SCO Xenix 286, and others. While this mode was useful for multitasking operating systems, it was of limited use for systems that required execution of existing x86 programs. The protected mode couldn't run multiple virtual 8086 programs, and had other limitations as well:
> 80286 was a 16-bit microprocessor. Although in protected mode the CPU could address up to 16 MB of memory, this was implemented using memory segments. Maximum size of memory segment was still 64 KB.
> There was no fast and reliable way to switch back to real mode from protected mode.
More about OS/2 on the 80286:
https://www.landley.net/history/mirror/os2/history/os210/ind...
From the article:
> Despite the best efforts of many dweebs, one handed text entry has never caught on [2].
One-handed text entry may not have, but putting one hand (usually your left) on the keyboard and one on the mouse, and using your voice for text related communication has taken over PC gaming. I wonder what an equivalent productivity app (or set of emacs modes) would look like.
[1] https://www.youtube.com/watch?v=yJDv-zdhzMY&ab_channel=Marce...
The answer is never. The reason is that different kinds of UIs are effective for different kinds of tasks.
I can't remember who I heard this from, but they asked the audience to imagine a speech-based steering wheel for cars. "Turn left... more... more... MORE!!! LESS! LESS!"
Speech also has problems with sensitive data in public places ("Please enter in your password" or "Your stock portfolio is now worth ..."), has problems with noise, and can also interfere with some parts of your cognitive processes. In addition, speech is serial, which means that you can pretty much only process one stream at a time. Contrast this with a web page that you can skim, or a visualization that can show a lot of data.
Different kinds of interactions have different tradeoffs. Speech is good for some tasks, and not very good for others.
Star Trek fans probably would know the correct answer: don’t use a steering wheel; that’s just a means towards a goal, not the goal itself. Everybody who ever used a taxi could have figured out the proper speech UI for a car, too: “Take me to the airport”/“follow that car”/etc.
We don’t have touch-based ovens where you prod the fire with your bare fingers or cars where you have to tap a screen to inject gasoline at just the right time, either.
Actually, the ideal UI is like a good butler: you don’t even have to tell him what you want; he just knows.
But yes, the answer likely will be ‘never’.
http://worrydream.com/ABriefRantOnTheFutureOfInteractionDesi...
As long as there are minds there will be text.
As long as text can be complex we will have to consume it in blocks and probably visually (or the direct brain interface equivalent)
I believe you are correct, in a sense, that as long as there are minds there will be text. However, I believe it should be considered from a slightly different angle. As long as there are minds there will be language. Text is one representation of that and it is a representation we can manipulate easily computers. This makes text a powerful and simple medium that can be used pretty much everywhere (f.g. Unix-like systems; networking protocols; data storage).
Language is universal. Intelligent entities will create and use language to communicate with other entities or with themselves (i.e. sticky note). We have been proven capable of learning language even in some of the most input restrictive states, such as the case of Helen Keller. Language can be seen, heard, or felt. I wouldn’t be surprised if some creatives come up with a way to smell or taste it. Computers can be read, heard, or felt (f.g. Braille terminals). Text is a useful medium for storing the basic information that is then transformed into the output most suitable for the final representation given to the end user.
Let me bring this back around to your post. Because of the aforementioned reasoning, while I generally agree with your post I do disagree that text must be consumed visually or in blocks. The information contained in text should be transmitted in whatever way fits the user best. I realize this is vague, but I think this a starting point to thinking further on how to best get information to the user. Personally, I've come to the conclusion that a flexible core application should send and receive inputs and outputs to a seperate interface application. This may take the form of a typical client-server style system. This way, the interface can be more easily to the user's interfacing preferences and the core application only concerns itself with its task. Large blocks of text may not be well suited in certain instances, such as the case with screen readers or Braille terminals.
I'm trying not to be too general or vague here. I don't feel I've presented enough to really provide any actionable insight here. But I see this post is getting long, and I will leave it at this for now as potential food for thought for yourself or others. It's a concept I'm still thinking about how to best implement, and the details of implementation across various OS's and devices complicate matters.
Not in our lifetimes.
People have been promising speech recognition since at least the late 70's. Today's "digital assistants" are little more than parlor tricks, on those occasions when they work.
They don't seem much more accurate than my old Covox Voicemaster.