Gio – write immediate-mode GUI programs in Go
gioui.org
gioui.org
The 'looks native' thing is mostly noise
From the positive side, I am glad more and more hackers on HN focus on that matter. Also worth noting big players on the market really do take that point seriously, e.g. Flutter has a Semantics tag [2].
[0]: https://github.com/rxi/lite
[1]: https://news.ycombinator.com/item?id=26223380
[2]: https://api.flutter.dev/flutter/widgets/Semantics-class.html
Which it is. Apple has been praised for the accessibility of both macOS and iOS for years if not decades, accessibility concerns have been baked into Cocoa forever. And accessibility is literally a top level category of the Settings app.
And I know that Windows has been improving by leaps and bounds.
> The OS can, for example, run OCR on the entire screen and turn bitmaps into selectable text, read text out loud, etc
That exists (look up VoiceOver Recognition). However it can not be reliable, and will never be anywhere near as good as actual semantic annotation.
Image recognition has no way to understand that the physical UI layout has no relation to its logical setup, nor does it have any way to differentiate between semantic and decorative content, or to see through invisibility to know that a button is a menu versus an action.
Usually it means that you need a deferred mode semantic tree of the application for accessibility UI to walk through, though.
Unity has no official support for screenreaders or colorblind modes. Unreal has both. If an major game engine developer like Unity cannot be bothered to add support, can we really expect the little developer hacking away on some OpenGL side project to?
Most people working with OpenGL use GLFW (or sometimes SDL), neither have any cross-platform API for supporting screen readers. Why? Because not only do Windows and MacOS have different APIs, but different screen readers, braille displays, etc. have different APIs.
But okay - let's look at the web. It's the most accessible platform of them all, right? What have Google/Apple/Mozilla done to support developers adding screenreader support from WebGL and Canvas-based applications? -> not a whole lot. You need to inject text into a hidden div, which also has performance implications so you need to build a UI to toggle it on/off or detect accessibility the way Flutter Web does.
We should have better accessibility, but as long as we expect developers to run through several hours/days of hoops to get even something basic working - it's just not going to happen, and that is mostly the fault of extremely poor platform APIs.
Games are special because the means of interaction is often a core part of the experience. An accessibility mode for a game can be a huge amount of work. Not to say that some people don't do that huge amount of work (which is great) but games lacking accessibility is no excuse to use game-like development practices to write apps that could otherwise be relatively easily made accessibile.
As a former member of the Windows accessibility team at Microsoft, I'd appreciate your thoughts on what makes the platform accessibility APIs extremely poor. I have my own thoughts on what makes them difficult to implement, but I'd like to hear your perspective first.
I've been thinking about this for a while. Below is an edited excerpt from a message that I wrote to a colleague a few months ago, about how I think I should go about such a project:
---
I've been thinking about the problem of multiple programming languages, and even multiple programming styles within the same language. Do we implement a C library that would make hard-core Unix and Linux folks happy? A library in a subset of C++ that some developers of games and graphics would be comfortable with? A more modern style of C++ that would make some other developers happy? And where would that leave folks working in Java, or C# like the Unity crowd? Of course, different platforms have varying levels of support for different languages. And there are more languages coming into popularity or on the horizon, like Swift and Rust.
So I think what we really want is a cross-platform message format or protocol, with multiple implementations: multiple providers on the application/toolkit side, and multiple client libraries on the platform side. The client libraries could be written in each platform's native language (e.g. C++ for Windows, Java for Android, Swift for Apple platforms, or JavaScript for the web), and separately, provider libraries could be implemented in multiple languages. All major programming languages can work with binary data buffers. And since none of us want to work directly with those raw bytes, we could use an existing standard like Google Protocol Buffers, which already has implementations in several languages. Initially, for desktop and mobile platforms as well as web applications, this protocol would just be used internally between components in the same application process. But I can dream about platforms themselves adopting the protocol someday. There would have to be glue layers between cross-platform providers and platform-specific clients, but if we design this right, the glue could be kept pretty thin, with most of the complexity being kept on one side or the other, so we don't have a multiplication of effort for n toolkits or applications on m platforms.
I think the protocol should be push-based, rather than pull-based like Windows UI Automation and some other accessibility APIs. That is, the application/toolkit would push full information about objects in the accessibility tree when it first creates that tree, and incremental updates when objects are created or destroyed, when properties change, when text content changes, etc. That's probably the only thing that's going to work for the web platform, and I think it's a model that other platforms would do well to adopt. (It's probably too late for Windows, but I dream of replacing the current accessibility model on desktop Linux someday, after my non-compete with Microsoft expires.) The challenge with a push model is making it efficient; we don't want to re-push the whole contents of a large text box when the user types a single character, and when we do need to push all contents, we need to do so efficiently. And sometimes we really do need to push a lot of information. For example, for a text box, we need the screen coordinates of every character, plus all of the boundaries between words.
Luckily, a push-based accessibility architecture has been done before, internally in the Chromium browser engine. As you probably know, Chromium has a multi-process model, where web pages are rendered in sandboxed processes, which communicate with a master browser process that interfaces with the OS. The browser process is not allowed to do blocking IPC requests into the renderer processes, so it can't implement synchronous, pull-based accessibility APIs like UIA in the obvious way. So the Chromium team implemented a protocol where the renderer processes push their accessibility trees, and incremental updates to those trees, over to the browser process, which can then store the trees in memory and then provide information to UIA or other APIs on request. The pushed trees are comprehensive, including the information I mentioned about text. Chromium does this using its own binary protocol called Mojo, which is kind of like Protocol Buffers but strongly tied to Chromium. So, while I'll take design inspiration from Chromium, I probably won't take the actual protocol or code.
I also want my protocol to scale down to embedded platforms. Accessibility on devices running embedded software (i.e. not Windows or another general-purpose OS) is basically an unsolved problem; as far as I know, device makers have to implement their own custom self-voicing interfaces, if they do anything about the problem at all. But imagine a standard where a user can pull out their smartphone, connect to the specialized device over Bluetooth or WiFi, and get an accessible interface to the device on their phone. Yeah, I'm swinging for the fence with this project. To pull this off, I think the protocol would need to be not only push-based but streaming, allowing providers to send out accessibility information without having to build up and maintain much extra state in memory. This would also help with the immediate-mode GUIs that are used in some games and game development tools. If we design the protocol right, these immediate-mode toolkits should be able to push out accessibility information and events at the same time that they're making OpenGL (or similar) calls to draw the current frame, again without having to hold much extra state in memory, which is something that these toolkits try to avoid.
The Linux desktop environments could really use more people working on the accessibility stack, it's really outdated at the moment.
Worse still, while the Gio layout engine is pretty good at culling non-visible stuff, ultimately the framework doesn't know what will be visible until we're in GPU land, so there would have to be a ton of bookkeeping added to test if, for example, some text was clipped or if things were overlapping.
In practice, the best way to model things would be to just allow the developer to do something equivalent to pushing fields full of ARIA information like in the browser. But now you're losing all the advantages of the immediate mode and layout engine, etc.
It's a lose/lose situation, unfortunately.
I think the real solution is somewhere in the middle. I think screen readers need to grow some OCR abilities, and libraries like Gio need to learn some better navigation tech, like supporting tab, keyboard nav, etc. Something like that.
Just IMO, it's unrealistic to expect to have an accessible GUI that stores no state, does none of that type of bookkeeping, has no OOP model, and only outputs to pixels. The point with these type of assistive technologies is that the user can't work with that type of visual data, they need more state presented to them. Sure, it makes everything easier to develop when you cut it all out and only use immediate mode, but that's exactly the problem: everything else has been cut out, on purpose.
https://machinelearning.apple.com/research/creating-accessib...
Forgive my ignorance, but what should happen here? Let say I have a scrollable text pane with an entire novel in it.
I agree there should still be a mechanism to pass more accessibility data than OCR could extract. What is the bare minimum information that is required? Pretend there is no UX model -- arbitrary things could be presented (like in a video game).
In particular, with Gio there isn't necessarily a single set of widgets or UX. Gio has a small set of Material Design compliant widgets, but is mostly a library for composing immediate mode graphics. For example I have several custom widgets, some that only interact with the keyboard, some that don't have any text at all (just graphics or animations), one is a akin to a 2D-scrollable map, click-to-drag Google maps style. I'm not really sure where I'd begin in making these accessible. How should something like Google Earth be made accessible, ideally?
I'm not positive what you want them to provide-- some kind of "fake GUI?" I might not be imaginative enough, but I can't imagine how that would work.
Also, test in high contrast modes, and screen magnifiers.
2. https://www.freedomscientific.com/products/software/jaws/
Windows also has the Narrator screen reader built in, and in Windows 10, Narrator is a decent screen reader IMO. (Disclosure: I used to be on the Windows accessibility team at Microsoft, where I worked on Narrator.) To enable Narrator, press Ctrl+Win+Enter. Narrator has a built-in tutorial (not written by me) to help you get started.
Mac and iOS has VoiceOver built in. To enable it on Mac, press Command+F5. On iOS, you may be able to triple-press the home button, on devices that have one. Failing that, you can find it in Settings.
Android has TalkBack built in. You can find it somewhere in Settings; the specifics are device-dependent.
It's not at all clear to me that the right approach is to have both the accessibility interface and the GUI be generated from the same components. It seems to me that it's just approaching the lowest common denominator of a passable GUI and a passable accessibility interface.
The traditional widget-style GUI paradigm asks you to essentially copy all your application data into their format and to keep this format in sync with your own data. This is tedious, and I believe hinders the creation of better GUIs. On the flip-side, owning your data structure means it can indeed hook into the accessibility API for you.
Immediate mode GUIs provide an alternative approach, where they only handle rendering stuff on screen. Since all the data structures are now properly controlled by the programmer, it is much easier to create dynamic GUIs. Taking a step back like this provides an opportunity to further the state of the art.
The solution I see to providing accessibility in immediate mode GUIs is that the programmer has to hook into the accessibility API themselves -- if it makes sense. It's not clear to me that certain advanced GUI programs can ever be very accessible. Yes, this is more work, but there's also an opportunity to create richer experiences for people with disabilities.
I can only talk about my own motivation for repeatedly bringing this up. If a developer is unaware that they're choosing an inaccessible GUI toolkit for their application, and that application is then required for a particular job, then that developer may end up unwittingly preventing some people from doing that job. And this isn't just hypothetical for me; I know of a blind person who lost his job (luckily only temporarily) because of an inaccessible application. So I think it's entirely reasonable for an application developer to dismiss a GUI toolkit because it's inaccessible. I wish more developers would research this on their own before choosing a toolkit for their applications. But since many don't, I feel obligated to draw attention to it. I'm glad I'm not the only one.
As for your suggestion that applications should implement the platform accessibility APIs themselves: First, it's impractical to expect application-level code to do this directly. The Windows accessibility APIs are notoriously hard to implement correctly. And my understanding is that the Mac, iOS, and Android accessibility APIs are only easy to implement if the application is written in the platform's native programming language. Now, these problems could be solved with an open-source wrapper library -- the SDL or GLFW of accessibility. And I'm planning to implement such a library myself sometime. But my plan is for that library to be used by toolkits, not applications. Because if every application has to directly implement accessibility in parallel with its GUI, then it's a safe bet that even fewer applications will be accessible. We want to get to a place where as many applications as possible can be accessible by default, with minimal effort on the part of the application developers.
A graphics application should be able to submit draw calls or queues to a OS subsystem that, for example, generates a comparable audio space, performs colorspace/size transformations for improving visibility, etc. Developers should just be responsible for plugging in to these systems, and the systems themselves should be as standardized as OpenGL, Vulkan, DirectX.
I don't know if there is a "Khronos Group" for Accessibility Standards. Maybe there is.
I suspect most Devs would never link up and keep up to date an alternative accessibility interface. Also there is a lot of benefit to the accessable and standard interface being close, to make it easy to cooperate.
My opinion for the last several years has been, if you use these toolkits, you might just have to accept that your app is not going to be accessible and is not really suitable for anything beyond highly specialized use.
It's more accurate to say that the vast majority of GUI toolkits that have ever been written are inaccessible, and most OpenGL-based toolkits fall into this category. Accessibility is tedious to implement at the toolkit level, especially in a cross-platform toolkit, so generally only big toolkits with corporate backing get it.
Also, on the web platform, if you abandon semantic HTML and use canvas or WebGL, then you're forfeiting the accessibility that your UI would normally have by default. As far as I know, the only way to make a canvas or WebGL-based UI accessible is to construct a parallel HTML DOM tree (and presumably use CSS to hide it behind the canvas).
Thanks for being curious about this.
The same valid comment was made when nuklear was discussed a few months back.
Many would love to unshackle themselves from the limitations of working with UI toolkits and it’s increasingly difficult/impossible to fulfill all the requirements we place on software, especially for solo developers. I don’t know how many one off tools I’ve written where the only interface is the command line over the years. It would be nice to have a simple step up from that as offered by Imgui tools.
Triple click on the text. Expected whole paragraph to be selected.
In input field click Cmd+A to select all text. Does not work as expected on mac.
Copy paste does not work.
Start selecting text but release mouse outside of the widget. Text selecting gets stuck.
Text navigation in text field does not work Cmd - left/right.
And these are from 5 minutes of playing with it.
ps: im on mac+safari
Double click to select a word and triple click to select paragraph doesn't work for me either but can't reproduce some of the other issues running it natively.
My pet peeve with most "home-grown" GUI frameworks is relatively subtle and probably hard to implement, but it's an example why developing a GUI framework (especially a platform-independent one) from scratch is hard: no subpixel rendering for text (only "classic" antialiasing). This makes these UIs look strange on a "normal-DPI" monitor (which are getting rarer these days, but I still have some and I'm not gonna throw them away anytime soon).
I do think it would be interesting to define a minimal set of user interface conventions that would forego a lot of features but would allow developers to stand a fighting chance of building a UI toolkit. If it became popular enough, users could come to understand its conventions even if they aren’t platform native (not ideal, but something has to change).
Has anybody carried out a larger traditional (i.e. non 3D focussed) application with an immediate mode GUI toolkit?
In many ways immediate mode UI's is a relative of declarative UI's like React and I think we might see hybrids or implementations of either on top of the other.
Why? Right now immediate mode UI's often lack accessibility,etc and since it's often game/graphics focused there has been little push to fix this afaik. But there really is little reason not to implement an im-gui that renders to a diff-able tree that then is translated to "Native UI components".
Imo one big reason why declarative UI's hasn't caught on is because C/C++ really don't have the best ergonomics for building "flexible" data structures or immutable updates, by turning it around to using im-gui style rendering it could probably be solved fairly smoothly.
https://games.greggman.com/game/imhui-first-thoughts/
The term "immediate mode" is only about the API, not about how things are implemented under the hood, e.g. an immediate mode UI can map to traditional stateful UIs, it's just not done yet very often because of ImGui's roots in the gamedev world (where it is important to easily integrate with existing game rendering engines).
Still the fact is that this is the only one I know of and in general accessibility hasn't been a big concern so far (And the imgui name isn't an accident considering the history).
Anyhow considering how laborious "modern" API's like UWP is, how badly declarative API's continue to map to C++ and the frequency at which gamedev rooted developers are asking for how to do "native UI"'s in C++ the time is definitely ripe for a serious one to emerge, the question is though if there will be one with enough momentum to become truly useful or if we'll end up with a bunch of small hacks that kinda covers some bases but don't end up being a good option for the majority.
I suspect it has more to do with the sheer amount of effort to build a new UI toolkit (text rendering itself is an insane amount of work) and so very few UI toolkits actually make it to a usable MVP whether they are traditional/widget-based or reactive.
As someone coming from gamedev, I'd prefer something like this to be few files (maybe even single-file-header lib) that mostly wraps win32,osX,etc core API's for a native experience with good accessibility,etc support. My declarative-style experiment was maybe 6-7 files medium sized files (and I think the implementation ran away with C++ bloat and complexity trying to get the API to mesh updates with C++)
On the not 100% native side there’s Fyne which seems to only depend on xorg development headers. It can do a lot.
https://caseymuratori.com/blog_0001
http://sol.gfxile.net/imgui/
http://www.johno.se/book/imgui.html
https://github.com/ocornut/imguiBecause searching "gio menuitem" returns a lot more pizza than development docs :). Prefixing my search request with glib (which gio is a part of) solved that for me.
It'll get there, with time, but for now I'd avoid it for anything serious.
Tinygo is a bit better in this regard.
&errors.errorString{s:"no support for OpenGL ES 3 nor EXT_sRGB"}
Old computers with GL 2.1 adapters are fucked.(I'm not claiming it's the best, just that it works well. Technically, I'd still prefer Ada's extreme modularity with explicit interface definitions.)
Which they failed to meet. There is nothing special in golang that is designed for large codebases. The opposite in fact, things like how they designed interfaces is counter to large code bases. There are much better languages that do well in large code bases.
That's true (so far).
> and its language features make it very poor to work with for large code bases.
Not in my experience. We have quite a large codebase at work and we implement new features just fine. At least it looks better than in my previous C++ projects, where after a few years the code structure became a big rigid.
See how much money might already been spent into Flutter development.
But in my eyes, Go+Gtk should be a great combination for the desktop, as the bindings have a nice API and Go is really approachable for newcomers - it was one of the design criteria to not have too high learning efforts. If anything, there could be more documentation on how to use Gtk. And so far, my Go+Gtk programs can be easily recompiled and run on the Mac :)
Just like Win32 object model, is based on message dispatch across windows classes, not necessarly inheritance.
Just like game engines based on component architectures.