Sharing Screen with GPT 4 vision
loom.com
loom.com
Photoshop? KiCad? Final Cut Pro? Blender wins as the app that I have struggled the most to master.
The excavator is a good analogy. The best-designed most graceful excavator will still be _hard_ (at least for more complex tasks).
There are 14 year olds making great use of it [1].
[1] https://variety.com/2023/film/news/spider-man-across-the-spi...
Blender is only "not intuitive" if you already learned to use a different program in the same or adjacent domain (say some CAD tool, or Unreal Editor, or even Paint 3D or Sketchup), as it's not going to be similar enough. Similarity to existing software is a good thing, but not worth it if you can offer much better ergonomics otherwise. Blender could, and did.
Blender "shortcuts" are most of the time one key, followed by others, alternative variations can be achieved with modifier keys (makes sense).
It takes very little time to see what the shortcuts are. Most of the time you can just hover by a tool icon, other times, menus have them clearly labelled by item.
At this point, you know the basics. Your human eyes are very good at perceiving changes in your FoV, and after inputting a key, the status bar presents information on alternatives modes or filters available for that tool.
Eventually, you start assuming (correctly) that other tools behave in the same fashion for a variety of things.
For example, you press "Y" (y-axis) after "S" (scale), and you scale on that axis! If you followed that by a number, you scale by that scale factor! And this can be applied to every other tool. Moreover, this and other combos make sense, are easy to understand. Whatever you may imagine as the effect they have on other tools is most likely the exact outcome.
Blender is very sane, it does exactly what you tell it to do. You do have to make the call, but when that is at a distance of a finger, it isn't an issue.
You can learn by just using it. Information is laid out to you clearly. Modifiers/filters are consistent, making their knowledge easily transferrable to different tools.
It cannot get more intuitive than this.
BforArtists specifically exists to provide better surfacing of interaction to Blender https://youtu.be/0vEtTP0C0Cs?si=comeTyStz98t9-a0
Granted Blender 2.8 onwards and even 2.x onwards are huge steps up from the 1.x days, but it still has one of the most opaque interaction models of any of the common 3D DCCs
Just having a better way to customize hotkeys would go really far, and I want to my mouse and movement controls to be roughly the same as unreal engine. Until those two things happen I use blender but I'm always angry when I do.
A few days ago, I tried to have it produce an o-ring. It didn't work out so well https://i.imgur.com/zYDJXpT.png
Having it use Python behaved much better: https://i.imgur.com/8uQSFtZ.png
I also gave it a more complicated problem, and it didn't do too bad before it forgot all about the instructions. It clearly lacks 3d understanding, but I don't know if I could do better given what it was given: https://imgur.com/a/8GCjmlo
While doing this, I got a lot of "There was an error generating a response" responses, especially when the screen was mostly full of nothing. I don't know how to avoid this, but it definitely struggled picking out the few pixels of relevant detail from a mostly-irrelevant screen.
It's currently not deterministic enough to entirely replace other kinds of Ui
Ideally, could be tested with something that's less straightforward, and requires understanding the data presented in some windows on the screen. Like "how can I fix this error"?
I assume most people in these comments don't understand 3d modeling, or they are seriously optimistic about THE IDEA of vision assistant AIs, but this demo is not exciting at all. In fact is detrimental to showcase real utility
This is just an early taste of a potentially powerful use case.
I understand the vision API doesn’t have memory, so each screenshot it takes is like an entire new context. If the script/application is able to send WHAT application it’s in, and has some RAG database in the backend to pull knowledge from, this would be incredibly useful.
Of course it’s slow now. If you’re legitimately stuck, a couple seconds for a personalized answer is a perfect trade off. It will get better.
There seems to be a trend in HN in general I’ve noticed of “it’s cool to be optimistic” and I don’t really like it. I mean sure I’m optimistic when it comes to human beings around me but when discussing technology and people who are out there to make money (or not) I don’t feel the need to be overly optimistic. Plus, competition is the very foundation of innovation and harsh honest feedback is a very important piece of the puzzle.
My little startup has system that can stream apps to an ML processing service with both a video of what's on screen and things like the context around it and what you clicked on. We run an LLM on top of this after a bunch of other processing (OCR, delta change detection, speech reco, etc.) for our knowledge capture purposes for our single app. It would be really straight forward to make such a platform available to others to build apps like what you're showing on more then just a browser, pretty much any application you run on a desktop or working across multiple of them, we haven't since before the new GPT4 with vision, most people weren't working on anything where that would help.
Anyway, I think we've solved a bunch of the heavy lifting to make this possible, feel free to email me, or anyone else reading this who like this space, if you're a company or dev that might want that layer so you can build something cool on top of it like this (diamond@augmend.com).
The app used is Blender, a 3d modeling application.
But hey your negative and simple comment could also be true who knows.
I'm guessing probably, it's just sending a screenshot of the screen right after the voice input finished (there are easy ways to recognize pauses) and sending that to the multimodal version of gpt-4, the one which is able to work with image data.
ChatGPT properly recognizes context of the task: that's the typical newly created document in the 3d software Blender. Since it starts out with a box, the user wanted to shape it into a sphere. ChatGPT provides him with a list of operations: change the selection mode to vertices, select them all and apply a bevel function, which in effect, will cause a lousy spherelike object to be created.
Two, because what you mentioned; there are a couple of other ways (like your primitive, but also NURBS lathe of a half circle, even the box with a lot of smoothing steps, subdivided icosahedron with smoothed faces)
In all seriousness though - this is absolutely amazing. Imagine it in conjunction with the Facebook/Rayban glasses with integrated cameras and headphones. Now you can walk around an event and hear "this is John Doe, he's a VP at X Corp..." or you look at a product and hear "you can get this for 30% less at this store"...
I appreciate the concerns around privacy - but tech has steadily been moving in the opposite direction - so at least we're starting to get some value from giving up so much data.
I sure hope uBlock will work on it :)
Yeah, people walking around with little cameras recording everything they see and sending it to OpenAI sounds totally awesome and not like a dystopian black mirror episode at all!
This is the imagination of my nightmares. Surveillance and consumption. Useful tech might have you look at a product and say "you don't actually need that. you can use the one you have at home that works great" or maybe list out the amount of global energy it takes every time you and the world queries the model.
That's called an ad (in particular, a coupon) and could be done already with phone cameras and the pseudo-AR that's been trendy in the past few years, but it's cheaper done by simple contextual advertising or even a coupon site.
Not that it's actually useful. That "30% less" is coming from somewhere, and believe me, it's not coming from the sellers if they can help it. Someone's getting shafted, and as the saying goes, if you can't spot the sucker in the room, you're the sucker.
I would be interested in how well AI vision could extract written tutorials out of video tutorials, which could then also be used for Q&A.
What model is it? And how is it understanding the image and the question? Can anyone shed some light on this technical process?
Apart from just plain under-training, there's that whole 'grokking' phenomenon that's been observed in smaller models, where it looks like they're not really improving for a long time, but _eventually_ suddenly undergo a massive improvement. I don't know if anyone has been willing to set enough cash on fire to see if grokking can happen in a large LLM and/or how long it would need to be trained for.
There's also still a lot of juice left in improving the quality of datasets and in training on larger data sets. OpenAI, having built GPT-4, probably want to work on some applications of it given that they have enough of a lead on everyone else. I think there's definitely a few "burn piles of cash to train a better model" buttons at OpenAI right now, but they have no reason to use them when they're clearly in the lead anyway.
There's also the actual limiting factor: there are only so many A100s/H100s in the world, but the amount is growing.