BlenderGPT: Use commands in English to control Blender with OpenAI's GPT-4
github.com
github.com
Perhaps this is infeasible with the current cost per generation but a multiple choice display could be handy (perhaps split the screen into quadrants and pick the one that most fits what you prefer).
There’s two approaches I use, I’m sure there are more.
* Do multiple completions and, filter the ones that successfully run, and take the most common result.
* Do a completion, if it fails, ask the Llm to find and correct the bug.
It could be inaccurate with high precision - like a bow that consistently undershoots, and you would have a harder time trying to correct it.
than this allegedly miraclous students are wasting there talents studying in IITs!
IIT students are amazing, i've worked with them. Far more creative and competent programmers than the average Ivy League grad in my cohort.
Would citizens of the EU and US not want to go to one of the best schools regardless of locale?
Not that it’s the same, but have you looked at either of these countries’ pasts? Even relatively recently— the Marshall Islands, or still ongoing, in Mauritius/Diego Garcia.
I say this, having studied at a European university as well as visited IIT and a top US university. These are different traditions with different strength and weaknesses but my point was not to say that IIT is better, it was to put the OP comment into a context.
IITs are way more teaching focused than research focused, and are have a very light international student presence - two factors that enormously influence rankings.
So just because someone ostensibly "passed" an exam doesn't mean they actually did.
Of course, India's education system is far from the only one where cheating is rampant, which is just one reason to look on credentials and accolades everywhere with a skeptical eye.
[1] - https://www.economist.com/asia/2022/05/26/indias-exams-are-p...
Your rich dad endows a building in an Ivy and you are in - and apparently it's all good!
I just finished writing up a long post [1] that describes how this can work on local models. It's a bit tricky to do via API efficiently, but hopefully OpenAI will give us the primitives one day to appropriately steer these models to do the right things (tm).
"Monkey around with the prompt and pray" has replaced unit testing.
And that matters to who, and why, exactly?
Once enough of the economy is automated, there's hardly a reason for it to keep humans in the loop. It can just run in a circle, serving itself. Humans, meanwhile, will just have to barter for scraps amongst themselves.
After an initial injection of capital, yes they can.
In my experience the vast majority of people are excited about GPT including artists.
You don't understand the difficulties and problems that people in the profession face, which I think is also why so many developers are convinced they can replace/"disrupt" other people's jobs with software.
I'm not saying I know that's going to continue forever, but it might. If the cost to produce software goes down, the demand for software will increase. That's what has always happened, but maybe this time is different.
Everyone always thinks this time is different though, it's good to be skeptical of thoughts like that.
My take is that if you stay on the cutting edge and get good at using all kinds of tools with max productivity, you'll probably end up as the tractor driver rather than as an unemployed ox.
When that's possible, the world will be very different. Right at this moment, AI is still useless for unskilled workers trying to write software, it's just a productivity multiplier for skilled engineers.
Do you want to listen to a 10 brand new songs by 10 brand new artists using generative AI, or do you want to listen to 10 brand new Taylor Swift songs (that were created with the help of generative AI)?
While some people will be able to leverage this to good effect, I fear the established have much more to gain in this new world…
Depends on the software, and how much mediocrity the end user is willing to put up with.
A trivial prompt can spit out a web page with functioning JavaScript for a mediocre-but-playable version of Pong.
This may not be of interest to us, but our standards are not necessarily shared by normal people: in the wild, I've seen websites where the thumbnails were all loaded as full-sized images and merely displayed smaller, bottles on supermarket shelves whose labels had easily visible pixelation and JPEG artefacts.
Infamously, there's a lot of stuff done in Excel that really shouldn't be. Some genes had to be renamed because scientists kept using Excel, and Excel kept interpreting the gene's names as dates.
I get SMSes whose sender ID has obviously involved someone somewhere trying to record phone numbers as floats.
Even in places with high standards, the UI of the Calculator app on iOS still gets confused if I tap buttons too fast (before animations finish playing?).
Here’s a recent article from the seattle times: https://www.seattletimes.com/pacific-nw-magazine/as-tech-job...
For people with eternal youthful energy, good health, no family, and a single-mindedness toward work in life.
On the second point, I actually think my salary is inflated and we'd be in the dark ages if I took that for a reason to hamper technology. Not only am I not just a developer of software but also a consumer, so I benefit directly, but more importantly so does everyone else. If everyone operated on that logic I'd still pay 20 bucks for a potato and a hundred for a hammer.
Let's be real the entire point of software is to replace labor. The software industry has done it to many sectors of the economy and called it progress. Which it is. We have no right to start complaining now.
Eventually we need to stop and think about how capitalism is winner take all, or we're all in for a very bad time.
It's not a code writer though, that's not its sole trained task. Why do you think it's going to have a drastically harder time doing the fuzzier higher level work? Human preference and subjective work has a wider acceptance of solutions.
I can have it write abstracts and works of fiction and songs. It wrote a great kids song about bumlollies and their terrible flavour, explained syncitial nuclear aggregates to a lay audience as a jaunty pirate and created ember templates in our custom framework. Have you tried it with any architecture questions?
> Let's be real the entire point of software is to replace labor. The software industry has done it to many sectors of the economy and called it progress. Which it is. We have no right to start complaining now.
It's totally fine IMO to have the views that it's big and scary for me and also good for humanity.
Writing generic code that are more or less stackoverflow copy/paste ?
The interesting parts are coming up with the logic &c. not typing the code imho
As an example, I wont code at work anymore, I chatgpt all day. Just like that, overnight. And my productivity went 10x - though i was already 10x more productive than you.> What should have taken 1 week takes 1 day, sometimes less (it's new software im working on).
I would only code if I was doing great software, that is, for myself, we dont do that at work.
It's gonna build tension silently and then release: capital will be massively reallocated - that moment will be tsunami. Talk to your representative.
Huge numbers were left unemployed with industrial automation in the US and left unemployable. The technical term is structural unemployment. All it means is you cant retrain 10,000 factory workers to be front end developers, and that even If they find a job its often not as well paid.
The two greatest myths of modern capitalism are that free markets are good for everyone (they're not) and that automation doesnt lead to unemployment. Any reasonable assessment of the data will show both of these to be clearly false.
An important effect is the speed of change - it _might_ be possible to train the next generation such as those who would have been factory workers become front end developers, but it's an entirely different challenge to take actual factory workers and train them for another job. That is to say, even if in the long run automation doesn't lead to unemployment, the short term effect may be quite different.
It sounds more like you're just seriously underestimating the direction GitHub Copilot is headed.
By my estimation it seems probable that we'll end up in the near future with some Jira integration that has an "auto fix" button on tickets. Possibly PMs or managers will be empowered to replace a large chunk of work that is currently done by people like me... and these things are only the beginning!
If you were thinking of something more severe then I'm curious what you're referring to? Otherwise I'm not sure what your point is.
It’s not hard to see where this is heading. And the goal is to automate away everything so we humans can just kick our feet up.
What's the problem if the new "source code" is all prompts written in natural language and the classical code is just a build artifact?
Most clothes people wear are garbage compared to bespoke clothes. Most industrial food is garbage compared to what a chef might make. And yet.
https://www.reddit.com/r/blender/comments/121lhfq/i_lost_eve...
Edit; I see it was discussed indeed on HN yesterday. I have already seen tons more of these in my circles and not only 3d/artists.
The more I think about this trend, the more I think it might be good.
Bootcamp devs are no longer good enough for junior roles since a GPT could replace them. Digital media people who learnt via YouTube and have no real talent are no longer skilled enough. Writers who can only churn out mediocre blog spam are now jobless.
This seems like it might be a net benefit.
Before sometime hits merge, they need to understand if the code meets all requirements, and it doesn't really matter who or what wrote the code in the PR.
The demand for engineers, and therefore junior engineers, will be reduced. That is the goal.
I make 6 figures, never took any of these classes (though I did write my own OS, and compiler). Why do you want people to waste time and money? Perhaps it’s time to split out those classes into some degree that’s more relevant to that work?
Also, I don't have a CS degree and that's not what I advocate for. I advocate for developers who spend years honing their skills and learning fundamentals of the craft.
I have also written a compiler and my own language as well as an operating system, and its usefulness has come up exactly once in near 15 years in the industry, and that was a very niche topic. Explain why, exactly, you need to know compilers and operating systems to write any modern day app.
When you openly state that you dislike these types of programmers, and want to essentially purge them? Hopefully AI destroys your specific job and you can’t find work again.
The relevant comparison is with junior artists in the age of Midjourney, Stable Diffusion, ControlNet, BlenderGPT, etc. They're facing the same uncertainty as software devs, at the same time, so there isn't much insight to be gained here just yet.
Your assumption that this raises the bar for human 3D artists is correct today, but it won't be long before human 3D artists are seen as much slower and less competent than AI artists, and there will be no going back.
irrelevant.
my whole world view changed when I read https://arxiv.org/pdf/2303.12712.pdf
now pair that with a larger memory, backtracking/revising and on-the-fly weight adjustment (aka, real-time "learning") and I think it might be game over.
add goals and motivations and maybe a vision system? game over for meat bags.
These advances are not only possible, they're inevitable. It's just too tantalising to leave alone.
there is such thing as verifiable software.
Only some software is possible to verify, and there are many properties that it's impossible to verify because the software isn't the only thing that exists in the universe. No amount of mathematical proof on an ideal RAM machine will anticipate rowhammer.
And: just because something's theoretically possible, that doesn't mean an AI system would automatically pick up the ability to do it. Verifiable software in practice is still way behind what we currently know to be possible.
you can infer algorithm failure rate depending on other input factors as an input. Say you found algorithm will fails every 10e15 years of continues run, you can accept such algorithm as reliable.
You're assuming we've solved physics, and that it would be tractable to model all of that. We haven't, and it probably won't be.
That is of course, until GPT-6 surpasses them.
Why should someone hire someone experienced when they could just cruise on patchwork solutions made by inexperienced contributors using an AI model until the next gen of an AI model is released?
Can the experienced person really outpace the model development in terms of innovation? Is it worth trying to innovate in a niche?
My experience in the software industry (20 years now) showed me that the best ones were the ones who got into it out of genuine interest. They tended to write software as a hobby.
There was no shortage of CS grads who couldn't be nearly as productive.
The self-taught ones, or the ones with genuine interest who also completed a degree program were the best.
I wouldn't discriminate against "boot camp coders" or people who learn things from YouTube.
There's a lot of people who live in a different world where an expensive college/university education is not an option.
You have a beach, and you want a jetty, just grease pencil approximately where you want it and go Hey thats a jetty build it. and if its not quite right, generate me 50 different versions and ill pick the best.
If that job will even be available. The other day my sister sent me a photo of a little food delivery robot she spotted on the streets[0], and mind you, we're not living in Silicon Valley, but in Poland.
--
[0] - https://www.deliverycouple.com/ - based on the markings on that robot, its these people.
My predication: in the next 5~10 years, most artists won't be "prompt engineers". Instead they'll focus on fix small details on AI-generated art.
It's still kinda sad tho, because it's usually the most tedious and boring part of the process. Now AI is taking the fun part and leaving the unfun part to humans.
I hope that these task-specific implementations of AI can reduce the tedium in these fields, like the way PCs did. Certainly, this advancement is leading to the ability for practically anyone to program a computer, in the general sense. Things will shift, but there will be opportunities to exploit those abilities for personal gain.
I think many people hear and see these complaints and think that people are being luddites or being afraid of losing their job. People should be looking at it for what it is - someone complaining that their entire job is changing into something they no longer enjoy.
He offered the explanation that so much of their time is consumed with nit picking through purely aesthetic decisions that AI would not be capable of the artistic reasoning required to produce work that could even get to the "pass or reject" stage.
Text was paralleled first. Then sound was paralleled next. And now image is being paralleled. It will be on level with highest percentile of human ability just as text and sound were before it. Game devs and comic artists are already replacing texture and background artists with AI generated images, just as they used AI to create hundreds of thousands of lines of fluff dialogue before then.
You can see the prompt here: https://github.com/gd3kr/BlenderGPT/blob/main/__init__.py
It is really easy to build this kind of thing - I've got a very simple command line chatbot that should be very understandable and you can easily play with the prompt.
https://github.com/atomic14/command_line_chatgpt
I would also recommend that people try out the openai playgrounds. They are great for experimenting with parameters.
That said, through my testing I found it tended to trail off during longer exchanges if I only added the guidelines to the system prompt, so I ended up adding them to the first user message + adding a small reminder at the end of each further user message.
Why didn't you do the same for OctoSQL? Are you considering DuckDB for your compute engine for OctoSQL?
Thanks!
> Why didn't you do the same for OctoSQL?
I started with OctoSQL! But the SQL dialect is a bit non-standard, while DuckDB uses the popular postgres dialect. This came up esp. with more complicated queries, where GPT was generating queries that failed with OctoSQL, while it manages to do well with the DuckDB dialect (even though it sometimes needs 2-3 attempts, but those are automatic).
> Are you considering DuckDB for your compute engine for OctoSQL?
I've been exploring that. The main disadvantage is that OctoSQL has support for temporal features and live-updating queries (something like reactive materialized views, in a sense), which would not be possible to accomplish with DuckDB.
Moreover, if you're already using DuckDB for the execution engine, there's not much reason (other than the plugin system requiring you to use C++) to not use it end2end. I think in that case a more duckdb-native project would make sense to just fill in the niche of simpler plugin authoring for DuckDB. Something like Steampipe for Postgres foreign data wrappers, but for DuckDB.
For OctoSQL I'm experimenting with WASM now as a SQL compilation target, as that could be designed to support the features I've mentioned above.
See how this file quickly got hacked together all in one file. It's refreshing to see and note-worthy as it appears when new exciting world-changing tech emerges that makes programmers let-go and become hackers again.
Ultimately there's a relationship between the preciseness in which you want to control something and the underlying information, as conveyed in language, to describe such precision.
Whether you use plain english, or code - ultimately to do things of sufficient precision you will have to be equally precise in your description. I'm sure someone with more time and more knowledge on such things have already formalized this in some information theory paper, but...
The point I'm making here is that this is great because a lot of people are doing "simple" things, and now they will be able to do those things without understanding the idiosyncrasies of Blender APIs, but I'm convinced that this will ultimately turn into something equally difficult as blender APIs to do novel things. and WHEN (not if) that happens, I hope users are prepared to learn the Blender APIs, because it will be inevitable.
edit:
one other thought. I think "language models" are not the right solution ultimately. I think kind of like AI didn't boom until the proper compute was available even though theoretical models and algorithms existed, language models are the crud solution.
once we have a loseless way to simply "think" what we want, then a "large thought model" will have less trouble, as there will be less ambiguity in what you want to what is said.
right now it's thought -> language -> model.
later it will be thought -> model.
When I write a song I usually just noodle on my guitar with some pre-programmed drums in the background, I just play whatever comes to mind at the time, record it, then listen back to it, change a few things, add a few accents, decide to add another guitar line, maybe shift in fifths or sevenths to add more voices, add a few instruments that fade in and out like strings or brass, etc.
Some people might have more methodical approaches to art, and that's fine too, but in my case it's absolutely a 100% exploration effort about stuff that I don't even know I want until I see it in front of my eyes. These tools are amazing for this.
The issue is, sometimes those two things seem to be the same thing!
tl;dr: it's easier to tell chatgpt to "Rewrite this story: " and then feed back previous outputs when writing a story than it is to get to an acceptable output from massively detailed prompts or long chains of iteration; this trait has far-reaching consequence rather than just writing fiction.
I do understand , however, that 'long-term memory' is a very active point of discussion and development.
You can't teach it new information, and often that is required to solve a problem.
This isn't a given. Plenty of times this isn't true. You cannot convince GPT to answer every problem all the time.
For example, try and teach it a grammar. No matter how many times you try and work with it, you won't be able to.
You can teach almost anyone a grammar if they are inclined to try to fuss through it. Not GPT.
And yes I have used GPT-4 a lot, please don't assume I haven't.
And that’s fair, you’re exciting the existing network, not training/changing its values fundamentally. But you can do a lot with that excitation, because it’s already got a lot to work with. And I don’t think this is very far off from how people work with new ideas in the immediate term - when they first hear about them, they think of them in terms of things they already understand.
It has been my experience that it’s generally capable of getting closer to the target, but maybe I just haven’t tried to push it past its capabilities.
One interesting thing would be to describe a scene and get the rough print. But do it in sections, such that you can select and begin to refine sections and elements within the scene to whittle down to the preciseness that pleases you for each element...
What would be really interesting will be how geometry nodes can be managed using gpt.
Once you have the source code though, you can use a variety of tools to manipulate it and save the result. Using chatbots to make modifications under supervision is fine. You discard bad modifications and save the good ones.
This is using natural language for one-offs and source code for reproducible results. It's looking like they will go well together.
There are other reasons not to do it, like it being an external API that charges money. And even if you got something local and deterministic, the generated code would be less easily tweaked by editing the original prompt than by asking the LLM to make the change you want, or by editing the code directly.
After all, using the Blender GUI, you can do a lot using only a 2D mouse coordinate and two boutons. So 2D mouse coordinates and text could be better.
A nice evolution would be an AI model that can understand natural language instructions, while taking into account where your mouse pointer, how the model is zoomed and oriented, and that has geometric insight of the 3D scene built so far.
with a search plugin you can have it find the api docs and have it output interesting parameters and how to use them with examples.
with a python REPL plugin you can have it generate 10 variations and run the code for each.
with GPT4 and plugins you could describe the output you want to midjourney or something and give the prompt to it to generate it in blender(feed the outputs to some image similarity vector to compare) and have it search through parameter space(or vector space of the prompt) until it finds something pretty close to what you want.
given your budget of course.
Imagine Jupyter notebooks with this capability. Or Photoshop. Or Davinci Resolve. We live in amazing times.
[1]: https://github.com/gd3kr/BlenderGPT/blob/main/__init__.py
system_prompt = """You are an assistant made for the purposes of helping the user with Blender, the 3D software.
- Respond with your answers in markdown (```).
- Preferably import entire modules instead of bits.
- Do not perform destructive operations on the meshes.
- Do not use cap_ends. Do not do more than what is asked (setting up render settings, adding cameras, etc)
- Do not respond with anything that is not Python code.https://www.cs.utexas.edu/~EWD/transcriptions/EWD09xx/EWD952...
If this weren't the case then it wouldn't be possible for (e.g.) the software industry to exist as it does: non-technical folks using natural language are able to converse with engineers who take informal descriptions and turn them into code, often leaning heavily on iteration the bring code and spec into conformance.
There have been many, many cases where I was not able to get GPT-4 to "understand" my problem. No matter how much I tried (until I hit the rate limit for those hours, anyway).
People are throwing these absolutes around, and it's just not totally true.
Much of an engineer's job is to try and implement the correct solution for imperfect requirements, then to go back and quickly fix things to match the real requirements.
Or try formulating a math proof with natural language.
Edit: besides, if it could work you lose the competitive edge. I could describe a much faster more cost effective system which the machine can implement. And we are off to the races again..
It's being rumored that OpenAI is currently training GPT-5 which will be ready in December, and that many people in the company think it will be a human or better level AGI. Even if it isn't the consistently supersonic jumps they are making every generation suggests we don't have long until human brains are outmoded legacy hardware.
>We’d have other problems than scaling crud apps, I think.
Ever since I first interacted with the original GPT-3 in 2020 I've had the realization that our future was going to be curtailed and distorted into an inconceivable Escher piece. It seems that future is nearly upon us.
I'm all over the place with this. Some days I think it's no big deal, but sometimes I'll get angsty about it.
Today, for example, I'm using it to generate some animations and stuff I generally don't like dealing with (math problems). I remember spending hours on this and not getting anywhere. This thing makes all that effort seem like handcoding websites in the era of templates.
Whenever it hits my direct line of work I'm like "no way that thing works, see, it did this small thing wrong and it proves it is fundamentally incapable of anything". When I use it for domains outside of my expertise I switch to "yeah, sure, but this was either already exceedingly obvious and/or nonsense busy-work to begin with".
News at 11: developer is arrogant.
Besides maybe.. "make all these entrepeneurial types obsolete." Poof!
We as humans are not just in the business of solving general problems. We are competing with each other. We need to be faster than the slower ones to survive. (I like to change that but that is not a technical issue.)
One of the ways to compete is to “talk faster” with it. Iterate quicker than the competition. How? I daresay we might get there faster by talking in some sort of modified language.. a code of sorts..
Another way to compete is to become a deep domain expert. Expert of what exactly, if AI is doing it all? Human psychology?
I guess I am just interested in the competitive aspect of it. I have no idea what will happen, but definitely curious what will be possible.
The interface is not the ChatGPT text box; it's SQL. The ChatGPT text box is just an assistant to help you do the correct thing (or at least, that's the way it should be used).
"Make a juvenile elephant."
"Make his ears comically large."
"Bigger."
"Give him a little floppy hat."
I don't think it's out of the question for these kinds of commands to result in the correct outcomes. Now, maybe I can adjust the ear size more precisely with my mouse, but it probably saved me a bunch of work.
When a client wants a button on a webpage, they don't send the web designer a legaleze document describing the dimensions of the button. They usually don't even tell the designer what font to use.
The web designer pattern matches the client's english request to the dozens of websites they've built, similar buttons they've seen and used, and then either asks for clarification, or specifies it clearly to the machine in a more specific way.
Is that different from the chatGPT flow?
Honestly, we also already mostly use english for programming too, not just design. Most of programming now is glueing together libraries, and libraries don't provide a formal logical specification of how each function works. No, they provide english documentation saying something like "http.get(url) returns an httpresponse object or an error". That's far from an actual mathematical specification of how it works, but the plain english definition is enough that most programmers won't ever look at the implementation, the actually correct specification, because the english docs are fine.
The designer knows the context of the question, the website, the previous meetings about the design styles, possibly information about the visitor demographics and goals, knows the implicit rules about approvals and company hierarchy, knows the toolset used, the project conventions, the previous issues, the test procedures, etc.
The equivalent of telling a designer where you want a new button would be equivalent to feeding a small book of the implicit context into ChatGPT and without access to visual feedback you could still end up with an off-screen button that passes all the tests and doesn't do anything. The "fun" part is that for simple tasks 90% of the time it will work every time - then it will do something completely stupid.
> they don't send the web designer a legaleze
That's the implicit context. (And yeah, bad assumptions about what both sides agree on causes problems for people too)
Also, chatgpt will try to make you happy. You want a green button here? You'll get a green button here. A designer instead will tell you it's a terrible idea and breaks accessibility.
Now, that safely describes a modern, optimizing C compiler.....
On the other side of the coin, there’s C++, which is usually doing the heavy lifting underneath the underspecified-but-sufficient Python code.
My guess is that as LLMs evolve, they will more naturally fill this niche and you will have high-level, underspecified “code” (prompts) that then glued together more formal libraries (like OpenAI’s plugins).
When a char AI misunderstood you, it's often quite easy to explain where the misunderstanding happened and the AI will correct itself.
The most familiar formal language grammars to most people here are programming languages. The difference between them and natural language has been categorized as the difference between "context-free grammar" and "context-dependent grammar".
The most popular context-free language is mathematics. The language of math provides an excellent grammar for expressing logical relationships. You can take an equation, write it in math, and transform it into a different equivalent representation. But why? Arithmetic. The Pythagorean Theorem would be wholly inconsequential if we didn't have an interest in calculating triangles. The application of math exists outside the grammar itself. This is why you, and everyone else here, grew up with story problems in math class.
Similarly, programming languages provide excellent utility for describing explicit computational behavior. What they are missing is the reason why that behavior should exist at all. Programs are surrounded by moats of incompatible context: it takes explicit design to coordinate them together.
If we can be explicit about the context in which a formalism exists, we could eliminate the need for ambiguity. With that work done, the incompatibility between software could be factored out. We could be precise about what we mean, and clear about what we infer. We could factor out all semantic arguments, and all logically fallacious positions. We could make empathy itself into software. That is the dream of Natural Language Processing.
I think that dream is achievable, but certainly not through implicit text models (LLMs). We need an explicit symbolic approach like parsing. If you're interested, I have been chewing on an idea that might work.
Such a well written reply! This puts into words a lot of my thoughts around programming today and how NLP can help.
Sometimes smart people say dumb things and it takes a while to figure out they are wrong.
Considering they also predicted iPads, we might want to take it at face value.
Also, the wax tablets of antiquity also strongly resemble iPads. It's a fairly old invention. Might as well argue the ancient Greeks invented iPads.
An is-versus-ought style mistake.
"The Sketchpad system makes it possible for a man and a computer to converse rapidly through the medium of line drawings. Heretofore, most interaction between men and computers has been slowed down by the need to reduce all communication to written statements that can be typed" - Sutherland
It's not perfect, but LLMs appear capable of making the same kind of inference, sufficiently that it'll inevitably be possible to program in natural language sooner or later with minimal or no manual double-checking
We're nearing the precipice of more natural human-computer interaction that will need to rethink the interfaces and conventions.
Alexa and Siri seem like Model T Fords when there's a jet aircraft flying overhead. I'm thinking these agents need to be replaced by more natural agents who can co-create a language with their human counterparts rather than relying on fixed, awkward, and sometimes unhelpful commands. It would behoove us to expose APIs and permissions delegation in a more consistent and self-describing (OpenAPI + OAuth / SAML possibly) manner for all possible services one would wish to grant to an agent. If a natural language agent is uncertain, it should ask for clarification. And on results, it is necessary to capture ever-more-precise feedback from users because positive and negative prompts aren't good enough.
I think Douglas Adams was closer to the truth on the subject of AIs and tea. I don't want to be overly cynical but I suspect we'll just get used to saying "OK, close enough" when dealing with LLMs.
Is that somehow baked into the algorithms?
Are positive words of encouragement interpreted as "positive signals" by the inference pipeline? Or do they somehow influence the attention mechanism?
Because otherwise, you're just rationalizing completely random and unpredicted behavior.
When super intelligent AI gains power, I want it to know I've been a good boy.
It's a system where you can talk to it and make a photo realistic movie. The example I always use is, you're sitting at the computer, looking at a blank screen and you say something like:
"Ok opening scene. Dockside, London, early 19th century. Early evening. There are several ships docked, one being offloaded. Stevedores are working, some disreputable louts hanging around."
The screen is updating as I'm talking.
"OK make it grittier, more dirt and grime, let's have a fight break out in mid distance left of the screen. Now pan slowly right to reveal a bar called the Skull and Crown. Make the sign dirtier but let the last light of sunset glint off of the skull."
Screen updates. We are looking at what appears to be a Hollywood level period set full of extras who look the way they should, based on historical data that the model has.
"As we pan over towards the door Micky gets tossed out by the big burly barman. Make him younger, skinnier, he's about 17 years old."
The point of all this is, no, you don't need exact language to specify what you want. In the world of filmmaking you never do that. The screenwriter describes things in some detail, but always leaves a lot up to the interpretation of the director, the set designer, the costumer, the makeup artist, the casting director, etc.
The AI can take on any or all of those roles for us.
What I want is something I can control to make the movie I want to make. Then I want to be able to iterate on it: Let's make the main character a woman. Now everything gets changed to fit that. etc.
Of course the AI can replace the role of the writer too, and the director, and the producer leaving me with nothing to do. But the fact that I can bring my vision to the screen still makes it a great tool.
Seeing something like this makes me think that the arbitrary holodeck commands "Paris, 1950's, rainy afternoon" is suddenly not a challenging part of the equation. It's really exciting.
https://i.imgur.com/IYuh29H.png
Not perfect but man, we're getting pretty close.
i now ran into your comment (with a purple link) and did some reflection. upon reexamination, its clear that the picture is fake (because im looking for it) but when i wasn't looking for it, its interesting how all the "hot spots" or interesting pieces of the picture are pretty good and the (imo) lackluster parts are the "less interesting" pieces like the end of the roads where it it blurs out. i wonder if that bias is inherently ingrained in the system.
https://i.imgur.com/pPU7K0c.png
Things still get a little weird in the distance (particularly in photo 3), but I think overall it's a bit better. People who are really good at writing prompts could probably do even better, although one of the strengths of MidJourney V4 and V5 is that it can give good results without the traditional paragraph of "incredible, award winning, photo of the year" etc.
Subsequently, this applies to posters, letters, newspapers, and other types of text-heavy images, ultimately reducing the language modeling problem to an image generation problem.
Now it's clear that the shift is coming and it will revolutionize the way we interface with machines.
Studios like Wetta Digital / DoubleNegative etc.. are gonna pounce on this
I would love to be able to have GPT sketch math figures, which I then modify/perfect.
Note: this comment is partially inspired by the workflow of Gilles Castel — I’d love to be able to use GPT in the loop of note taking, similar to the system that Gilles setup to improve sketching speed.
I'd love to add this capability to our SaaS product, but I've waited for OpenAI to make GPT-3.5 or GPT-4 available for fine-tuning. (Cramming an entire API into the prompt does not seem feasible, not even with support for 32K tokens.)
Something like `Some text blabla.<span style="display: none;">Hidden text</span>` And when asked for something specific, GPT would output the hidden text.
So you could push code onto github with an exploit along to common usecases.
EDIT: found it: https://news.ycombinator.com/item?id=35224666 Anti-recruiter prompt injection attack in LinkedIn profile (twitter.com/brdskggs)
LOL
The trend with generative design automatically pushes the user towards high fidelity thinking. I am afraid that linguistic interfaces are in conflict with a natural human creativity. They are useful as a tool for specific use cases, but the idea that they will replace or augment the design process is ludicrous.
Another problem for me is the post effect of prompting. Millions of people will have little to no incentive to learn. This is gamification of the design process, with unknown social and economical effect. People have a tendency to search for the easy question and answer. This is not progress at all.
The push from A.I. marketing is immense and people are freaking out. This is the first tech product which is having a negative impact before even reaching broad adoption.
Suddenly A.I. ethics teams are fired and nobody has any issue with alignment and black box? Ok, computer.
For me, the responsible thing is governments to regulate the implementation process with frameworks which are not so hard to build on ethical basis. The Roman law will always give a loophole for exploitation by the big corporations.
Don't get me wrong. ChatGPT is a very powerful tool for summarization, sentiment analysis, text classification, codebase documentation etc. But the design industry implementation in my view is not well thought. I would like to have assistant, not generator. As a designer, there are a ton of use cases for automation, interactive help, etc. Sadly, we are going in a direction which will produce polished mediocrity on a grand scale. Soon we will need fact checking A.I. and A.I. content blocking everywhere.
The other day, I shot some footage with my Blackmagic camera and proceed to do some editing and color correction. Virtually nowhere, I had the need for linguistic interface. We have powerful tools in our disposal as it is. The content is the problem. Dopamine driven short forms are changing the way people interact with the world. The average attention span in 2000 was 12 seconds, in 2015 – 8.25 seconds, today is less. So the tech industry tries hard to convince all of us that the progress is in merging with the machines and living 24/7 in A.I. induced coma? No, thanks. Keep your SOMA for yourselves. We like it natural here:)
I started with a scene that had a camera and a cube. I told it to make the cube red. Result: "Error executing generated code.."
I told it to delete the cube. It failed again.
I deleted the cube and told it to make a cube at origin. It made a cube.
I told it to make the cube bigger. It made the cube bigger.
I told it to make the cube red. Fail.
I told it to make an animation of a spinning cube. Fail.
I told it to make a car. It made a rectangular cube with 4 cylinders, somewhat resembling a toy car made of wood blocks, laying on its side.
I told it to turn the car upright and make it more aerodynamic. It failed and I returned here to HN to see if anyone else had some advice.
Also FreeCAD/KiCad, I guess, if the resources are there, but Dassault and similar have the ability to bring to bear a lot of FTEs very quickly if they wish.
Never thought I'd see the day I'd pay OpenAI a cent given that I don't really agree with their level of "open"ness.
(Practically, you aren’t going to get quite the ideal even on the 8k context option because you are going to use at least some response tokens on each request, though I would imagine text -> software controls can be optimized so that the response is fairly token-efficient in most cases.)
Also, GPT-4 API access is in limited-access with a waitlist that is advertised as prioritized based on submitting AI evaluation cases to OpenAI’s repository.
I see this complaint a lot, but after watching a lot of interviews with the founders, I think they have by and large taken the right approach. They're trying to drip feed things at a rate that allows it to be digested enough by all parties before releasing more. They actually appear to be quite principled thinkers on this topic.
This seems like a rather shallow interpretation. What's the actual claim here? That Sam Altman, Greg Brockman, and Ilya Sutskever all have narcissistic personality disorder?
If you were in there shoes, what would you do and what would the rationale be? What might some of the consequences be?
From listening to interviews with them, they seem to be cognizant of the fact that all of this tech will inevitably get into the hands of the public one way or another. They seem mostly to be trying to give people as much time to process the shift that's happening, think through potential implications, and try to adjust as necessary.
In the long run I think that continuous training will be obvious with new models rolled out on a very high frequency if not continuously.
Any time we can get a program to do the repetitive work that is time consuming but not interesting/creative that is a huge win.
I knew that Blender has the possibility to script and access all the UI elements with python. So it was only a matter of time :)
The first is to just use the fact the GPT-4 will have seen a lot of blender code so just knows how to do it.
The second way is to tell GPT-4 in the prompt what the API surface looks like and have it script against that.