Terence Tao on GPT-4
mathstodon.xyz
mathstodon.xyz
So yeah, “only”.
University could have paid for Terence Tao’s assistant, but then they could have paid him more too.
If you expand your point of view to the interests of all mathematicians, or society at large, the waste here is even more extravagant.
To the GP's point, a friend of mine makes 80k but is still provided an assistant. That's in the entertainment industry, where PA's are dirt cheap and union laws prohibit multitasking.
Either salaries in the US are extremely higher than what I thought, you have a very idiosyncratic definition of "good" or you are significantly overestimating how much people make.
The point is that amongst people that I know who are making 500k, the majority of them aren’t having assistants. And I do think it extrapolate to the population of people with 500k income too (no proof here, just a hunch).
In the case of this post, one could argue that collating statistics should be done by ICM staff rather than the committee chair.
For example, this article [1] says that to be in the top 1% of earners in the US a family needs to make $600k or more. There are 330 million people in the US, meaning that 3.3 million people are in a household earning $600k or more. I think that's a lot!
[1] https://smartasset.com/data-studies/what-it-takes-to-be-in-t...
That is a bit more than 1/9th of what the head football coach at UCLA makes, who has two assistants (offensive and defensive coordinators) who each earned a bit more than Tao. Additionally, UCLA still paid more than $3M in 2021 to another head football coach whom they had fired in 2017. The university seems to have fairly clearly defined priorities.
https://www.collegefactual.com/colleges/university-of-califo....
But even IF this were a profitable activity, should a university really be in the business of running an enterprise that literally causes irreversible brain damage (Story about USC, but undoubtedly applies to all football programs: https://www.si.com/college/2020/10/07/usc-and-its-dying-line...)?
With respect to the brain damage, he spent much longer in the professional setting than college, it’s more likely the damage came from there. I don’t think colleges should be in the business of restricting what students should be allowed to explore based on the possibility they might choose a career where brain damage may occur. And if they did, Engineering should be the first to go. Just think of all the brain damage resulting from professionals engineers of the military industrial complex!
Every numbskull to every professor now has the ability to generate information on a scholarly level and solve problems that could take a human years of training and experience to produce.
I'm really trying hard not to get all doom and gloom about this stuff, but I'm honestly worried. It keeps me up at night.
Mass unemployment is a likely outcome, but modern states already have amazing safety nets in place and they're not far away from a universal basic income.
As technology has improved, so has life on almost all relevant metrics (education, longevity, poverty, crime etc.) and even though technology amplifies greed, criminals do not prevail because they are heavily selected against seemingly regardless of the level of technological sophistication.
As for paperclippers, there is actually no indication that larger neural nets misinterpret human intention. Rather, larger NNs mostly become better at understanding. Current AIs when prompted to make paperclips do not interpret this as us wanting to turn the entire world into paper clips, but in terms of (the more sensible interpretation) of meeting the economic demand for paperclips, hence giving advice on how to optimize the machinery for producing paperclips etc. Small models do tend to find weird, unintended shotcuts, but my hunch is we are seeing this less with very large NNs.
Lastly, we can always employ AIs to control other AIs. AI scammers will be counteracted by scam detection AIs. Same for fake news, mind control, drone warfare, propaganda, etc. AI will allow us to make better political arguments against mass surveillance and for freedom, and it will make journalism more efficient at exposing crime and injustice. Demonopolizing AI is hence extremely important, so that every misuse of AI or misaligned AI can be counteracted by other AIs.
very few western and east asian nation have that. 80% of population dont have safety net
I'm hopeful about the future use of AI, but this strikes me as potentially overly optimistic. Or possible misinterprets what academics do.
ChatGPT is very good at collating existing information. I don't know that it is anywhere near the level of creativity necessary to generate novel ideas (especially those that have to be rooted in reality to solve engineering problems).
For example, if you ask "How can we redesign a Wankel engine to achieve 2% better efficiency?" it does give a good summary of the main mechanisms that impact efficiency (thermal losses, friction, combustion efficiency) but it doesn't give any actionable, novel ways of implementing them. It's basically a 101 course summary of combustion engines. "Generate better materials" is not an actionable solution. So unless you think a freshman/sophomore understanding is all that's needed to solve our big problems, we've still got a ways to go before we can turn the reins over to AI.
I think a lot of people misunderstand what a scholarly level is. ChatGPT is not even close.
Unfortunately, professors currently have to spend a lot of time doing tedious administrative/organizational work unrelated to research/teaching.
That led me to a realization, this could mean that people will likely discard less known libraries or technologies more easily than before. Popular technologies will become even more popular since we’ll be 10x more productive with them compared to less known ones since we won’t have LLMs super powers for those.
When I ask it about Objective-C it bullshits and fabricates APIs.
When I ask it about python or bash scripts it is astonishingly good
Then I asked chatgpt to port it piece by piece to C.
Took a while to get a hang of gpt4's token length limitations, as well as fixes here and there... 1.5 hours.
But damn... Most programmers are just not capable of doing this at all.
What AI seems to be replacing is mostly bullshit jobs in this case, that should have been handled by someone else.
Our species has proven itself to be exceptionally bad at finding the middle ground between giving individuals the freedom to self-direct and allocating resources efficiently at scale.
These resource allocation problems exist at the national level, but also within even some of the smallest organizations and certainly everywhere between.
Outside of FAANG, how many development teams really have a dedicated DevOps teammate? A DBA? A tester? A sysadmin? Hell, a sysop?!
SWEs outside of FAANG are so accustomed to wearing dozens of hats that many would still consider the wisdom in "The Mythical Man-Month" a pipedream.
terry mentions running some checksums after processing on the linked page. I'd imagine most people working on serious tasks do something similar.
My current solution is pdfplumber → GPT-3 API. I played around with a few different options, and this is personally what’s worked best for my use cases.
[1] Well . . kinda. The model has some weird ideas about what an edge table is.
> with the only tedious aspect being the cut-and-paste between the raw data, GPT4, and the spreadsheet
I had a table in a PDF of registers and associated information, and I copied and pasted the text directly into ChatGPT as a big block, and asked it to structure the data as a table again based on its best understanding of the data, knowing it came from a table. To my surprise, it did a really good job. A couple small edits here and there were needed to change some formatting and it missed a couple values, but overall it took me a couple minutes to edit and I was on my way.
People keep saying it saves them time and look forward to the future.
For the life of me I can't figure out how to use it.
I recently used GPT-4 for matplotlib. I wrote a imple PDE solver and wanted it to create a function that saves a simple 3d array as an animatied 2d plot. It did it right away. I could ask it for improvements and it did it too.
Both of these tasks are easy, and if you are working every day in web development or with matplotlib I am sure you can do them in 5 minutes. But in my case, each of them might have taken half a day. And even if I could do it in 1 hour, that would be 1 hour of furrowed-brows staring at stackoverflow. Using GPT-4 is just extremely easy.
From my experience I claim that GPT-4 can also solve more complex problems. I think if I iteratively ask for features on top of what it has given me, I can get up to 3-4 times the amount of features before the whole code becomes too complex to handle. This is just a guess.
My knowledge of programming doesn't extend far beyond "the basic theory of programming" and it took about 3 days total. Without GPT4 it would probably have taken me 3 weeks. Nor would it have even of been attempted because the old tester "worked" with lots of manual intervention and frequent data losses (it ran on a winXP laptop from 2004, and relied on analog syncing signals between devices)
Since baseball is back and most of us are fans, we decided to write a baseball simulator. We each had a Friday afternoon to write one up. Half of us got to use the free GPT3, and half had just regular googling. After the jam, we'd compare notes at the bar and see what the difference, if any, was.
Holy cow, was there ever a difference.
Those without GPT3 got pretty far. Got the balls and strikes and bases and 9 innings. Most got extra innings down. One even tried the integration with ERA and batting stats in the probabilities of an event occurring but was unable to get it done.
The GPT3 group was estimated to be 2 weeks worth of work ahead of the googling group. Turns out, there is a whole python library for baseball simulations and statistics. The googling group didn't find that, but GPT3 just prompted it outright on the first query for everyone using it. This group got the basics of the game done in ~30 minutes. Managed to get integration with actual MLB statistics. Built somewhat real physics simulators of balls in play and distances, adjusted for temperature and altitude. Not all of them at once, but a lot of really great stuff.
Aside: Did you know that MLB publishes, in real time, all 6 degrees of freedom for a ball, from where it leaves a pitchers hand to where a catcher/batter interacts with it? They put out the spin rates in three axes! Wild stuff.
Our conclusions were that it's totally 'worth it' and is a ~20x multiplier in coding speed. It spits out a lot of really bad code, but it gets the skeletons out very quickly and just rockets you to the crux of the problems. For example: it gave out a lot of jibberish code with the python baseball library; like trying to pass a date into a function that only takes in names. But it gives you the correct functions. Easy enough to go and figure out the documentation on that function.
Like I said, it's a ~20x multiplier for our little experiment.
Action Items for management: Pay whatever you have to and let us use it all the time.
Also, as to what the GPT group was able to produce -- sure it was a lot of code, and apparently a quite a bucket of features -- but did it actually produce a usable simulation? Or even a coherent statement of what a "baseball simulation" should do, actually, and how its accuracy is to be measured?
I'm not casting aspersions here - I'd really like to know.
It's easy enough to try it out for yourself too! Give yourself a challenge and see where it takes you.
It's just that, if someone gave me 3 hours, and asked me to come back with constructive, actionable progress toward creating a simulator for X (where X is sufficiently rich and complex, like baseball) -- I wouldn't mess around with skeleton code at all.
Instead I'd try my best to come up with a statement of what the simulator should do, and why.
anyway, here's the text:
> Today was the first day that I could definitively say that #GPT4 has saved me a significant amount of tedious work. As part of my responsibilities as chair of the ICM Structure Committee, I needed to gather various statistics on the speakers at the previous ICM (for instance, how many speakers there were for each section, taking into account that some speakers were jointly assigned to multiple sections). The raw data (involving about 200 speakers) was not available to me in spreadsheet form, but instead in a number of tables in web pages and PDFs. In the past I would have resigned myself to the tedious taks of first manually entering the data into a spreadsheet and then looking up various spreadsheet functions to work out how to calculate exactly what I needed; but both tasks were easily accomplished in a few minutes by GPT4, and the process was even somewhat enjoyable (with the only tedious aspect being the cut-and-paste between the raw data, GPT4, and the spreadsheet).
> Am now looking forward to native integration of AI into the various software tools that I use, so that even the cut-and-paste step can be omitted. (Just being able to resolve >90% of LaTeX compilation issues automatically would be wonderful...)
Ironically, Tao's post convinces me that AI, though amazing, isn't really the solution. Better UX and data quality is. Why was the data so disjoint to begin with? Why is Latex so hard to work with?
In this case GPT-4 is used to solve a problem that shouldn't have even been one to begin with. The administrators of the ICM could've simply exported the raw data as a Google Sheet (for example) and his problem could've been trivially solved even without GPT-4.
You do know Mastodon is a federation of servers, right? I predict most people will move off this particular one to a faster one.
Anyone that got in can share what it says?
if you disagree with that then we'll have to agree to disagree.
Twitter is a far better user experience in that respect, to the point that I actively avoid clicking mastadon links now because they fail so often
incredible what mental hoops people will jump through to disqualify AI!
the reality is that a lot of things have a terrible UI with unorganized data. that's why this tool is so amazing - because it doesn't matter anymore.
how naive. how do you know it's right? ah, you have to manually do the calculation anyway to confirm, this is what Tao ended up saying in a reply asking as much.
AI is great, but it's not a silver bullet, since its correctness can never be 100% under the current LLM framework.
I don't understand how you can hold this position with AI considering it's only the beginning.
I suspect the satire whooshed over your head.
This AI future you're wanting is an Idiocracy-like world where nobody knows how anything works and everything is in decay.
Some of us are just burned out on the hype cycles and prefer not to count our chicks before they hatch.
You're making the same mistake you're accusing them of making - assuming to know the future at the beginning. You're assuming that fixing these issues will prove to be trivial or at least inevitable. Sure, recent progress has been swift, but if you recall, it was damn-near stagnant for decades. Some were even claiming we were in the middle of an "AI Winter" and could not see the spring!
Based on all currently-available evidence, the current techniques that we utilize for generative AI are unreliable, in terms of accuracy of derived facts. It will require either a different or complementary approach to iron that out, or we're going to have to start seeing some _very interesting_ emergent properties from scaling higher. This stuff could show up tomorrow, or it might never show up at all! But the _current_ LLM framework does not look like it can do what we're looking for here, not reliably, certainly not 100% reliably.
You have to do those same confirmation calculations anyway when you use a spreadsheet. In my experience the utility of something like what ChatGPT can do is still unparalleled.
Rich pro-ai argument.
The idea that everything would work great if only all of our data was structured and easily parseable everywhere just leads me to ask "Do you not interact with humans on a regular basis?"
I needed to gather various statistics on the speakers at the previous ICM
Why did this work need to be done? Are even the people who say they want this data actually going to use it for something productive? Is there something revelational in this data?
If you are producing data then exposing it in a nice programmable format is an extra cost and generally provides you no benefit. It usually hurts you, if people stop visiting your site and see fewer of your ads!
This is "really" a problem of incentives. It is usually not possible to capture any of the positive externalities of exposing your data. So maybe we could convince everybody in the world to switch to using different browsers with a native micropayment system; that might incentivize everybody to release all data as clean machine-readable tuples.
What I'm saying is, the phrase "Better UX and data quality" ignores just how hard that solution really is. It turns out training an LLM over most of the internet is _easier_ than global coordination.
I have asked a langchain bot about wikidata ids for specific places, links to the page, to read it and then to answer facts about places and got very good results instead of made up numbers.
Wikidata links to FIPS codes, OSM ids, GeoNames and that gives us an opening to link against the cool datasets from Flickr, Foursquare and others who have created gazetteers.
To me, Semantic Web was dead on arrival because of its UX, but now a semi-smart agent can help us get past the UX problems and jump from plain text to json output.
If somehow magically from the beginning of computing we had a natural language interface to a computer's operations, we would still have arrived at particular standards/specs for data formats. There would still be something like xml/json/csv. Indeed, I'd wager there would still at a certain point be some kind of high level programming (or otherwise formal) language adopted, to answer the particular clumsiness of natural language itself [1].
Putting aside any issues of reliability, its simply not sustainable (economically, environmentally) to put all our work into this stuff. Even if it does shine with one off stuff like this.
1. https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...
A woodworking example is that a planer is great tool that helps you make nice flat surfaces. But, to a certain extent, it's a downstream fix that wouldn't be necessary if a carpenter was using a better overall process. I.e., if their upstream process for cutting/ripping wood made nice flat surfaces to begin with, the awesomeness of the planer becomes moot. (Apologies to the legitimate woodworkers if this analogy is off).
Where tools like GPT becomes invaluable is when you have no control over those upstream processes but still need to get the job done. But leveraging a tool for a downstream fix when upstream fixes are possible is usually a less-good approach to creating good systems.
To torture the woodworking analogy, your assumption is that the carpenter has no control over ripping the boards. In some instances that may very well be the case, but there will also the instances where the carpenter does have influence over creating the boards, or even wholesale control over ripping them. In those cases, using a planer to fix poorly ripped boards may not be the best approach.
Even in your examples, yes, you have to work with other teams. "Control" doesn't mean you have dictatorial control over those teams. But it does sometimes mean you have build relationships, leverage what you can, and explain the value to those that do have some modicum of control. The idea that we just throw our hands up and jump to workarounds is often an excuse for taking the short-term easy at the expense of a better long-term solution.
Cost of implementing better process for all carpenters is significantly higher, than all carpenters still using bad process + _one_ AI being able to clean it for carpenter, plumber, translator, developer (you name it, you got it).
Not even entering laziness/corposlowness gardens etc.
"Corposlowness" is just another name for "bad processes". It supports the claim rather than negates it. Using AI to overcome bad bureaucracy makes it a workaround, not an idealized process. What often happens when implementing workarounds rather than good processes is that the workaround can create bloat and waste of its own and overtime, not really fix the problem. Like hiring more administrators for a large organization, they can take on a life of their own, eventually becoming divorced from the problem they were intended to solve.
Again, I'm not saying that AI is misapplied in Tao's case. I'm just cautioning that it's not a panacea for bad processes. In many ways, it can be misused as a band-aid for bad processes, just like creating excess inventory is a band-aid for bad quality control.
There are a couple of reasons why it would be hard to change the upstream process to not necessitate planers. The main one is that logs are typically ripped into boards when the wood is still green, and in the process of drying, boards change shape and dimensions: they bow, cup, warp, and shrink, and you might still need a planer to bring them back to flatness and to desired final thickness.
Do you imagine a future where machines output data that can be barely read by humans, but can only managed through the help of AI? Honest question.
There are high peaks and troughs in AI buzz right now.
Yes, on the one hand you've the but-can-it-dance crowd.
On the other hand, Terence Tao on GPT-4. I mean, I'm not weird for really expecting the story here either be about GPT4 helping Terence Tao on some difficult newfangled proof, or Tao talking about the math behind large language models. Instead this boils down to
GPT4 even does the work of some of the smartest mathematicians in the world[1]
1: by parsing some web pages and PDFs for their meetings
> 1: by parsing some web pages and PDFs for their meetings
Like that old joke about the guy who impressed people by claiming he had helped a brilliant mathematician solve a problem that had stumped him. And the punchline is something like "yeah, and it only took me a few minutes, all I had to do was replace his timing belt."
I think a lot of people miss that due to being shown in search instances and companion apps.
The hundredth time you do it you're going to be like "why is this so f'in annoying still."
Today's interface to language models is subpar for a lot of applications. Lots of room to improve that. A tool can be both amazing and still be just another step on the road to something truly seamless - just like how now it's "tedious" to use the computer for it instead of mailing/faxing paper forms around and filling out tables by pen and pencil.
If you don't like this tool, don't use it! If you like, it, use it!
Umm, and I want a pony? The world doesn't come in perfectly structured data formats. I was pretty amazed when I pasted in some of my medical lab test results and ChatGPT was accurately able to parse everything out and summarize my data. It worked extremely well.
In Tao's case, the data was organized already, simply not given to him.
In any case, Tao, per the post, had to manually calculate to confirm that GPT-4 was correct in any case (he implies as much in a comment asking if he checked for correctness).
Consider the time it would take to do manual entry for each of those examples compared with the time it takes to verify that the generated content is correct.
"Are these the correct state and zip codes?" is much faster than typing it by hand. You just ask yourself "is MA the correct state code for Massachusetts? Yep, next" rather than "Massachusetts that's... (lookup) MA, type MA; next that one is MI, already a state code ..." and down the list.
I would be willing to content that GPT will do the list faster and with better accuracy than a human doing the same work (that would also need to be checked for correctness).
I'm sure the company that did my medical labs also has my data in a structured format somewhere. So what, should I call them up and demand they give me my data in a spreadsheet? I can go down that useless path, or I can paste my data on ChatGPT and get results in 15 seconds.
Yes, go ahead and send OpenAI all of your HIPAA data.
It's my data. The entire point of HIPAA is that I own the data and I can send it to whomever I want if I decide to. I get value out of it, others may not choose to do it, that's their right. But I'm pretty sure sharing my CBC results is not going to be the death of me.
I think an important question is how much faith we give in the answer (especially with medical data!). There are lots of examples of great uses but also a number of examples of hallucinations and just plain bad summaries. When the stakes are high, the conviction needs to be couched in the risk of it being wrong. We need to be cognizant of the automation-trust factor that sometimes makes us place unwarranted trust in these systems.
Oh dear. Federation at its finest /s
If the instance is too small, it will easily fall over under heavy traffic. If it is too big, it is highly centralized and even then it cannot scale at a maximum of 300,000+ users at the same time.
Eventually, they would all be re-centralizing back to Cloudflare. Once that goes down, all of the biggest Mastodon instances go down at once.
As for GPT-4 in mathematics (LLMs in particular), in goes inline on what I said before. Fundamentally, these LLMs cannot reason transparently, nor can it directly bring up its own mathematical proofs with thorough explanations that allows mathematicians to work with. Only low hanging fruit of summarization of existing text.
If Mr. Tao can see its limitations, surely it puts the hype of unexplainable magical black-box LLMs to rest as being great sophists and bullshit generators.
"... automatically solving multiple challenging problems drawn from high school olympiads."
This combination of GPT and LEAN is the first thing that has finally got me really interested in the potential of LEAN. And LEAN very directly addresses the "bullshit" issue, since it formally verifies whether or not a proof is correct (it's like a compiler like Rust, but for mathematical proofs).
But the problem DID exist, and problems like that are likely to continue to exist for the foreseeable future.
I’m fairly convinced that for many problems (including this), AI is not the best solution but will become the preferred solution. “Best” doesnt get in the way of “good enough” when convenience is at stake (and making a good UX is often very very inconvenient)
I'm in the middle of reading Steve Jobs' very own words, Make Something Wonderful book. Apparently throughout his whole life he frequently mentioned and discussed a lot about providing and creating easy access to the computer because ultimately people do not want to program but want to use computer instead. He did mention about how Morse code was not popular as much of the later telephone technology because people refused to learn and use the unintuitive Morse code even though it only takes 40 hours for average person to learn the entire Morse code. As a trained communication engineers I can very much relate to this because even though we learnt much of the underlying technology of communication that enable the Morse code or the telephone to work across the Atlantic ocean most of us can't be bothered to learn the Morse code and use it if we can help it because the telephone make it redundant and it's counter intuitive to use.
Granted, now we know that telephone or circuit switching in general is a suboptimal in providing and solving the human communication problems and needs. As of now the Internet, packet switching and the multimedia approaches are probably the best solutions but telephone has served us a potent and still one of the best solutions for communication until very recently.
Because real life is messy.
> Why is Latex so hard to work with?
Because people that use latex want to be able to brag they use latex
(just joking for the second one.... unless ?)
This doesn't make sense - the data inherently is centralized. Tao implied as much.
> In this case GPT-4 is used to solve a problem that shouldn't have even been one to begin with.
Friction. These problems shouldn’t exist, but they do anyway, and they’re everywhere. Anything human is inevitably going to be imperfect and messy to some greater or lesser degree, introducing friction into dealing with it. Especially as we produce more and more unstructured or semi-structured data, which AI is particularly good at wrangling. If AI can help us cut through that friction significantly faster and more accurately, that’s a win.
Coworkers, associates, clients, family - everyone sends us funky data. Sometimes it's a PDF of text instead of a text file, or a picture of their computer screen with a paragraph in it, or copy-pasted text with unwanted formatting, etc, etc.
And the point of this article wasn't really about the tedious job of copy/pasting, it was that once the data was ready, instead of hunkering down and having to work through a problem with an unknown number of variables, GPT-4 can drop an answer in seconds flat. Maybe a few more if you have to tweak your prompt a few times.
Part 2 here is that we have a problem to solve, and that's how data is sent and received, and that it all needs to end up in a structural format that's usable by our AI tools.
I am thinking that writing a script that can dump textual data from literally anything I'm looking at so it's AI-ready makes a lot of sense, then we have step 1 taken care of as well for my own personal GPT-solvable situations.
Because the world is not a database. His source were formats meant for Human consumption, not machines. This will never change, people are lazy, greedy, or fear "data leaks", so we will never have a machine-first-format that everyone will use.
I mean, that guy is using latex, others have markdown, or org-mode, or Microsoft Word, those all are not meant for easy scripting. This is why AI will be such a dealbreaker, because it will close this gap, make human-formats to machine-formats when neccessary.
> Out of curiosity, are you sure that GPT did it correctly? If yes, is it because you were able to "spot check" it in a few places? Or have you used GPT enough that you trust it with a task like this? Or is this for some kind of internal use where a few small errors is unimportant, and you only need the broad strokes to be correct?
> Yes, yes, and yes. (There were some obvious checksums, for instance by computing the total number of speakers several different ways.)
Still cool, but the circumstances above aren’t always true. In fact, in a lot of meaningful work, none of them probably are.
If GPT is like an assistant whose work you always have to double check in order to be sure of the validity of the results, that seems to be a pretty big caveat to me. In that case it’s probably better to just write a script whose output you know to be deterministic, even if it takes slightly longer (and the bigger your dataset to verify, the more likely it is that it’ll take less time to write a script than to validate results one by one).
The scary part is if/when companies/bureaucracies/governments/etc start using GPT for all sorts of Important Tasks and skip the validation part because they assume the machine will always get it right.
Sounds like every junior developer fresh out of college I've had to work with. Interns are even worse. Even senior devs mess up from time to time. That's why we do code review and other QA. Trust no one.
Another key difference is that a deterministic script that has determined to be valid will be valid for all subsequent runs, regardless of input size. That is not the case for ChatGPT, which might be correct on some runs and not others, or do well on small input sizes but start messing up when thousands (or more) of data items are involved.
Again, I still think ChatGPT is really cool, but thinking about all these nuances about the nature of the work that is actually being done when you use ChatGPT vs writing a script vs handing it off to an intern seems crucial to the debate.
In this case, AI was the tool used to have a better UX. People build something with use-case X in mind, and use the best tool for that job. Piping that into use-case Y requires some duct tape and plumbing. It turns out that AI is great at that sort of repurposing.
It's great if someone using the best tool for their job is nearly the best tool for the flexible infinity of other use cases, which is what flexible enough duct-tape gets us.
Do you mean build better, specific, UX and normalize data for only this task? Why spend that time, even as a developer of whatever improved app/system you're thinking of, when he can just turn to AI?
I'd rather lean on GPT4's ability to generalize highly specific technical work rather than ruminate on each individual app's lackings and how it could be better than using AI if we just... put more effort into it?
I think the reality is, there are cases where it matters, and cases where it doesn't matter (at least, not very much).
I work on aerospace software systems. I suspect that most software developers would be surprised at the process and rigor involved. That does not mean that humans are flawless; of course not. But there are going to be a lot more hurdles to replacing all of the humans with non-deterministic machines.
There are numerous software components on a modern airplane. How good is good enough? If the overall system works correctly 99% of the time, is that good enough? There are about 16 million flights annually in the United States alone. If 160,000 of them crash, is that good enough?
I mean, like, if Microsoft Word failed to save your document 99% of the time, you might be irked. If Professor Tao's data was only 99% accurate, he probably still has a reasonably good picture of the information he needs.
Regardless of if a human or a machine is writing the code, there are different levels of software criticality. As humans, we don't apply avionics software development process to iOS games because it's not worth the extra time and expense. Likewise, there will be a spectrum of where it makes good enough sense to automate tasks with AI and where it doesn't.
Which is not to dismiss AI! It's amazingly useful technology. But there really are places where above average accuracy is needed, and AI that works great for handling some tasks might not suffice for other tasks. If you truly need all of the results correct, then a nondeterministic neural network is probably not the best path to be on. There are other methods; even other automated methods.
I did this once, a fuel estimation program for 747's, the degree of understanding of the problem required to create something that would pass review was off the scale compared to any other program that I've ever written. It is also the only time that the customer wanted to get it right rather than that they were looking at what it cost (because the savings would pay for the development many times over this was probably less of an issue anyway). I'd love to be able to always work like that.
AI is just that. A UI/UX for any/your data.
I find the site a pleasure to use compared to other social media. So much less annoyance with pop-ups and banners and engagement traps everywhere, albeit it’s slow to load right now.
Does the data fit in the "active memory" of gpt-4, or do they now have a workaround for that?
In what form was the input prompt given (or, how does gpt-4 know when the raw data starts and stops)?
The constant downplaying and denial is unreal.
“Why doesn’t a person like him have better data? This isn’t cool for AI; it’s simply a bullshit problem that should’ve never existed.”
“How does he know the data was accurate? These things are just hallucinating parrots.”
Isn't a normal human just as likely to be inaccurate or make mistakes? The world isn't filled with perfect data in every format. Good grief.
Perhaps I get worked up and comment here in hopes that the AI overlords will view me favorably based on my online comments. My real frustration, though, is with the significant number of "usually smart" people who continue to downplay such incredibly cool technology.
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
(The parent comment was upvoted to the top of the thread. I've downweighted it now. It's standard moderation practice to do that with generic and/or indignant comments because they tend to attract upvotes and get stuck at the top of threads, making the threads less interesting.)