I built a plane spotter for my son in minutes
louison.substack.com
louison.substack.com
https://github.com/search?q=plane+spotter+language%3AJavaScr...
This topic has been beaten to death though, and it isn't changing. Companies have little incentive to have explainable AI since it would open them to infringement claims.
At least this person does not fetch all planes and then iterate over all of them but sets boundaries in which area to get all the planes.
I recognize that there are questions about how that works in the longer-term, but for now, it's useful.
Usually the original answer has some responses attached either clarifying something important or outlining why this approach is bad and linking to a better answer.
With GPT you only get one answer Saïd with complete confidence.
It's saved me quite a bit of time in terms of understanding how to use APIs as a client. I find the API docs that Apple provides are so bare bones and don't illustrate how the APIs are meant to be used, and you really have to search around for example projects or That One WWDC Video Where All Is Revealed. Asking ChatGPT first saves a bunch of time having to do all that searching.
This response in particular.
I immediately thought of this project (that hit the front page of HN):
I do see how this is a little too idealistic, though.
They don’t do half as well in large codebases that use ad-hoc frameworks. For example, it has no idea how to retrieve the currently-logged-in user object for a new endpoint you want it to build.
The solution there is to fine-tune it on your codebase, but that’s likely a few years away for the average LLM user.
It knew the data model and methods that existed in code, and used these.
In this video they say "Custom-tuned models on code" are coming soon, but aren't there yet: https://youtu.be/qbxj_JDQ4Lk?si=qSTzMBpmqqJD5o9u&t=1632
The names are very similar.
Just builds a local vectordb with your code base
Currently working on an Elixir project with Ash Framework, and I wouldn't trust a single Chat GPT output for either of them.
Greenfield? Unless we disagree on the definition of greenfield (no prior work), then a plane spotter app is anything but greenfield. GitHub returns 11 repos for JavaScript alone.
The wider context window and better tooling are all being worked on.
A more true title would be "Displaying nearby ADS-B data for planespotting in 120 secs...".
Flight TAP1369...
Flight TAP23NP...
Flight TAP1691...
Flight RZO134...
TAP1369 (ICAO) is TP1369 (IATA); similar for TAP1691 (ICAO) and TP1691 (IATA). But TAP23NP (ICAO) is TP1823 (IATA) and RZO134 (ICAO) is S4 134 (IATA). What's broadcast in ADS-B is very often very different from what's shown at the airport and not what someone would recognize.Doing this mapping is a pain. The OP is using OpenSky Network who say:
callsign
This column contains the callsign that was broadcast by the aircraft. Most airlines indicate the airline and the flight number in the callsign, but there is no unified system. In our example, the callsign indicates that this state vector belongs to UPS flight 858. By looking up the flightnumber on services like flightaware.com, you’ll find out that this flight goes from Lousville to Phoenix every day.
i.e. go look this data up elsewhere.
I asked ChatGPT:
Write me some JavaScript that can map an ICAO callsign broadcast by ADS-B to an IATA flight number for the same flight.
Mapping an ICAO callsign broadcast by ADS-B to an IATA flight number is a complex task that typically requires access to a comprehensive aviation database. Such databases are often provided by aviation data providers and may not be freely available. However, I can provide you with a simplified example of how you might approach this task if you had access to the necessary data.
I asked it to create me a python library for the GT1151 touch screen controller and it came up with working code on the first try. Then I asked it to add support for that chip’s gesture mode and it came up with completely reasonable looking but nonfunctional code, because it didn’t understand the gesture mode implementation on the chip.
similarly, what about the vast number of times a prompt like this this _will_ have worked for other people?
we don't know either of these two numbers, especially not which one is larger than the other
But I also suspect that it overvalues individual insight and undervalues the normal process by which most complex ideas with application to the "real world" actually get created - through back-and-forth interaction between people.
I’m not sure if I agree with the second half of that statement. Several times I’ve asked GPT-4 to suggest ideas, and it came up with some very good ones that I would not have thought of myself. Here’s an exchange I had with it just now:
https://chat.openai.com/share/d23ce0d2-ed60-4259-9176-73b590...
I’m also not sure if I agree with the first half, particularly the “great” in “great at executing.” While I have also been amazed a few times by code it wrote for me, I have run into more cases where, despite repeated back-and-forths with me, it was never able to produce code that ran.
Maybe better, in my experience, would be “ChatGPT can often come up with good ideas, but it is only occasionally able to execute them.”
I checked the city I live in and three more cities I know quite well, and I have to disagree. This reads/feels like a content farm, which (if GPT wrote the content, as you seem to imply) is exactly what it is (as I have excessive Adblock I don't know if there are ads, but I don't think they are needed for the definition of a content farm).
And there are no ads. It's really more of an experiment rather than a real product.
A nightmare because we have “devs” using it to develop applications laden with vulnerabilities. As an example we found a web app with account enumeration/forced browsing vulnerabilities that would have led to a sizable chunk of our user base being exposed. It would later come out that he had copied and pasted everything from ChatGPT.
When I asked him if he trusted everything he found on StackOverflow he said no as if I was an idiot. When I asked him why he trusted ChatGPT he said “because I asked it to check the code for any vulnerabilities and it said there were none”.
But, like I said, it’s also been a dream. I’ve seen an uptick in sites with poor practices on bug hunting platforms and I’ve banked about 35k. So that’s nice.
Asking it to split it into smaller functions so it can be emitted one at a time had mixed results.
But the future is indeed bright for such tools.
When you pay, the results come much quicker and are far less likely to timeout.
The knowledge resource is no longer centered around the teacher, but rather on the student. The key is to have a teacher that is keen on noticing what interests the students, then converting those interests into meaty projects that students can build.
The key is going to have very small learning groups, maybe max 4 per adult. The adult needs to be able to quickly learn new topics with the help of LLMs, and be able to articulate what the students interests are.
The value of teachers who can effectively leverage LLM is going to sky rocket.
And I could refine the generated configuration with 2 or 3 other types of loglines... I'm now a convert.
> Imagine envisioning Airbnb, and having its whole frontend and backend done within 1 minute.
...sigh...
You're seeing the tip of an iceberg, and you're like, "wow, it's cold"... but you honestly have no idea.
Any tool that you spend less an hour using, you have no idea about.
That's it. There's nothing more to say.
Use it for longer. Try building larger things. Give me a considered opinion when you've formed one, not a reaction video.
There was a spate of this kind of article when chat-gpt first came out, and at the time, it was like: why are we seeing these "I spent 10 seconds with a LLM and I made a potato!" and "I spent 20 minutes with an LLM and made a one-page HTML website!" articles, and none of the "I spent a month with an LLM and build a new programming language", or "I spent two weeks with an LLM and I build a raytracer" articles?
Oh they said, "It's too soon, give it time... it's only been out a month. 3 months... 6 months...".
You still don't see them.
...because they don't exist.
No one has done anything impressive with this stuff; it's fundamentally limited in what it can produce, and the massive productivity benefits you see (30x faster!) are for trivial tasks, not difficult tasks.
...and the modest productivity benefits you get from using an actual copilot don't make articles that are nearly as interesting to get as many clicks.
Look, AI can seem magical, but when you have something that seems too good to be true, after using it for an hour, or even a day, maybe it's not the right moment to drop a blog post about how gosh darn amazing it is?
...and even if they were, wouldn't it be nice to focus on solving the other problems you mention instead of writing the code?
- https://smaller.fish/posts/l_plus
- https://www.kdnuggets.com/2023/04/datalang-new-programming-l...
write a simple raytracer in python
into chatgpt and it will return fully working code of a simple raytracer that can handle spheres and Lambert shading.To be fair you can type the same thing into google and get a link to fully working ray tracer in python back, so maybe not the best example.
edit: What would be interesting would be to see how far you can push ChatGPT by asking for more and more advanced features before it stops giving valid answers. Basically what is the most advanced raytracer currently 'stored' in ChatGPT
It pulls in the unknown unknowns.
However, there's many reasons it would be good at trivial tasks, and arguably most programmers spend most of their time on said trivial tasks, its just the count of them that's overwhelming; switching between multiple different languages, looking up syntax (at least for me!), and deciding on what's a sensical solution for the use case. Nobody cares about the performance characteristics and just wants it to work.
So anyway, I don't think a blog post about how you build a new algorithm to come out of any of the LLMs of today, but that doesn't mean I find them underwhelming.
At least that's how I use it and I'm super happy with what it gives me.
I've built entire GitHub Actions to help improve team effeciencies. Built validation rules. A bunch of Python scripts. Parsed a multi-level JSON response that would have taken a few hours, etc...
All in minutes. Imagine knocking off like 15 of these tasks before noon.
What I find most useful in this space is twofold:
1. Specificity: so yeh, I can look at SO or wherever and see responses to a semi-specific query, say about querying a database. And yes, I could post a question and maybe after a few hours if I'm lucky get a response. But with ChatGPT I can throw in the actual table and column names and get a specific response back. Most often it needs honing, sure, but it's most often a great starting point.
2. Being able to back-query. So the fact that history is maintained means I can say "actually, add in a column called X and can you make sure that the query only returns unique IDs" and - again - I'll get an immediate response without having to go all the way back through the questioning journey again.
As with any tool, the extreme responses on either end seem obviously flawed. On the one hand: no, it's not going to replace all developer jobs. On the other, to just completely ignore it, not try it, always assume that "the documentation" is better - this seems blinkered and foolish too. Somewhere in the middle is a sweet spot where human brains collide with these tools and we manage to tease out code that is more quickly researched and saves time but also well considered and treated with a degree of healthy scepticism before deployment.
‘Meh. Let me know when your cat can recite T.S. Eliot’
> Imagine envisioning Airbnb, and having its whole frontend and backend done within 1 minute. We’re not there yet, but we could be there soon.
The difference in scale between toy plane spotter and an Airbnb front end and back end is simply enormous, with a combinatorial explosion of possible states and failure modes. Our existing tech will not scale to that level by just adding more parameters and GPUs. We may or may not have a major paradigm shift in the future that enables it, but this talking cat ain't producing Airbnb in 1 minute.
While its knowledge is impressive, it continues to spout nonsense fairly easily.
You can write programs in markdown, it's a lot like literate programming that Knuth envisioned. I've been using this technique a lot for side projects lately --> https://swizec.com/blog/programming-in-markdown/
ChatGPT is able to analyze JSON dumps of dependency graphs and identify communities of files that should be modules. This sounds promising and I'm figuring out how to make it more practical for real use --> https://swizec.com/blog/finding-modules-in-a-big-ball-of-mud...
And while I haven't written about it yet, I've been able to get GPT-4 to analyze a blogpost and suggest phrases you should link to your past articles. Needs a little more tweaking with the prompt before prime-time because right now the result, while good, lacks some zhuzh. Whole thing was cobbled together using the markdown programming technique I mentioned in the beginning.
I'm at the point now where I'm so tempted to turn markdown programming into an executable script that can save me some copypasting and perhaps where I can say "Inject other markdown here" (basically functions), or "Go look up documentation at <link>" (that's been a painpoint).
So yeah there are dozens of us spending time with these things. Dozens!
Speeding up trivial tasks like this is game changing for many things. The amount of "trivial" things I've built with chatgpt that would've just stayed as an idea in my notes is crazy already.
The thing?
Depending on how charitable I'm feeling, either "internet comments" or "mansplaining".
But even if its expertise mainly opens the door to trivial stuff, that alone is a huge boon for accessibility for a lot of folks.
> I spent a month with an LLM and build a new programming language
For the past few months I've more or less been doing this. I'm self-taught, so my CS knowledge is sketchy and has lots of holes. I never took a compilers or programming languages course. I've done lots of reading but it never fully clicked for me. Having an LLM outline the process for me at a high level, provide pseudocode or even code snippets for less clear bits, and also (try to) answer whatever questions I throw at it related to my external reading has been pretty incredible -- things suddenly clicked. Another way I use it is to ask for important introductory texts or papers when exploring a new area
Yes, occasionally it spouts nonsense -- you have to evaluate its responses critically. If you ask it to do anything complicated or involving fast-moving APIs or less well known languages or frameworks, it will straight up make shit up. But we shouldn't throw the baby out with the bathwater here
A snippet of a conversation from earlier this summer:
Me
In Pierce's types and programming languages, he describes the untyped arithmetic expression language.
In one of the definitions by inference rules, he defines the language as the minimal set T such that
- 0 is in T, true is in T, and false is in T
- If t is in T, then succ(t) in T
- pred(t) is in T, and iszero(t) is in T.
That makes sense if t is numeric but how are these defined if t is Boolean?
For example, true is in T, but I don't see how this definition can work
if T is to also allow boolean values, like how is succ(true) well-defined?
GPT In Pierce's presentation in "Types and Programming Languages", the untyped arithmetic language allows you to form terms like succ(true), even though they don't have a meaningful interpretation in "everyday arithmetic". When we say the term succ(t) is in T, we're referring to the syntactic formation of terms, not their semantic meaning.
The language defines the syntax of valid terms, but it doesn't give them any inherent meaning at this stage. This is why we can syntactically form terms that don't make sense semantically. When we actually evaluate these terms (if we were to interpret them), many of them would result in "stuck" states, where we can't proceed with evaluation. For example, while we can form the term succ(true), there's no evaluation rule that tells us what this term evaluates to.
The definition you cited doesn't discriminate based on the nature of t. All it says is: if t is a term in T, then succ(t), pred(t), and iszero(t) are also terms in T.
Remember, this untyped language serves as a stepping stone. It illustrates what can go wrong without types. The oddities you're pointing out motivate the introduction of more refined, typed systems later in the book, where such nonsensical combinations are explicitly prohibited by the type system.Okay, I'll bite. I think quite often solving difficult tasks contain lots of trivial tasks. And speeding the trivial tasks can speed the difficult task up considerably when you can outsource the trivial tasks to an LLM.
I do not know how to make a catchy short video on my last side project scratching my itch (VR environment/metaverse to build 4 dimensional objects and move around them in 4d, see https://www.brainpaingames.com if interested). It was multi-year thing to get it where it is now (not pretty yet, I admit...), and it may of course just me being dumb, but yes, I consider it was/is difficult. Especially given I have never done anything in either C# or Unity. The speed-up ChatGPT gave to this project during last half a year was absolutely not trivial. I guess I must have saved years of my limited side-project time to get this far.
The whole AI hype reminds me of the ancient Greeks going to the Oracle at Delphi to obtain answers.
From the perspective of an outerspace alien, what is the difference between asking a god like figure for answers or an AI is some data center?
How maintainable is software that was not fully understood when it was built? When it breaks in 6 months because some api changed. I guess the solution will be to get AI to fix it ... Sigh.
I thought maybe ChatGPT would replace me, but so far the same people come so I can prompt it for them.
I've spent more than an hour with the tools, ChatGPT and SD being my two big buckets and it's great for productivity within some areas, mostly in fast mockup/art and creative writing, mostly fiction level.
I've use it to format and transform my own work, and as a tool it's useful.
Like any of these tools, early adopters get an edge for a bit and then the market adoption rolls out and we equalize.
I see it mostly filling in the 80% of the easy stuff, with humans stepping in for that last 20%, which has been my approach with good success.
However the hype train around it has gone past annoying and into downright scammish/spamish marketing.
2. Automating away well-defined trivial tasks frees up more time for figuring out the difficult tasks.
3. If you're running solo, being able to ask for alternative perspectives or to validate ideas - even if that's against a compressed archive of Reddit answers - is still valuable.
4. How is this better or worse than making progress through asking questions on X/Reddit/mailing lists/IRC/Usenet/your local library? I'll tell you: it doesn't irritate other people as much, and it's likely a lot more efficient. I get it, "I spent 20 minutes with an LLM and made a one-page HTML website!" doesn't sound impressive, until you compare it - about a year ago - "I spent two days going through awful ad-laden tutorials and made a one-page HTML website!".
The big cognitive gap people might be running into - and I think you might as well - is that people think these LLM tools are there to replace people and to do things that weren't possible before. They specifically aren't. They're tools to do stuff we know how to to do, but a bit faster.
Nobody is making explicit claims of magic. Nobody - other than breathless media - thinks this will replace people doing the jobs they do today. It's just going to make it a bit faster and quicker to get some stuff done.
They're better screwdrivers, not replacement geniuses.
May I ask, have you tried to do something difficult with it? Have you concluded it's impossible, or are you guessing that it isn't due to a lack of signal? I think it might actually be impossible, but I still think they're valuable tools for the reasons above.
I think this is the money-shot. LLM's (specifically ChatGPT) have helped me debug weird issues, and help get started with new technologies / libraries where searching for the issues on Google did not yield (good) results.
Could you or anyone provide examples of this?
Because in my experience, it's not true. I've found that difficult tasks may have parts that are trivial, but always have a truly difficult core, which is what makes them difficult in the first place.
Whereas there are plenty of long/boring tasks that can be decomposed into individual trivial tasks, but the whole point is that nobody's calling them difficult though. Just long and boring.
But maybe I'm misunderstanding, so I'd love any counterexamples.
I'm going to suggest hard problems are solved by breaking them down into a set of smaller problems.
Perhaps a hypothesis or two and some experiments that need designing. Or perhaps you need to understand a problem from different perspectives, or research whether something similar exists in a different domain.
Aristotle, Euclid, Copernicus, Galileo, Darwin, Tesla, Edison, Turing, Einstein... not one of them had an entire solution to a hard problem revealed in a single gulp. Every single one of them took small iterative steps and needed to break a problem down and approach each part of it as an individual problem in its own right.
You might be doing an experiment nobody has ever done before to test a hypothesis that nobody has ever considered in human history, but LLMs can still help you determine if the hypothesis is framed correctly, understand if your experiment has parallels elsewhere in the knowledge its exposed to, research how to validate your results, and help you with a ton of other small mundane tasks needed to do good science. Doesn't make it a scientist, just makes itself useful to scientists.
Likewise, you might not know how to write a Phoenix LiveView application to predict in-play odds of your favourite sports team and identify value in online sports books (trust me, this is a hard problem), but it can help break that down into the individual pieces, let you get started with small utility functions that you can build on, and work with you to make you more productive. Doesn't replace the work you need to do as an engineer, just makes itself useful to you as an engineer.
I don't think that's true at all. To the contrary, it was a ton of thinking and experimentation and then getting really lucky with major flashes of insight -- which is the very opposite of breaking something down into tractable parts.
Einstein coming up with general relativity wasn't something that he did, or anybody could have done, by gradually breaking the problems of gravity down into parts that were then straightforward to solve. That's not how his discovery worked at all.
Difficult problems are difficult precisely because they can't be solved in an easy straightforward way of breaking them down. They seem quite impossible to solve until you try a bunch of things, sometimes for years/decades, throwing stuff at the wall, and you hope you get lucky. But many times (usually?) you don't. That's what makes them difficult.
... is both breaking things down into small steps, and something that an LLM can help with.
> They seem quite impossible to solve until you try a bunch of things, sometimes for years/decades, throwing stuff at the wall
Do you see what you did there? That's breaking things down and trying lots of small things.
What you seem to think I'm suggesting - which I'm not - is that solving hard problems is linear once it's fragmented.
I'm suggesting hard problems are only solvable through fragmentation - breaking them apart - but I'm not in any way suggesting that this means they're solvable through a simple linear thinking process. Fragmentaiton and linearity are not the same thing.
If anything, LLMs can help make non-linear thinking more efficient by getting you out of small areas of focus, and as I said originally helping you explore a problem through metaphor, different perspectives, different domains, and so on.
> Do you see what you did there? That's breaking things down and trying lots of small things.
No it's not, that's my whole point.
If you're trying to find a filament that will work for a commercial light bulb, then testing out 1,000 materials is not breaking anything down to solve the problem. Instead, it's trial and error. They're literally the opposite approaches.
Some problems are solvable through fragmentation. Many others are just fundamentally not, and these obviously tend to be the more difficult ones.
> If anything, LLMs can help make non-linear thinking more efficient by getting you out of small areas of focus
That's literally the opposite of breaking things down into small steps. So now I don't even know what you're arguing anymore. But also, I don't see how LLM's help that in the least. I have not seen any examples of LLM's demonstrating "non-linear thinking". They literally work, well, with linear token prediction -- one token at a time.
No academic paper ever written, invention ever created or piece of art I can think of or heard of (including general relativity, which was referenced earlier in thread), was solved by just sitting there and thinking about the big hard problem and solving the big hard problem in one go.
The closest I can think of what you're referring to is an accidental discovery, so I think that's where our lines are crossed.
At its core I still don't buy that LLMs are useless in the context of solving hard problems. You disagree. Time and experience will tell, but it seems more likely that successes will be attributed to them than not.
And that's before you get into all the issues of suddenly steering the ai after a long game into providing scoring reliably, sometimes it confuses the scenario from the game proper, or hallucinates actions the user never did.
I had the prototype coded in a evening. But the rest has been a month of banging rocks to get it to work reliably start to end (well I've yet to solve scoring, but still, the rest works.)
LLMs are incredible. They really are. They can think fuzzily about human concepts, and interface those fuzzy concepts with cold hard computer things like programming languages and databases. But we need a whole bunch more ancillary technology before they'll really start shifting the social needle.
I have near-zero web dev experience before this project. After about a month, I started to rely on it less and less. But I still use it to debug when I hit the wall.
Already, I use ChatGPT almost daily to help me spot bugs. Doesn’t do a perfect job but it is at least 10x smarter than the rubber duck on my desk.
I used it to port 7 modules python code into Node/JavaScript. I’m familiar with JS. It did a reasonably good job.
And I use it all the time to prototype. Four weeks ago, I knew nothing about chrome extension development. Three weeks ago I released an extension on the chrome store. It’s not that ChatGPT wrote the extension… far from it. But it got me past the initial knowledge gap pretty quickly so that I could focus on the parts of code that I already know how to do.
> It did a reasonably good job
…and you probably had to fix a few things, but over all it saved you time and offered a modest productivity boost.
You have clearly actually used these tools.
I’m not being negative about LLMs; they’re great tools; modest, helpful, non-world ending tools.
That’s been my experience, and I also use them daily.
…this “2 minutes and you have an Uber”, unbounded productivity for small tasks = unbounded productivity for large tasks narrative is just people failing to understand what they’re using, because they haven’t used it enough to gain an intuition of how vastly distant where we are now is from that end goal.
That’s the point.
It’s not pessimism about the future; it’s more like:
Pay attention to what we have right now by actually using it, because the shiny examples people give you are, bluntly cherry-picked.
Don’t believe the hype.
It’s good. It’s not that good.
And so you see, that’s why you and I don’t have youtube channels, and stock images of us making comically exaggerated gestures. Just so that we have a thumbnail template for:
“Is Elon actually right? LLaMama autogpt 5 about to go supernova?”
We knew this was a goldrush. At least it’s still pretty funny from a Carlin-esque perspective.
Fifty years later I was directing construction of systems of systems that could remove a tumour from a little girl's brain by taking 2D X-Ray images, reconstructing them as 3D density maps, identifying the organs in those maps, identifying the tumour in those maps, computing the best way to align proton beams in order to irradiate just the tumours, and then position the patient and beams, operate the accelerator to generate the protons and zap the tumour. It involved, literally, billions of logic statements - every one of which was trivial in its own context. Spacetime is also flat, everywhere, at the right scale. But we could orchestrate those billions of mathematical and logical trivialities into something very nontrivial because of huge productivity gains in producing them and stringing them together during my lifetime. A 30X productivity gain in producing even trivial (and the example give was not trivial) elements, is part of the path to radical innovation.
Anyone who tells me the latter: nope, it's as good as a university student in an industrial placement/internship.
Anyone who tells me the former: nope, the algorithm definitely had to learn a lot to be able to produce even a half-baked mess, and even 3.5 clearly knows more than I do about almost every subject.
So…
> it's fundamentally limited in what it can produce,
Indeed
> and the massive productivity benefits you see (30x faster!) are for trivial tasks, not difficult tasks.
For many people, programming at all is difficult.
Even for us, switching to a completely different language and API isn't a matter of moments.
Don't get me wrong, it absolutely screws up; but at 80% right-first-time (in my experience) going up to 90% just from giving it back the error messages from the compiler or interpreter, that's better than most students and much better than coming from a different domain.
The same in reverse means you should never ever rely on it for anything critical, even though it can still save time if you can perform fast reliable tests on the output.
Where I fundamentally differ in my value judgement is this:
> 80% right-first-time
Even taking that number at face value, that is really bad. Really, really bad. Not because of the number itself, but for its consequences:
(a) It's downright harmful as a teacher or tutor. If I'm learning something with little or no prior knowledge, I'm in the worst possible position to question what I'm being taught, and those numbers mean it's going to routinely make mistakes.
(b) If these numbers stay in the same ballpark, it's impossible to compose tasks. (80%) ^ 10 ~= 10%. Make that 95%, (95%) ^ 10 ~= 59%. Unless we see a 3 order of magnitude improvement in the error rate, interactive usage constantly supervised by an expert is the only possible way to use this technology.
(c) Even if it's something I know very well, is basic enough that the model can reasonably do, and not critical enough that an error will be catastrophic, a 20% error rate means I have to keep my eyes wide open.
Unless we see fundamentally new developments, it's a nice to have, but not a game changer. Intellisense level of useful, not git level of useful, let alone Internet level of useful.
I don't know how far we are from that.
But yes, for now it's a footgun for whoever tries to use it as a replacement rather than an assistant.
For how long, I know not.
But I suspect it is on par with git, despite the flaws: LLMs make you the manager, git helps you collaborate with other humans.
But how do I know that the code doesn't do something malicious? As in, calling a .ru or .cn flight tracker (and thus giving it my geo position, and the geo-positions of everyone using the app I had designed via ChatGPT) instead of a genuine, non-adversarial one?
You could say: "it's using opensky-network.org, not a .ru nor a .cn domain, it's right there int he code, it's genuine" which would be correct on the surface, but as someone who has had no connection to this space how do I know who opensky-network.org is without first going on their website and checking who they are? And thus increasing my development time and approaching said development time to a real-life solution.
And I only gave one simple adversarial scenario, one that can be easily checked and debunked, but there are other many scenarios more convoluted than that.
I suspect something interesting is going to happen the price point of tailored individual LLMs running full time drops below the threshold where modeling an existing business unit’s org chart as a collection of interacting LLMs becomes viable.
Can you have a Director of Engineering LLM managing a pool of project manager LLMs and platform engineering LLMs and system administrator LLMs and infosec LLMs and software engineering LLMs and frontend engineer LLMs etc.
Can all of those LLMs create organizational plans, give each other guidance, self correct each other, give each other feedback, and work towards a solution?
For the low end of the market, can you create a software engineering organization with one well compensated Lead Engineer/CTO managing a huge collection of LLMs as a business unit? Or will the error rate compound making the ultimate output of the business unit just noise?
Even in the space of work - We need a lot of plumbing knowledge to get access to a domain e.g. learning about the basics of beekeeping or how to build a nodejs based webapp which makes an API call - a lot of knowledge in a domain is just grind and grunt work which is being automated.
We now need to peer beyond and that's where true exploratory work in any domain lies and is useful.
I can't see the full details of the states/all API call in the code screenshot but it appears to be iterating through the entire dataset of Open Sky Network flights and for each one determining if it is within the observable radius.
That's a lot of work and inefficiency and could hit the API rate limits.
A more efficient approach would be to call the states/all API passing a bounding box ( supported by the API ) and then just eliminating any aircraft that fall within the disjunction of the bounding box and the observable radius ( the corners ).
When I ask GPT4 questions about Python, it is almost like talking to a pretty knowledgeable person. It offers opinions on best ways and eve no what NOT to do.
This might sound like a trivial conclusion at first, but as a community we should find some kind of common understanding without playing down its potential usefulness nor riding the hype train some people want us to jump on.
Disclaimer: I am the maintainer of ntfy.
[1] https://ntfy.sh
But it’s really just a really effective search engine. It works great for info that’s plentiful online and does a great job with aggregating it for specific examples. But it isn’t anything you can’t just google yourself with some more effort.
It's not a good example.
I think if proponents would use different language, there would be less complaining (but also less hype).
Most of the reason for building something like this yourself is to learn more about the APIs of the data sources, the browser, the language, etc. You get none of that when copying someone else's code, or when asking ChatGPT to do it for you. So why bother?
Credit to the creators of these systems for getting it even close, but until they can actually correctly cite their sources and provide followup references to real content instead of hallucinating titles and emitting hyperlinks to unrelated medical research papers, I'll stick to using it as a very rough overview of a subject, just to get the right collection of search terms for following up in traditional sources.
Right now, without even being allowed that extra digging, it is darn difficult to be sure about its high-dimensional musings.
What I would like to see most is the actual word/phrase/source distributions for a given prompt, in order to judge the sparsity of the underlying training data and subsequently, 'crowd source' the remaining gaps.
The other day I was struggling to parse a Japanese sentence, a particular grammatical construction made no sense to me. I wrote the sentence in ChatGPT, asked it to break it down for me, and it came up with a plausible-sounding explanation. Problem was, I couldn't find any hit on google when I searched for the thing it was talking about. So I asked ChatGPT to give me more details, tell me what I could search for, and it would insist that its explanation was correct and then gaslight me by telling me that the reason I couldn't find anything on Google was because it was a niche subject not usually taught in grammar books.
After some more searching around and double-checking it turns out that I had misread a kanji and the sentence I typed into ChatGPT was complete gibberish as a result. ChatGPT's explanation, while sounding very plausible, was complete fabrication.
The idea that some inexperienced people are shipping software using this tool is insane to me.
For example I had a massive sql query that was loads of statements unioned together, I said “for each statement remove this filter and add this filter” and it would go “certainly, here are the first four, I have used an example table name feel free to change it” then I’d say “can you do it for all the statements, not just the first four” and it would say “of course here you go” and just give me the first four but also makes them useless by changing the table name!
I’ve got great hopes that one day I can get it to help me shape the code in a way that the jetbrains ide’s can’t today - for those I have to choose from a set of available operations - I want to talk to it and get it to change the code in a set of operations that I choose!
One day maybe :)
Instead the model decides to make stuff up and pretend that it knows. That's vastly worse.
It reminds me of the early days of DuckDuckGo, when if you searched for something obscure with no matches online it would still fuzzy match some garbage like a binary blob in a Chinese PDF while Google helpfully would just tell you that it couldn't find anything.
Does the model know it doesn’t know though? Does “know” even make sense as a concept here? I don’t know if it can really introspect like that, but of course it would be so much better if it could can have some sort of confidence score with each answer.
I really think you're trying to compare apples and oranges, in multiple ways. For one, we can test the software by running it, which is a pretty different problem from asking language questions, with a much slower ability to verify correctness (based on what you describe and what I imagine).
I'm not saying your experience is invalid. In my own adventures, the equivalent of what you did was my writing some incomplete bash in an existing script, wandered off to another part of the code. I then came back to that incomplete snippet, and though it was some unfamiliar syntax someone else had written (vim even highlighted it like it was special!). Naturally I went and asked ChatGPT what that snippet did, and wasted 15 minutes trying to corroborate it before checking the git history or something and realizing my own error.
As long as the tests are not also written by ChatGPT...
Many critical security issues require a deep understanding or the code or some intense fuzzing to discover, it's not enough to ask ChatGPT "write me X" then superficially glance at the output to validate that it looks correct. That's the part that worries me. Completely broken code will be caught immediately, but subtly broken code may linger for a long time and make it to production.
And from my limited experience with ChatGPT, it seems very good at making up broken things that look superficially correct.
I don’t notice nearly as many errors when I’m asking it about things I’m not already an expert in. The most likely explanation is that it’s fooling me.
I pay for gpt4. People using chatgpt to “learn” are absolutely slurping up incorrect information without knowing it.
Imagine trying to teach yourself physics with textbook where 10% of it is completely but convincingly wrong.
This is basically how many people already use the internet, read information on random blogposts, stack overflow and more, then take that it's true for granted. ChatGPT isn't really different than reading those things in the end.
What has to change is how people treat information that they read, no matter if it's from a blog post, ChatGPT or a friend. Verify everything, before you'd bet on that it's true.
I've used GPT4 to understand topics I've had no exposure to before, and it's true, a lot of the things GPT4 writes isn't accurate in the end, but lots of things are accurate too. That you can mold the information in a different way than static browsing, makes a lot of difference.
Overall, even with some false information, ChatGPT personally saves me a ton of times and I find myself only using Google for verifying information now, not for finding information.
TLDR: trust your gut feeling.
It's great that we've got it down to 2 minutes.
However anyone who has worked in a rails shop has experienced that somehow that initial set up seems to make everything else down the line get a bit more time consuming.
In fact, I would posit that there might be an inverse correlation between how quickly you can get a blog up using framework X (including an LLM) and the long term scalability of that project as you start adding features and engineers.
Legacy rails apps can be terrifying, I can only image the horrors that will be legacy LLM apps if they can even manage the structural cohesion to become legacy apps.
The more magic is involved in getting everything set up, the less well you understand the system you just built, and the harder it will be to do what you need to after the initial scaffolding. Wiring up the initial boiler plate is tedious but gives you a complete picture of the system.
It's called "on Rails" for a reason.
I think it's an exciting time where lots of people are getting computers to help them with their problems in a way I haven't really seen since HyperCard.
I feel like this is very much in the Hacker spirit.