I recorded a screen capture of a task. Gemini generated code to replicate it
twitter.com
twitter.com
I've been scared of AI since seeing chatgbt a couple of years ago. I feel like it's only a matter of time until a dev can feed an AI machine it's entire code base and business requirements. And then a separate AI could carry out the manual/integration testing tasks. AI could potentially cut down the number of devs required to maintain a webapp or ios app after it's built.
I feel triggered by this post especially because I've made it a career writing automation code haha.
Who can say how AI will develop, but beware of happily-ever-after stories. It could be a nightmare for all of civilization, for just engineers, or just not develop much futher.
Ideally, skepticism has none of that: Denial is arguably not part of skepticism - it lacks skepticism of yourself (i.e., denial includes certainty). Skeptics has nothing to do with anger, or bargaining (you can't negotiate truth); and depression and emotional acceptance also are out of place.
Skepticism is unemotional in that its definition doesn't reach into emotion. But skeptics - all humans, as far as we know - are emotional creatures, and beyond a doubt those emotions play a role in driving much skepticism.
Which returns me to my point: Many responses to AI are driven more by those emotions than by skepicism.
Or it will reduce the demand for inexperienced juniors and offshore teams since you can replace their slop with AI slop.
Every company is looking exclusively for seniors or at least mid level. They don't care that someone needs to train juniors in order to become seniors as long as it's someone else who has to do it. Companies can stay irrational longer than you can stay solvent.
So everyone keeps telling me how hot the SW dev market is right now due to all the openings and the high demand, meanwhile I'm only getting rejections because I don't have 3+ YoE in AWS, Kubernetes, Django/Flask, System Design, etc.
The gap between seniors and juniors just gets bigger.
Seriously, the demand for skilled handymen in the cities is insane as people can't do shit anymore. As per the South Park episode, you have to treat them well if you want them to pick up the phone or return your calls.
Unfortunately your status in society and the way you get treated comes from how much money your profession makes. SW devs were also not very popular until google and meta started the bidding war.
What's the useful part of my job? It's bringing values to my customers: they want to do something, and they need help. But today, unfortunately I must always tell them to reduce their ambition because they wouldn't be able to afford it or to be stuck waiting for the project to deliver even if they could fund it.
I've seen a enormous shift in developers ability to actually ship valuable products to customers when opensource became mainstream, thanks to github and things like npm: no more custom half-backed libraries for doing all the features necessary for the product to work, we could just re-use an existing library. More than half of our job disappeared these days, yet nobody regrets the days when you had to write your own code for absolutely every features, and the number of programmers exploded since then[1].
I wish AI assistant could be as impactful as github and npm, and I'm pretty sure they will eventually, and that day will be a great day for developers, not a bad one.
We're not going to lose our job, because from the perspective of the dude who hold the money, our job is to be the weirdo that talks to machines to deliver what the his ego want his company to be. Hence, the more you are able to deliver, the happier the money holder, and the more money you make. Our job will be threatened when the guys with money will be willing to actually make the effort of talking with the machine, but I don't see this day coming anytime soon. The ambition of man is unlimited, but his will to make efforts is in scarce supply.
The only realistic risk with AI, is big corporations grabbing all the benefits of the added productivity, this is a serious risk, and a very good reason not to be using OpenAI so we don't trigger a self-reinforcing feedback loop that give them a monopoly position where we all end up losing because we depend on them.
[1]: this has caused lots of sustainability issues for the library authors, but not for the developers using the libraries.
There's plenty of managers who prefer hacking an excel and vba than relying on developers; those will probably be more than glad of using some IA tool, to do something more advanced; of course mostly just for internal stuff, in the short term.
Anyhow many managers have a distaste for developers; they consider them overpaid slackers who'll lose time on anything, rather than producing money for the business.
They'll be more than glad to turn the job to the cheapest unqualified guy around, if with an LLM he can make something that looks passable
It's extra funny when you think that resilience was the first thing Sam Altman answered when asked what kids should be learning today .https://youtube.com/shorts/OK0YhF3NMpQ
Seeing the same patterns with AI. Every startup is now incorporating “AI” or “deep learning” or “OpenAI” into their decks/motto/pitch.
Have yet to see anything worth using beyond the initial hype. Using AI to me is like learning another programming language. Same shit. Different interface.
Proven to do what exactly? A solution for which problems?
Cryptocurrency had billions invested into it and the most valuable product that came out of it was cartoon apes. There's not a single blockchain that replaced the existing financial system, an area it was supposed to disrupt. There's not a single blockchain in use anywhere in logistics and supply chain, an area the blockchain was allegedly going to revolutionize. There's not a single blockchain being used for identity management by serious enterprises.
Most new technologies are overhyped, but blockchain/cryptocurrency is the only one where you can look back and be astonished at how virtually nothing was created. Its most important lasting contribution to society is providing incontrovertible evidence that just because "thought leaders" and deep-pocketed investors say something is going to change the world, it doesn't make them right.
Snowden is correct here: https://twitter.com/Snowden/status/1759304612664779247
AKA micro transactions in games.
Barely, if at all. Technically it can be decentralized, but Bitcoin's design promotes centralization, and the result is that currently Bitcoin can be controlled by either two entities (AntPool and Foundry USA) or three (one of the former ones plus F2Pool and ViaBTC) - and it's not going to get better.
I don't think the thesis of most Bitcoin critics is that the math and social engineering involved aren't interesting.
I used to use BTC quite regularly for payments back in like 2014-2015. But now, it's too expensive to move around due to fees so I just hold it. The same can be said even of ETH - the fees are too high to interact with a lot of the DeFi stuff.
There are 'layer 2' solutions that tackle the fees issue but uptake of these has been slow.
It's why they pivoted to calling BTC a store of value years ago. The payments side of crypto is kind of disappointing.
What's the risk-adjusted return and correlation with other securities? Those are both more important than share price.
Since it doesn't pay dividends in USD, its USD value is only achieved if you go around convincing everyone else to stop selling it so that you can sell it. Which I think conflicts with it being used for payments.
1. type out yourself 2. copy and paste from StackOverflow 3. find a library that does the thing you want
It's not like CoPilot is any better. It's like when Microsoft and Google tried to force text completion on emails, it just gets in the way and makes me lose my train of thought.
AI is really great for very specific tasks that would be difficult to incorporate into a traditional algorithm. I really like Photoshop's background removal tool for example. But general purpose AI to me is blown out of proportion in terms of hype. Not everything needs iPhone levels of scaling. 3D printing, VR, Web 3, Cryptocurrencies, NFTs, Metaverse, AI. The list goes on. These things have niche use cases. AI is great for a lot of niche use cases (video upscaling, for example). But for general purpose software interaction? Maybe not.
I have been using a paper notebook to take my notes for a while and I like that I can remember what I scribble spacially.
Recently I decided to also use the notebook to sketch the really important e-mails - the ones you send to people either really high up or that you value a lot but can't reach often - in paper. I have been able to scribble rather quickly in paper and come up with concise but also complete write-ups and also noticed I am happy to not be looking at the computer screen.
I started this because I noticed there was a lot of noise going on when using outlook with all the notifications popping-up and the hard to understand new interface that just scales like ass and becomes unreadable in my 4k laptop, and the autocomplete kept axing my thoughts.
Amongst other things, one of the tasks I tried to put ChatGPT to was writing scripts in a not so popular dialect of a very popular language: PowerCLI (PowerShell). Gemini is even worse!
The issue is of course the relative lack of PowerCLI vs the huge body of PowerShell generic stuff. Hallucinations include invented function parameters and much worse. It doesn't help that PowerCLI and MS's Hyper-V effort (whatever that is) both have a Get-VM function etc.
These things are "only" next token/word guessers. They are not magic and they are certainly not intelligent. I do get great results in other domains and with a bit of creativity but you have to be really careful.
No need to feel triggered. Use these tools as best works for you and crack on but do be careful to be an engineer and critically examine the output from the tool.
This is precisely the kind of vacuous "this is technology, I know technology, this is simple" hubristic underestimation that's being called out.
There is no upper bound to the intelligence of a "next token/word guesser". You can end up incorporating an entire world model to your predictions to improve their accuracy, and arguably this has already happened, to a currently-unreliable and basic level. It is possible that no technological advances are required to reach better than human intelligence from this point -- only more compute, bigger models and datasets, and (therefore) better next-token predictions.
So I am devoid of anything? Nice. I slapped "only" within quotes to imply that there is more going on and a lot more complexity than implied by a naked reading of my comment. I'm sorry you missed that.
There is no notion of a bound or even intelligence for a LLM. It is a tool and no more - we know how they work - that is defined and we run our own. We can marvel at what looks like intelligence from the outputs but it isn't that. They can be enbiggened ad-nauseam but I very much doubt we'll get intelligence per se.
You might disagree with my arguments but please don't describe me as vacuous.
The completion was "vacuous [..] underestimation". Something you appeared to be doing in this one phrase you wrote, not something you are. I don't know anything about you as an entire person. Please try not to take criticism of some of your written thoughts so personally. I continue to take issue with characterizing the LLM as "a word guesser" because it implies a limit to capability that I don't think actually exists.
> There is no notion of a bound or even intelligence for a LLM.
When LLMs are outscoring humans on many/most standardized tests, including for tests where the questions are novel, I also disagree that there is "no notion of intelligence for an LLM". It feels like goalpost-moving, to the extent that I now have no idea what you actually mean when you say intelligence.
> we know how they work
I think this is also a hubristic statement. The researchers working on these systems do not speak like this. They say things like "we did reinforcement learning on question-answering in English and it turns out it answers questions in French too now and we were surprised and can't explain why that happens".
Services! All sorts of things are going to become cost-effective that currently aren't.
Want a personal trainer? Motivational coach? Someone to sit next to you slapping your phone out of your hands every time you open social media? Personal runner doing errands? You can afford that now!
We're going to have an ever-shrinking pool of highly critical un-automatable people with an army of support folk keeping them running at peak productivity at all times.
You already see this trend in people who think of themselves as a business. They hire everyone from personal assistants to nannies. All in the name of "Well I make $200/h and there's this chore that costs only $50/h to delegate ..."
OpenAI + any half-assed data broker could easily infer the company, as I am sure they have already done.
All hail Microsoft, I am glad I chose the right AI megacorp overload early.
You either let a robot tell you how to do your job better, or someone else will.
Now that the means, motive, and opportunity are there, combined with the a general uneasiness regarding employment opportunities, gains have definitely been on a sharp uptick.
Whether that's a first-order effect of people using generative models or a second-order effect of people believing they will be replaced by those who do; either way, the pressure is real, and the gains are material.
It may take a larger timespan and more samples, but I have little doubt middleware and other glueware is being rapidly "no-coded" by GPT models on the private computers of contractors.
What's amazing is the social impact this has - often people don't believe it's real. It feels like when I had to explain to my parents that in my online multiplayer game, that the other characters were other kids at home on their own computers.
I think it's a matter of denial. Yes, software is made for humans and we will always need to validate that humans can use that software. But should a human really be required to manually test every PR in 10k person teams?
Again, as a founder of an AI Agent for E2E testing, we work with this every day. If I was a QA professional right now, I would watch the space closely in the next 6 months. The other option is to specialize in the emotional human part like in gaming. You can't test for "fun."
1. https://testdriver.ai. Demo: https://www.youtube.com/watch?v=HZQxgQ1jt4g
Sounds intuitive, but there are gaming researches working on that regard. Two related terms (learnt from IEEE Conference of Games) that come to mind:
1. Game refinement theory. The inventors of this theory see games as if they were evolving species, so this is to describe how game became more interesting, more challenging, more "refined". Personally I don't buy that theory because the series of papers had only a limited number of examples and it is questionable how related statistics were generated (especially the repeatedly occured baselines Go and Mahjong), but nonetheless there is theory on that.
2. Deep Player Behaviory Modeling (DPBM): This is the more interesting one. Game developers want their game to be automatically testable, but the agents are often not ready or not true enough. Says AlphaZero for Go or AlphaStar for StarCraft II, they are impressive ones but super-human, so the agnet's behavior give us little insight on how the quality of the game is and how to further improve the game. With DPBM, the signature of real human play can be captured and reproduced by agents, and thus auto-play testing is possible. Balance, fairness, engagement, etc. can then be used as the indirect keys to reassemble "fun."
but this solution only appears to do e2e testing ignoring api and unit testing. Additionally, automated test are mostly used for regression testing not exploratory testing of new features where most bugs will be found.
Why do we need web apps or iOS apps? They're just task specific computer interfaces.
It's possible that this kind of AI tech eliminates the need for task specific computer interfaces at all.
You don't need to tell the LLM to code up a TODO list app so you can sell it in the App Store.
The user doesn't even need to tell an LLM to make them a TODO list app.
Given an LLM that can persist and restore context, the user can just use the LLM as a personal assistant that keeps track of their TODO list.
Whatever the software is that we're working alongside our AI colleagues to build in ten years' time, I don't think it's going to be automated tests for apps and websites.
I would expect better results from LLMs using programming languages, perhaps ones tailored to LLMs, to prepare tasks on behalf of their users.
(Also, LLMs doing anything direct are an incredibly inefficient use of computing resources. There are quite a few orders of magnitude of difference between the FLOPs needed to do basic calculations and the FLOPs needed to run inference on a large model that may be able to do those calculations if well trained.)
We force accuracy on them when we computerize them, because computers historically haven't handled ambiguity well. That demands the skill of 'programming' - interpreting fuzzy real world problems, making them precise enough to be modeled in a way that classical computing can handle, and making computer routines to help.
But the underlying problem humans are looking for a bicycle-for-the-mind to make easier often didn't start off 'precise' at all.
AI is a slave in service to corporations, and those corps have a culture of three decades of depraved sociopathy, with comical levels of contempt for customers, people, humanity, and civilization.
AI is not your friend. Whatever service it provides is a Trojan horse or, at a minimum, surveillance.
> Case in point: the Arc Browser.
> For years, The Browser Company has been promising to save the internet. Its Arc Browser is a smart refresh of what a modern gateway to the web should look and feel like and it generated a lot of goodwill with early users. And then, earlier this month, they released their AI-powered search app, which “browses the internet for you.”
> The Browser Company’s new app lets you ask semantic questions to a chatbot, which then summarizes live internet results in a simulation of a conversation. Which is great, in theory, as long as you don’t have any concerns about whether what it’s saying is accurate, don’t care where that information is coming from or who wrote it, and don’t think through the long-term feasibility of a product like this even a little bit.
> But the base logic of something like Arc’s AI search doesn’t even really make sense. As Engadget recently asked in their excellent teardown of Arc’s AI search pivot, “Who makes money when AI reads the internet for us?” But let’s take a step even further here. Why even bother making new websites if no one’s going to see them?
wayfair laid off a bunch of US IT staff and are trying to build an AI in india to replace workers (I think, based on events and new job opennings).
edit:typos
adt'l: tbf, AI hasn't actually replaced jobs there yet (afaik)
Art in general seems to be a high risk occupation because the stakes are low. It's a similar case for stock photography. In the gaming world, Embracer recently fired a bunch of artists with the intent of having fewer people be more productive with the use of AI generated content.
Unsure if the code was provided (logged out I only see the one post) so unable to speak to the code or how it runs, but judging by the poster’s profile I think I can trust they know when code runs correctly. Though it does appear they’re currently working at this specific agent’s parent company.
Honestly, this impresses me more in the work the selenium team has been doing than the llm that used their api.
I can’t work out how to full screen the video if it’s even possible.
The resolution on the video is too poor for me to be able to read the code.
There are ads and troll comments everywhere.
It seems to want me to create an account.
It’s just horrific.
Is there really no better way to showcase some cool observation or story about Gemini doing something useful?
https://www.microsoft.com/en-us/power-platform/products/powe...
You can write just a few sentences, or even upload a drawing of what you want your app to look like, and it'll create it for you.
The catch is it's only for corporate apps. These can't be publicly facing. Everyone needs to sign in with M365 and have an individual license.