I work for the client side and this bothers me a lot. It's very hard to get a true honest value analysis done with all the sales influence and office politics going on.
I work for the client side and this bothers me a lot. It's very hard to get a true honest value analysis done with all the sales influence and office politics going on.
The problem with LLMs is the lack of reliability and coherence. If it gives you a wrong answer and you ask it if it's sure, in most cases it will show another wrong answer and you need to go through multiple of these hops to get something fine.
I think the go to is everyone said the same thing about the internet. Look where we are now.
AI is also different in the sense that it has already gone through several hype cycles, each followed by an "AI winter" of broken dreams. Clearly large DL models are a breakthrough, but the amount of hype and hot air is entirely out of proportion to the actual results.
That being said, these initial results are not reassuring.
1. Take the documents, chunk them up on paragraph then sentence then word boundaries using spaCy.
2. Generate embeddings for each chunk and cluster them using the silhouette score to estimate the number of clusters.
3. Take the top 3 documents closest to the centroid of each cluster, expand the context before and after so it's 9 chunks in 3 groups.
4. For each cluster ask the LLM to extract the key points as direct quotes from the document.
5. Take those quotes and match them up to the real document to make sure it didn't just make stuff up.
6. Then put all the quotes together and ask the LLM not to summarize, but to write the information presented in paragraph form.
7. Then because LLMs just can't seem to shut up about their answer make them return JSON {"summary": "", "commentary": ""} and discard the commentary.
The LLM performs much better (to human reviewers) at keyphrase extraction than TextRank so I think there's genuinely some value there and obviously nothing else can really compose english like these models but I think we perhaps expect too much out of the "raw" model.
It's also all long context models getting pages of data, which even for these flagship ones is certainly just RoPE or similar which is a cheap hack but isn't super accurate [0]. 4o is best and still showing haystack benchmark accuracies below 80% and Gemini is just completely blind. That certainly needs fixing up to 100% before we can say for sure that nothing will ever get skipped.
[0] https://preview.redd.it/rlgauej7ve4d1.jpeg?width=1086&format...
For my uses, I find AI to be a much better search/answer engine than Google ever was. It can produce answers with hyperlinks for further reading much better and more efficiently than any other option. I no longer have to read through a bunch of seemingly random Google search results, hoping that my specific question is addressed.
Nonethless, try looking more left and right. There are really good opensource solutions which can surfe the web for you like stormai, Or you can use anythingllm and give it all your local files.
You can write your tips and tricks, start commands, upgrade procedures etc. in Markdown and reference it through your local LLM.
I like it for coding specifically for languages or things i write seldomly (i'm not coding every day but did for 15 years).
Nonetheless, googles internal code review tool is already suggesting things which are getting accepted by more than 50%. Thats a lot and will only get better. GitHub with Copilot will also just get better every day too. They probably struggled (as the whole industry) with actually getting used to having ML stuff in our ecosystem. Its still relativly new.
That being said, I often come in a situation that its not the code that is bad but the solution. Where I need to tell the LLM that it's not logical to do it in a certain way leveraging on my own knowledge.
As it happens, I've just been using it this weekend to write code for me. Two things:
(1) it's not business critical code, it's a side project that I want to get done but otherwise wouldn't have energy for. Especially not in this heatwave in a century old building that has no air-con.
(2) my experiments are a weird mix of ChatGPT wildly messing things up and it managing to get basically everything done to an acceptable result (not quality of code, quality of output). Sometimes I have the same experience as you, that it's just aggravating in its non-comprehension, sometimes it's magical.
I don't know if it can be magical more often if I was "better at prompting". But I do know that I also get aggravated (less often) by other humans not understanding me, and those are much harder to roll back to a previous point in the conversation, edit the prompt, and have them try again :P
As for code quality… well, sure. Stuff I'm asking ChatGPT for is python and JavaScript, and I'm an iOS dev. I can't tell when it's doing something non-idiomatic, or using an obsolete library or archaic pattern in those languages.
Most non-tech people think AI is no different to “algorithms” and is just another IT buzzword that means “computer people click a few buttons and it does all the work for them”.
AI is basically ml at this point. And it did already A LOT.
Whisper, great jump in quality for speech to text, segment anything, AlphaFold 2, all the research paper Nvidia publishes regarding character movement, AI Raytracing, Nerfs, all the medical research regarding radio imagin, advances in fusion reactors...
We have never been so close to a basic AI/AGI / modern robots. we have instructGPT which allows for understanding 'steps' easier and more stable than anything we developed before in multip languages.
ChatGPT and LLM advances are great and helpful.
Image generation is already poping up in normal life.
AI is not wildly overhyped at this point. We are in the middle of implementation after the first LLM breakthrough and a LOT more money is funnelt into AI/ML research now as it was 10 years ago.
The future is, at least for now, really interesting and there has not been any sign of a wall we are hitting.
Even the missing GPT-5 might feel like a slight wall, but we just got GPT4-o mini which makes all of the LLM greatness a LOT cheaper and a lot easier to use.
We switched from text parsing and avg bad results to just using llama3 (with a little bit of saveguarding) and its a lot better.
Really? where ?
In my company we even have a LoRa for a specific company style of images (icons and similiar)
really curious about this. do you have an example link by anychance
Its a german news site for it people.
But i have seen ai image already on the street, unfortunate i was in public transport and not able to take a picture fast enough when i saw it and i currently work most of the time from home.
Your account is 3 days old, you haven’t read anything over and over again.
I just create a new account to get away from fomo and add a little bit more effort to commenting.
But as you can see, it doesn't work very well
It is still unclear to me whether this is a deficiency in all ChatGPT models or this is just one of the many.
It's both, the level of criticism is also overwhelming and delusional.
I hear things like, AI will never be as smart as me!, It will never take over my job, Look it got this and this wrong, No chance it can ever be as smart as me! It's always a comparison to their own abilities so I think a lot of it is just an attempt to stay relevant in a world where technology is about to replace all of us.
I think by now everyone knows the limitations of current LLMs. The result of this "in depth" analysis is not only expected but very obvious. Nobody is surprised and none of this is suppressed because it's so obvious.
What's delusional is the amount of criticism around the performance. The denial that AI is more and more matching human intelligence. Of course it's not there yet, but LLMs made a giant leap and bridged a huge gap.
AI is no match for humanity now, that much is obvious. I think the delusion lies in the fact that many people are trying to deny the trendline... the obvious future and trajectory of what progress has been pointing too. Milestones are getting surpassed at a frightening pace. AI is now running circles around the turing test and I can now ride a car with no driver and it's normal in SF.
We are here bitching about the fact that AI shortens a text rather then summarizing it without remarking at the fact that you can actually bitch to the AI directly about this fact and demand the AI to stop shortening the text and start summarizing it.
It is true that LLMs appear to be an exciting technology, but it's also delusional to assume they're following a positive trendline. Performance between GPT-3.5 and GPT-4 was like a 25% improvement that took 2500% more resources to train. It's clear that there's diminishing returns to bigger and bigger models, and the trend we've been seeing in the industry is actually smaller models trained for longer periods in an effort to bring down inference costs while maintaining current performance.
More intelligent models may require new techniques and technologies that we don't have yet. I'm sure it will get better but the path to improving isn't as obvious as you're making it sound. Making comparisons to Moore's law is also disengenuous because we're actually running into physical limits on how dense we can make chips due to the size of atoms themselves, so past trends for technological development may not continue to bear out.
You are doing exactly as I said. Focusing on obvious criticisms on the current state of the art. We know the obvious pitfalls of LLMs. It's completely obvious nowadays.
I also never pointed to an obvious path forward. I pointed to the an obvious trend that indicates whether you like it or not, we will move forward.
>It is true that LLMs appear to be an exciting technology, but it's also delusional to assume they're following a positive trendline. Performance between GPT-3.5 and GPT-4 was like a 25% improvement that took 2500% more resources to train. It's clear that there's diminishing returns to bigger and bigger models, and the trend we've been seeing in the industry is actually smaller models trained for longer periods in an effort to bring down inference costs while maintaining current performance.
AI is following a trendline. I never said specifically LLMs are exactly on this trendline. LLMs are only a part of this trendline along with other technologies and part of the overall progress deep learning is making. It is extremely likely that there will be an AI that will solve all of current issues with LLMs in the coming decades. Whether that AI is some version of an LLM remains to be seen.
>Making comparisons to Moore's law is also disengenuous because we're actually running into physical limits on how dense we can make chips due to the size of atoms themselves, so past trends for technological development may not continue to bear out.
I never made a comparison to moores law... are you replying to me or someone else? AI with the performance of a human brain is highly realizable despite physical limits because intelligence to the level of the human brain already EXISTS. The existence of humans themselves is testament to the possibility it can be done and is not a fundamental limit.
I think the tide is turning for sure. Metaverse scam unraveled pretty fast after building for months.
Perhaps everywhere else. On HN mostly what I see are these fairly shallow dismissals TBH. It is a natural reaction when your own livelihood is affected, old as tech itself to be sure [0]. Still, it's getting tiresome. A new technological revolution is unfolding, and the people best positioned to lead and keep it in check are largely balking.