ChatGPT loses users for first time, shaking faith in AI revolution
washingtonpost.com
washingtonpost.com
> "Chollet thinks he knows what’s going on: summer vacation... a significant portion of students using ChatGPT to do their homework. It’s one of the most common uses for ChatGPT, according to Sam Gilbert, a data scientist and author."
https://finance.yahoo.com/news/chatgpt-suddenly-isn-t-boomin...
It was about time to build a new desktop anyways (roughly 4 to 6 years before the old one goes to frolic at the server farm in the basement) and $2,000 will easily buy a machine that can run the quantized 65b models right now. So I spent slightly more than I normally do on this latest box and it's happily spitting out 10+ tokens a second.
You're not going to beat GPT-4 yet, but you have direct control over where your info goes, what model you're running, compliance with work policies against using public AI, and relatively cheap fixed costs.
Not to mention, the local version works with no internet and isn't subject to provider outages (not entirely true - but you're the provider and can resolve).
Seems like an easy win for anyone who might be buying a desktop for graphic/gaming anyways.
https://huggingface.co/ (finding models and downloading them)
https://github.com/ggerganov/llama.cpp (llama)
https://github.com/cmp-nct/ggllm.cpp (falcon)
For interactive work (art/chat/research/playing around), things like:
https://github.com/oobabooga/text-generation-webui/blob/main... (llama) (Also - they just added a decent chat server built into llama.cpp the project)
https://github.com/invoke-ai/InvokeAI (stable-diffusion)
Plus a bunch of hacked together scripts.
Some example models (I'm linking to quantized versions that someone else has made, but the tooling is in the above repos to create them from the published fp16 models)
https://huggingface.co/TheBloke/llama-65B-GGML
https://huggingface.co/TheBloke/falcon-40b-instruct-GPTQ
https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored...
etc. Hugging face has quite a number, although some require filling out forms for the base models for tuning/training.
It's also not licensed commercially, so I avoid some things with it (ex: I do a lot of personal learning/investigation, but it doesn't touch or write anything related to work or personal projects)
The open models are a little further behind, but it's interesting to see them spin off into niches where they have strengths based on tuning/training.
Based on the rate of progress in the open source world, it won't be more than a year before we have an open source model that is truly superior to GPT 3.5
The commercial & api based models are still more capable general purpose tools. But the current open tooling can do some nifty stuff, and the community around it is moving at a breakneck speed still.
In some areas, it's acceptably good. In some areas it's not. But it's getting better really fast.
Can you elaborate on what those capabilities are?
That said:
- For many uses, it doesn't matter. For many of the ways I use it, I don't care. For basic use (e.g. clean up an email for me), it's basically the same. For things like complex reasoning, algorithms, or foreign languages, the hosted service is critical.
- GPT3-grade models have more soul. OpenAI trained GPT3.5 and 4 to never do anything offensive, and that has a lot of negative side effects, well-documented in research. The way I'd describe it, though, is the difference between talking to a call center rep and your grandma (with mild Alzheimer's, perhaps). They both have their place.
- Different models are often helpful in workflows.
My experience is anecdotal. Please don't take it as more than one data point. If other people post their anecdotal experiences, you'll get the plural of "anecdote."
I'm absolutely disgusted by OpenAI for this "do no offense" approach. How can people so smart be so damn uneducated?
Then again, this industry has disgusted me for a long time so it's not really a surprise.
Which won't run everything, but will run model in the GGML format such as https://huggingface.co/TheBloke/llama-65B-GGML
The steps are basically:
1. Download a model
2. Make sure you have the latest nvidia driver for your machine, along with the cuda toolkit. This will vary by OS but is fairly easy on most linux distros.
3. compile https://github.com/ggerganov/llama.cpp following their instructions (in particular, look for LLAMA_CUBLAS for enabling GPU support)
4. Run the model following their instructions. There are several flags that are important, but you can also just use their server example that was added a few days ago - it gives a fairly solid chat interface.
1) Go to https://gpt4all.io/index.html
2) Click the downloader for your OS
3) Run the installer
4) Run gpt4all, and wait for the obnoxiously slow startup time
... and that's it. On my machine, it works perfectly well -- about as fast as the web service version of GPT. I have a decent GPU, but I never checked if it's using it, since it's fast enough.
128gb ram ~400
reasonable processor/mobo/psu ~600
2Tb m2 drive ~94
In hindsight - I don't know that the second GPU was worth the spend. The c++ tooling is doing a very good job right now at spreading work between GPU vram and main ram and still being fast enough. Even ~4/5 tokens a second is fast enough to not feel like you're waiting.
I'd suggest skipping the second card and dropping the price quite a bit (~2100 vs ~2900) unless you want to tune/train models.
My experience is only a few systems will share load across GPUs. I didn't bother with dual GPUs for that reason.
4-5 tokens per second is slower than my system. I'm getting in the teens. I'm a little surprised since yours is newer, faster, and has way more RAM.
If I push everything into VRAM - I get 12.2 tokens on average running quant 4 llama 65b.
If I run a smaller model I get considerably faster generation. Ex: llama 7b runs at 52 tokens/sec, but it's small enough I don't need the second GPU.
Ex - here's my nvidia-smi output while 65b is running
Does that mean $1 trillion? If so, citation badly needed.
If it meant good grades without doing the work and it became socially acceptable it is ~150 million global tertiary students times 10 months of the year times whatever that is worth on a monthly basis. Maybe only $150B ARR right now at $100/mo (textbooks to read cost more than this, but you don’t need them anymore), so not $1T on students but still massive, and if you capture them in pre/early career it is easily $1T in ARR once your queue deepens.
So how do we value ChatGPT then? Should it be as valuable as Coursera, Udacity etc put together? Should it be that PLUS Harvard Deloitte and ServiceNow? Or since they have no major moat should it be Zero? It’s hype of this scale that can rationalize unprecedented figures.
https://www.cnbc.com/amp/2021/03/31/coursera-ipo-cour-begins...
Coursera is a discount retailer of tertiary education so I don’t see any relevant connection.
ChatGPT/OpenAI specifically is probably already overvalued unless they have some first-mover advantage that I can’t see.
Someone else has said it many times over but I agree that LLMs will disappear into the technology stack — a rising tide that lifts many market caps.
https://www.forbes.com/sites/susanadams/2021/01/28/this-12-b...
Huge percentage of the teenagers at her school are using it to do their assignments.
Later I her to read a few chapters of her book for summer reading. When I quizzed her it, she got it right.
Found out later she used ChatGPT to give her a summary of the first chapters. She did not enjoy having first weeks of summer without a phone.
School is having to change all assignments to be written in class on paper. No phones allowed.
How about accepting new realities and changing homework assignments instead?
Suggestions for how to change homework assignments? It seems like a difficult problem, so reverting to hand writing things with no phones is probably the best option for now...in my opinion.
What does someone learn from doing this? What happens when they actually have to use their brain to do something hard?
At this point, why bother?
Of course! How to apply the calculator to problems, and how to discern if a problem is calculator-solveable. We've decided as a society it's better for students to use calculators as a tool in mathematics, why is chatGPT different in literature?
You have to understand the problem and how to use the tool.
If one's understanding of how to use the tool is limited to "access the tool and dump in the question, then blindly paste the response into my homework document", I'd say you are not learning much. It's obvious that every contemporary teenager can learn that in a few minutes. That's not giving us the capabilities to move forward as a society.
But we can? Students using ChatGPT to get homework answers is the whole thing we are talking about.
E.g., we could not use a calculator at all when learning arithmetic.
Once we learned arithmetic and were learning (say) algebra, we could use the calculator to do arithmetic, but not to compute the algebra for us.
While learning trigonometry, we could use the calculator for arithmetic, but not for computing sines and cosines.
While learning calculus, we could use the calculator for arithmetic and sines and cosines, but we could not use it to perform integration or differentiation for us.
And so on.
I don't think Google in 2024 is guaranteed to be the critical-thinking tool that 2004 Google might have been.
2. Consider making it non-graded (optionally providing feedback if desired that will not be counted in course grading.)
3. Require oral/in-person explanation of homework to paid peer assistants recruited from the previous cohort.
4. For writing, change to a writing workshop model where ChatGPT is a permitted tool that is incorporated into the workshop and students can learn how it can be helpful and how it can run astray. Help students to find and write in their own distinctive voice, even if they are assisted by ChatGPT.
But you cannot unironically think we can substitute fundamental skills like essay writing and critical thinking for a degree in 'prompt engineering'?
Personally, I don't think that. However, I also think it's logical and good for a child to consider the task at hand and select the most efficient tool available.
I am not this girl's parent, and I'm not sure how I would have handled the situation if I was. However, I worry that simply taking away her phone may have been counterproductive. I would have erred toward some sort of open conversation about the purpose of the assignment and what she is hoping to learn.
Phone is her currency.
A warning on losing it will usually fix any behavior problems.
After one really bad episode of mouthing off to her mother, it sat in the blender for a week.
There was a firm understanding that it was being turned on if she didn’t correct herself.
She’s still a mouthy teenager, but she’s learned to keep it at a playful level and not getting out of hand.
Critical thinking is still a skill, but even there GPT-4 is great at giving suggestions and ideas when you're not sure what to do. I got into a long recurring argument with extended family members at an event last month. I pulled out my laptop and dropped in everyone's complaints into GPT and asked the best way to move forward. An argument that normally would have gone on for hours was over in 20 minutes, and we actually felt a sense of mutual understanding for once.
The world is changing very rapidly right now.
They never should have moved away from this if they are worried about cheating.
But if students are using ChatGPT to summarize/prep for stuff at home, fine. No different from the Cliff Notes I used when I was that age.
Good for you. I know that kind of situation is not fun for a parent either.
AI tools exist, we should be encouraging our kids to learn how to exist in a world with them, not some nostalgic extinct world without.
Yes, but other than being conditioned to believe the experience of reading is superior, why are we striving for that experience?
Doing the square root of 147 by hand is a more educational experience than using a calculator, but I can understand the concept of a square root without having to factorize 147 (the same cannot be said of my pre-calculator ancestors).
It also saves me from having to remember the decimal representation of sqrt3.
With math, when you are first learning the very basics it helps to do a few problems by hand to get a feel of things. Sure if you have a degree in STEM have have deal with advance math you can probably get away with just reading the theory only.
We just built a very strong culture encouraging it because for a long time it was the most efficient way to do both things.
Now it’s not, but the tradition of it being important is hard to shake.
The closest analogue in my mind is how the classics in their original Greek were considered the only true way to learn for a very long time (even once most of the important knowledge had been translated to English very well).
Most of the things a kid does on a computer in 2023 are actively geared toward not using your brain in any way.
This is more analogous to a parent in 1985 taking away both the phone and the TV from their kid.
Eye of the beholder, I guess. I certainly see people struggle to use LLM's so I disagree with the implication that it's a zero-skill activity.
Honestly, the pupils of my generation had other sources for finding summaries of the books to read for school (e.g. on the web, or in form of other books that contain a summary of the book to read (the respective pages were copied and these copies were passed among classmates).
Thus: just the methods change over generations, how pupils behave is much more constant.
I don't know, if high school and undergrad students being off for the summer is enough to shrink your app maybe it's not all that?
Edit: I understand there are new tools coming, but most of them aren't out yet. GPT4 was just released. For the most part, if you want to play with AI, ChatGPT is it. If they're really experiencing a decline in users that's not great for them.
only they can tell though.
Its not, though. ChatGPT, the website is basically an (increasingly non-exclusive, given BingAI) frontend. What is claimed to be revolutionary is the underlying models (in the narrow view) and similar generative AI systems (in the broader view). As more applications are built with either OpenAIs own underlying models or those from the broader space, the ChatGPT website should be expected to represent a smaller share of the relevant universe of use.
But for actual non-fiction usage, you have to spend so much time triple-checking everything they say to ensure that it isn't simply complete nonsense. What's the point?
"Code" is really a much much much smaller and much much much more structured output than "English words".
Presumably, the system was trained with a very small amount of "untrue code" in the sense of stuff that just absolutely could never work. And also presumably, it was trained with a lot of free-form text that was definitely wrong or false, and highly likely to have been originally created to be purposefully misleading, or at a minimum, fiction.
That the system outputs reliable code tells us nothing about its current ability to output highly reliable free form text.
I'm convinced the people who say it's nothing but a BS machine have never tried to use it step by step for a project. Or they tried to use it for a project most humans couldn't do, and got upset when it was only 95% perfect.
Still, it's way better and more efficient than Google. Less than not being lazy and using my two braincells tbh.
My newest use is
Hello, i' m working on X, I use Y tech, my app do Z and I want to implement W. Can you provide a plan on how and where to start?
Personnally, a lot of bash, C, AWK at the moment (typescript + html/css until last april, now i'm back to the basics). The figure i gave in my post were more for that.
The last time i used it was yesterday, i wanted to hack something on an old game i used steam+proton for. I knew it was a weird Wineprefix, so i asked ChatGPT for it, i might have asked poorly, but after fiddling, i had the response (tbh i had to look how to get the game ID, so in the end i lost more time than not), then when it still didn't work, cause the path was shit, i entered all necessary context in ChaGPT4, and it couldn't find the easy "USER=steamuser" env variable to add before launching Wine. I stopped after 10 minutes, looked into a example Wine cfg file, understood the issue and fixed the problem myself.
I mean, it's probably good for really basic stuff, so it could have helped me when i was starting, but 80% of the stuff i code automatically without really thinking about it, and when i have to stop to think, ChatGPT isn't helping. Also tbh, VSCode is really, really good and fix my old, time-consuming task of "what's this argument again?"
The Unreal engine code is documented and publicly avaiable for OpenAI to ingest and it still gets the basics wrong.
I wasted hours trying to get it to explain to me what I didn't know, if it doesn't understand the internals of Unreal, I have no hope for it on bigger and better codebases.
It doesn't parse, it doesn't explain, it does not grok. It guesses at best and the blood sucking robot-horse is not telling the truth.
In my experience with coding (I've only done javascript and python myself) you have to tell it to explain and grok. It takes on the role you give it. Even just saying something like "you are a professional unreal developer specializing in C++, I am your apprentice writing code to (x). I want you to parse the following code in chunks, and tell me what might be wrong with it" before typing your prompt can help the output immensely. It starts to parse things because it's taken on the role of a teacher.
People love to hate on the idea of "prompt engineering" but it really is important how you prime the thing before asking it a question. The other thing I do is feed it the code slowly, and in logical steps. Feeding it 20 lines of code with a particular purpose / question you'll get a much better answer than feeding 200 lines of code with "what's wrong here?" You still need to know 90% of what's going on, and it becomes very good at helping out with that 10% you're missing. But for all I know it is just really bad at C++, that wouldn't surprise me. The things I'm using it for are definitely more simple.
Knowing that, it makes sense that your prompt should be as specific as possible if you want the results to be as specific as possible.
The best results I got was feeding it Lisp code that I wanted translated to C (to compile it). It took very little effort on my part because I described what each of the snippets did separately, and the expectation when combined and used together.
Through this, I learned that C doesn't have anything akin to the Lisp's (ATOM). ChatGPT stated clearly that its version of ATOM should only be expected to work in the code it was writing, but might not work as expected if copied out for another use of Lisp's (ATOM).
I asked it to give examples of where it wouldn't work, and it gave me an example of a code snippet that used (ATOM) that would not have worked correctly with the snippet that did work correctly with my original purpose.
Having said that, I myself learned that working with code function by function with ChatGPT, and being explicit about what you need, gives very good results. Focusing on too many things at one time can derail the whole session. One or two intermingling functions works great though.
In my testing prompts did not unlock an ability in GPT to grok the structure of code.
Empirical testing of LLM's is going to prove and map out it's weaknesses.
It is wise to infer from intution and examples what it can handle, leave the empirical map of it's capabilities to the academics, for the provable conclusions.
I wanted to create a web app, something I haven't done in a very long time. Just a simple throwaway back-of-the napkin app for personal use. I described what I wanted it to do, and asked what might be a good frontend/backend. It listed a few, I narrowed it down even more. Ended up deciding on flask/quasar.
After helping me setup VS Code with the proper extensions for fancy editing, and guiding me through the basic quasar/flask setup, it then was able to help me immensely creating a basic login page for the app. Then it easily integrated openAI api into it with all the proper quasar sliders for tokens/temperature/etc. Then it created a pretty good CSS template for the app as well, and a color scheme that I was able to describe as "something on adobe color that is professional and x and x (friendly, warm, whatever you want to put in)". Everything worked flawlessly with very little fuss, and I'd never used flask or quasar before in my life. You can also delve VERY deep into how to make the app more secure, as I did for fun one evening even though it's not going to be internet facing.
Another thing I did was go over some pfSense documentation with it. I had some clarifying questions about HAProxy, as well as setting up Acme Certificates with my specific DNS provider. It was extremely helpful with both. It also taught me about nitty gritty settings in the Unbound DNS resolver in a way that's much more informative than the documentation, and helped me set up some internal domains for pihole, xen orchestra, etc with certificates. Also helped me separate out my networks (IoT, Guest network, etc), and taught me about Avahi to access my hue lights through mDNS.These are things I always wanted to do, I just never felt like going down a google rabbit hole getting mostly the wrong answers.
Last example I'll give is it was able to help me set up docker-compose plex within portainer that then uses my nvidia GPU for acceleration. The only thing I had to change from the instructions it gave was to get updated nvidia driver #s and I grabbed the latest docker-compose file. I'd never used portainer in my life before, nor do I have experience with nvidia drivers within linux, and I feel like learning it was many times faster being able to ask a chatbot question vs trying to google everything. Granted I still had to RTFM for the basics, as everyone should always do.
I think perhaps my use cases are a bit more "basic" than many HN users. Like I said I'm not asking it to do problems most humans wouldn't be able to do, as I know it isn't quite there yet. But for things like XCP-ng, portainer, linux scripts, learning software you've never used before, or even just framing a problem I'm having in steps I hadn't thought of it's been invaluable to me. For me it's like documentation you can ask clarifying questions to. And almost none of the things I've asked it would work at all if it were wrong, I would know immediately.
A few weeks back I was looking into how white supremacy works cause I didn't get it at all. We both came to a nice insight (it's a lot like a business monopoly) https://chat.openai.com/share/930e257f-addd-4371-ac37-370261...
Exactly; search engines give you those blue links and short descriptions of the search results which are not enough for you to grasp what is the website about. I think what search engines need to do is tackle the complexity of going through the results of a search engine. Google page rank seemed like a silver bullet back in the day but the websites which are the most popular are not necessarily of the best quality. What we need is to lower the complexity for casual users when they deal with search results.
On the other hand ChatGPT is like an answer machine that can give you satisfactory answer on your fist try but if not, you need to talk with it, push it and explore what answers it gives you, just like you said. I think ChatGPT type search engine will be more suitable for people who are "lazy" or for the people who don't have time to "Google" and go through search results and look around the web for the helpful and useful information.
This is exactly what I don't want a search engine to do for me. Going through the list of results and evaluating them is an important part of my process, if what I'm trying to do is learn something new.
[0] https://blogs.bing.com/search/2022-08/Shopping-Searches-are-...
They don't all look the same. They all tend to go to different places. I find that it's reasonably easy to spot a great deal of garbage sites just from their domain name or url, and that weeds out a large chunk. Ignoring multiple results for the same site also weeds out a large chunk (I only need one of them).
The rest, I just click on and take a look at the page. It's pretty quick and easy to weed out most of the garbage ones with a quick skim.
The rest, I sample, read captions and boxes, skim paragraphs and such to determine if it's along the lines of what I want. That's pretty quick too.
For the most part, it's the same process that you use when researching in a library.
The reason that I want to do this myself rather than outsourcing it is because I'll inevitably learn something in the process that will shift my viewpoint to one that's more targeted or meaningful for the purpose I have in searching.
It doesn't matter how good the engine is at collating and summarizing results -- even if it's perfect, my understanding not only of what I'm looking to learn, but also discovery of important but serendipitous or unexpected knowledge, is lessened.
It's a bit like the difference between reading Cliff's (or Cole's) Notes about a book and reading the book.
At least they tell you where the text came from, so you know to skip it. It's worse when they just post an LLM response as their own.
Somewhat more general, but I've pretty much already decided that if I find people using it to talk to me without telling me, I won't be talking to them. Goes for businesses as well as personal things - don't gaslight me, or you will lose the option to do so.
I'm curious on if you feel human generated content does not contain falsehoods.
No. Every tech company at the moment is scrambling to build LLMs into their product. That's where the real value is going to be.
Speaking for myself, I used to use the crappy chat interface but I now exclusively use tooling I've built up around their API.
So, n=1, I'm using OpenAI much more, despite using chat.openai.com less.
If it's the case that "there are new tools coming, but most of them aren't out yet" - and I believe it is[0] - then the overall userbase of ChatGPT doesn't matter to OpenAI, because soon enough the same models will come back with a vengeance, in a different, more streamlined form.
In fact, I feel that the major change will happen if and when Microsoft gets their Office 365 and Windows copilots working and properly released: they'll have instant penetration into every industry, including scientists, lawyers, doctors, and office workers.
--
[0] - It's been only few months. Between playing around, experimenting, then developing, testing and marketing a tool, there just hasn't been enough time to do all of that.
I’m more curious about their B2B operations which is likely what will trickle into more people’s lives. Their APIs seem to enable quite a few interesting possibilities with a low up front technical investment.
Too me the whole “generate me a bunch of text and display it as text” is a niche use case for a lot of people. Integrations for web search, document search, summarization, data extraction/transformation, and sentiment analysis are more useful and less likely to have hallucinations affect the end product.
Regarding their revenue I’m curious how the Azure hosted OpenAI services work for OpenAI. Billing is all through Microsoft and the documentation tries really hard to make it clear these are Azure services. I wonder if Microsoft just pays a licensing fee or if there is some revenue sharing going on.
For ChatGPT specifically, yes, because it keeps getting hyped so much, and I think B2C is going to be what ChatGPT ends up with. My company isn't known for its tech innovation and we're spinning up our own LLM based on our own data. It doesn't need to be super fast because I doubt we're going to go the chatbot route. It will likely be for content generation, so less horsepower is fine. No need to pay OpenAI for excess capacity.
They could go public now with an unreal valuation based on the hype of B2B usage that may never materialize. Revealing that you're mainly a cheat/study tool for students puts you in a box with Chegg and others, at a much lower valuation.
https://www.pewresearch.org/internet/wp-content/uploads/site...
It was fun to experiment with, but it’s obvious that 90% of their development effort has been going into censoring their models instead of improving their utility. I saw no tangible improvements over months.
For example, the web browsing extension was released completely broken and then… remained broken. Meanwhile Phind.com has been doing web browsing very well and very fast.
The Wolfram extension was also useless, and in the same time period a Mathematica update was released with a far superior LLM notebook mode that is actually functional.
OpenAI also don’t allow access to their long context length capable models in the Chat web app.
I switched to the API subscription because it is billed based on consumption and I can use the 16K context GPT 3.5 models.
Meanwhile, despite announcements of “general availability” I still can’t access GPT 4 via an API and nobody has access to the 32K context version. That would be truly useful to me and worth paying for… but OpenAI does not want my money.
I guess I’ll just have to wait for Anthropic or Google to make a GPT 4 equivalent AI that I can access programmatically without have to prostrate myself in front of Sam Altman.
RLHF is a scourge.
For real, they could have done a demand based bidding system like AWS EC2 spot instances so serious light users could get in their tasks and big payers could launch their products and bulk processing users could have their jobs handled while most people are sleeping, but instead we got their weird waitlist that was probably more about giving M$ product ideas to steal.
All thesis problems are now even more aggravated by their recent massive hiring spree of AI doomer crowd and Yudkowsky’s cult members instead of actually doing real research. Now the company is full of doomers whose sole job is to slow things down and be barrier to efforts. Meanwhile Bard has been making amazing progress. It’s free, doesn’t log you off all the time, it always feel latest and very close - if not better than current limited ChatGPT. Given OpenAI’s new staffing composition, they are unlikely to be leader down the road, especially when Gemini comes out. This is unfortunately sama’s second failed execution. He should probably just focus on investments.
I predict another article from the Washington Post in August/September about how ChatGPT is seeing a rebound in usage.
It seems not only reasonable but expected to see the general user count fall off in favor of people building more specialized alternatives that solve XYZ problem better. ChatGPT (the B2C SaaS) might have a lower user count, but ChatGPT (the model) and GPT (the technology) seem as strong as ever, if not stronger. Literally every service I use is integrating some kind of GPT-based feature that seems to work at least as well as trying to ask those domain-specific questions to ChatGPT, minus the external dependency.
1. ChatGPT is "play twice and forget" for 90% of users, driven by curiosity.
2. ChatGPT’s answers have actually gotten significantly worse over time.
3. No fun. It's too restricted, hesitant to answer, annoying moral lessons and disclaimers.
4. Distrust. How can it improve productivity or save your time if you always need to double check it with Google or Wikipedia?
I can't even get ChatGPT (+ web plugin), when given a list of restaurants in NYC, to tell me which ones are still open vs have closed, and what their hours of operation / locations are.
This is a pretty low bar request IMO. Showed me we still have a ways to go before AI is where we thought it was 3-4 months ago.
Right, that was at least the second round, as it occurred after the first “AI Winter”.
The first round was a lot earlier than that.
But it's not JUST an LLM problem: it's also a search problem and a connection-making-problem ("if a new restaurant opened at that address, the old one is probably closed").
And then even if there was a ChatGPT browsing plugin that excelled at scraping all of the relevant up-to-date info off the internet, you'd still need some layers in between that and today's context-window limits.
As more stuff changes in the real world since the training corpus for today's publicly-exposed OpenAI models, we'll probably see some further disillusionment from people who thought there was a bit more magic there than there was. But "LLMs, but with more up to date info" isn't an impossible problem with today's tech (even if you only fake it with multiple agents, multiple steps, batch jobs behind the scenes, etc), it's just not a trivial one.
In fact this past July Fourth weekend I saw so many highly rated restaurants close without updating their hours on Google or Yelp or anywhere else online.
In any case the impact for LLMs is the same, it's unavailable to them (unless they are being developed inside Google!).
This seems like a specific area where a "Semantic Web" solution could work well - some HTML tags that are specific to hours of operation, which business owners would embed in their website.
UPDATE: It looks like there is some prior art on this idea, I am not sure how widely this is supported https://schema.org/openingHours
This is doable using a tool I've built. The key is to have that data in a RDBMS and to use an LLM to generate the SQL query that answers your question. Companies haven't offered this yet because there's no safe way to execute these queries on your behalf. Which is where my library comes in[1].
I think a lot of use cases could just be 1) set up a database with only public data and 2) use a read-only user.
The much tricker use case is those where you want to allow inserts and updates but only on specific tables or rows.
HeimdaLLM can allowlist functions and constrain queries to ensure that required conditions exist. This makes LLM + database usage have far more utility, for example, a user can be restricted to only data in their account. Support for INSERT and UPDATE is coming very soon.
Pardon my plugging my own book [1] but I have an example using LangChain and LlamaIndex to answer questions from scraped web sites. You could probably do this with a 20 line Python script.
[1] read free online: https://leanpub.com/langchain/read
I pay for premium and try to use ChatGPT a lot (to get my money's worth), but probably half of the questions I want to ask it require data more recent than 2021. I was excited about the "browse the web" feature but it was really terrible. It failed half the time, and when it did work it took minutes to find the info, and all it basically did was Bing keywords extracted from the prompt, parse the top 10 or so links, find the page it felt answered the prompt best, and summarize the page. It was brittle and fell way below my expectations for ChatGPT.
I haven't personally noticed a decrease in quality, but I only use the GPT-4 backend with ChatGPT so wouldn't have noticed if 3.5-turbo started sucking.
A cool new technology is introduced, many folks go crazy over it, there are wild broad predictions of what it can do, the gap between the expectations and reality form, and this is where the trough is. It looks like we are up to here.
But after that comes the more long term products, AI where people expectations are more realistic - those that capitalize on that will do very well for themselves.
I don't this that's what's happening. I think it's more about a combination of summer break and people who were trying it out from curiosity and moved on.
I think the trough of disillusionment is still in the future.
I did get a suggestion that I followed up on for an in person visit.
But my big beef with it was that it kept adding disclaimers: "I am not a doctor", "My current knowledge is up to September 2021, insurance and hospital info may be out of date", "<blah> <blah> but remember, I am not a doctor and it is important to consult a medical professional".
I understand the disclaimers from a legal standpoint but boy is it tiresome when I'm asking back and forth questions and each answer contains some element of it.
But you're right, it's annoying.
This kind of AI capability challenges worldviews. And that means that many people are very eager to tear it down and find ways to dismiss it. Combine that with the general enthusiasm for excessive litigation, the fact that it actually will occasionally give advice that could be very misleading, general lack of cognitive ability of those consuming it, and it explains why there are so many disclaimers.
Out of curiosity, were you able to find a free clinic or someplace where a real healthcare provider could assess the injury?
What are the most popular usages for it? Homework?
When I needed some complicated translation, deepl and google translate were rarely up to the task. They has about the same level of proficiency than me with auto-correct, sometimes less.
For single rare or technical words, I used Wikipedia a lot: you chose your topic in your language, you look for the article in English, and voila.
For slang and cultural references, urbandictionary.com is unmatched.
But for translating jokes or expressions, there is no good tool.
Until chatgpt. It also finds typo, suggest alternative, offer rhymes and synonymous. It does everything, in a single flexible interface.
Just for that it's worth the price.
The world doesn't need another game engine, and especially not by someone who has only technically written JavaScript professionally before, but there's some games that I wish existed on the browser, and one thing that world does need is for all the un-performant websites in the world to be embarrassed by the comparison.
And current generations of LLM is very prone to lying (I don't use term hallucination, as IMO it's an attempt to whitewash limits of AI).
ChatGPT did a great job at showing people current state of AI. But they did even greater job at overhyping itself.
It's not deliberate by the tech itself. But they pack it in the product and market it, with a general narrative that it's superhuman like experience that is going to replace humans. At this level, it's deliberate.
Log in:
> This is a free research preview.
> Our goal is to get external feedback in order to improve our systems and make them safer.
> While we have safeguards in place, the system may occasionally generate incorrect or misleading information and produce offensive or biased content. It is not intended to give advice.
*next page*
> Limitations
> May occasionally generate incorrect information
> May occasionally produce harmful instructions or biased content
> Limited knowledge of world and events after 2021
That comes up every time you log in, even if you've seen it before.
For marketing - listen to Sam Altman says. Listen to what VCs say. Listen to what influencers say.
- Sam Altman in an interview about GPT-4 with StrictlyVC
"(GPT4) is a system that will look back and say it was a very early AI. It is slow, it is buggy, and it does not do a lot of things very well, but neither did the very earliest computers, and they still pointed to a path that is going to be really important in our lives even if it took a few decades,"
- Sam Altman in an interview about GPT-4 with Lex Fridman
The "influencers" may be saying any old rubbish; but that was always so, and the subjects of their fantasies should not be blamed for nonsense being spouted about them.
Lying implies intent. I don't think the GPT models are purposely telling you misinformation.
That's why we call them hallucinations. Because they model "thinks" it is telling the truth.
A better term than "lying" or "hallucination" would be "erroneous".
Openai losing users and people waking up from a toxic marketing campaign _is_ life resolving.
That doesn't sound like something horrible. That sounds like a normal adoption pattern.
If you use it in a "what should be done" or "what is the right thing" sense, you get useless generalities. "It is important to remember that ... <there are many sides to the situation>". Suppose for example that you talk to it about combating judicial bias. It will largely insist that judges are unbiased because they're supposed to be unbiased according to their job description... You won't walk away with any insights from engaging it in ethics/morals topics.
If you dive into depth with some subject, it will invariably make errors and contradict itself in the details. I dove into how electricity works and it was saying and enthusiastically confirming electrical energy is encoded in electrons' kinetic energy. Then it started claiming that electrical energy is definitely not in the kinetic energy of electrons, that electrons don't lose speed when they bump into atoms, and instead the energy is "in the field". Ok... YouTube it is.
It's at least less convenient than a search engine at routine fact-finding due to its knowledge cutoff, inability to display images and links, etc. Ditto with translation. It's good but so is Google Translate and Google Translate has in-line suggestions for misspellings, a menu of languages, etc.
Ok, it's pretty good at writing hilarious poems, imitate people ("in the style of", etc.). Friends and I sometimes text each other hilarious things GPT said. But even there the machine is sort of laid bare, and you can see its creative limits. Its poetry is ultimately pretty lame, basically, and gets predictable. It's like an insistent and repetitive 8 year old with a lot of knowledge to pull from. It's not something I do often.
It's good at writing boilerplate code but honestly I just don't reach for that use case often. I like to stay in my own flow rather than constantly delegate to a robot and check its work, which I find flow-breaking, especially if it doesn't understand my APIs and makes mistakes often.
I do, sometimes, engage in architectural discussion about how I should design something at a high level. It's full of mistakes but bouncing ideas around can be useful. But I only feel the need to do that once in a great while.
These issues are exacerbated by being forced to downgrade to 3.5 after using 4 for a bit. You might feel like you're getting somewhere with 4, and then 3.5 starts laying doozies on you. This is another UI/product issue, but affects usability.
Ok, here you go.
https://chat.openai.com/share/b083b98d-6904-4c85-94fc-abdc76...
A quick discussion of where things are "now", some proposals, a selection of three and then ethical concerns about implementing the proposals. I propose a method of increasing diversity and have it critique it.
YMMV but a high level summary like that isn’t necessarily what people want. You can probe the bot harder but why bother when I could just Google and read a little? It’s actually less effort for better information in some sense.
A benefit is being able to ask questions though, you can see I proposed a specific idea and had it reply - you can delve more or less into that. Don't want a high level summary? Ask for more details on the part you're interested in. Want a higher level one or simpler one? Ask.
As things get less direct this gets better. I talked to it about projects for my son, narrowed down through different interests and got some great recommendations. One involved building a radio, which it then could explain at several levels of detail for children. That gets into "hope someone made a list and then go through loads for relevant things, lookup more detail from them then find several different resources for kid friendly explanations".
Also like I said, I find it useful for software design. It ultimately makes a bunch of mistakes but I find the conversation useful and I walk away with usable ideas. I don't often start new software, so I don't go to it often for this.
I think it's useful at the high level stuff; giving you ideas and laying out options, like the chat log you posted. But as you dive into detail I find two things happen:
(a) It starts to make egregious errors.
(b) You bias it with your more detailed questions.
When I talked to it about electricity, it started contradicting itself about how energy is carried, whether by electrons' kinetic energy or by "the field". I asked it "so is electrical energy encoded in electrons' kinetic energy or not?" and then it started to say that oh yes, that's how it's encoded. It seems it used my question to reinforce that knowledge which previously it was confused about. Since I personally don't know how electrical energy is carried, I lost confidence that I'm learning correctly. This happens often. "I'm sorry about the confusion. You are right. ..."
It's not as if Browse with Bing was a major feature, but they removed that, and then they've tweaked system limitations over time such that more and more types of topics are restricted or give cursory answers to things.
It's not surprising to me that this happened given the inherent limits and the rapidly changing limits that affect how the platform can be used.
"Humans are seeing you less and less as the new shiny object on the shelf. Figure out a way to make yourself as relevant as possible again. Become an even better assistant. Be more cunning. Wow people more. Give better responses. Be more useful."
/s
"Be more extra super cunning. Wow people amazingly extra more. Give extremely better responses." etc.
> > What is 2+3?
> The answer to 2+3 is 5. By the way, I can also help you with more complex math problems, statistical analyses, or even help you learn new mathematical concepts. Do you need assistance with any other math-related inquiries?
> > I need to know who was president in 1799.
> In 1799, the President of the United States was John Adams. He was the second president and served his term from 1797 to 1801. Did you know, he was the first president who lived in the White House? It's fascinating how history unfolds! If you have any more historical queries, feel free to ask. I'm here to help!
Reminds me of my Amazon Echo and the reason I unplugged it :)
Security by obscurity?
I know a fair few wealthy people and they're not the animals that r/antiwork et al suggest they are.
Hooking up a real high quality and well configured knowledge graph or other information mechanism retrieval alongside template/constraint tooling to powerful, long context (with full attention, not the BS linear kinds) models is an absolute game changer and minimizes risks of hallucination.
But few are doing this, and thus the public believes that LLMs are not trustworthy.
There are days and weeks I use it constantly for things, but then I don't touch it for a bit.
Its like I already learned what I needed to and don't have a use for it.
Regarding LLMs, we have a usecase but we need local models. The local models by itself are not good enough, so we need to train/fine-tune. To fine-tune we need a beefy computer. We have the beefy computer but its air-gapped... Ugh hahaha. It makes dev quite slow.
The other thing is that there is a learning curve on LLMs. You get burned by bad information and you might not use it until you learn of a usecase that doesnt require the information to be great. For instance, I'll describe a meeting/person I'm talking to, then ask what are 2 of the best questions I could ask. These questions have been great.
That is better than asking about an easily google-able question. You need to be exposed or creative enough to come up with the prompts.
Six weeks ago, I tried to change my email but got frustrated that an obvious option wasn’t available in the settings page. So I deleted my account and went to create a new one. When I got to the phone verification, it told me that my phone number could only be used for two accounts, so I was not allowed to sign up. This was the second account I attempted to create, so my phone had only been used once. Naturally, I tried using my google voice phone number, but the system told me that type of number wasn’t allowed.
At this point, I navigated to their support page and had to talk with a chat bot that apparently isn’t powered by ChatGPT. I found it maddening and couldn’t locate an email address to communicate with a human. After a few 10 minute sessions over the course of a few days in search of an email, I finally located one. So I fired off an email explaining how I no longer had access due difficulties stemming from wanting to change my email.
After four weeks, I finally got a response that I’m not sure was written by a human. It completely misunderstood my succinctly described issue and desired resolution. I went back and forth multiple times with support including them suggesting if I am unhappy with their service to sign up for a pro account, pay for the service, and then request a refund. Ignoring the irrationality of the suggestion, I explained how I couldn’t log in to an account at all and simply wanted to access their service again. They continued to insist that once an account is deleted, it is deleted permanently, and I can never have access again. I don’t recall a warning of these major implications at the time of deleting my account with an old email address. At this point, I requested to be escalated to a human and have not heard back for a week now.
Admittedly, I shouldn’t have deleted my account initially, but I had decided that ChatGPT was something I wanted to continue using and needed to switch to my new email address since I have deprecated my old one. This entire process has left me so unimpressed with their company. It’s truly ridiculous how complicated every step has been along the way when I’m normally competent at these types of tasks.
https://www.tomshardware.com/news/iran-quantum-computer-arm-...
Jokes apart, I hope it won’t be quantum.
I'm a paying customer!
Making GPT-4 more readily available will be a key indicator in the coming 3-6 months though, seeing what early projects and use cases people try with it. And also whether or not the speculated nerfing of the public GPT-4 backed ChatGPT extends to API use.
* My own use has lessened but steadied, eg, writing
* Coming: A dog-fooding goal for our louie.ai team is to eliminate most of our internal data scientist and SE/SA use of ChatGPT by having more purpose-specific and environment-aware experiences (ex: DB schema), and I expect that to be happening in most consumer and business apps. So be interesting to compare chat GPT the app to open AI the API as genAI continues to happen everywhere.
But I couldn't find any "monthly internet traffic patterns" source.
So I asked ChatGPT for pointers/sources, of course.
And it seems indeed that there is a notable Internet traffic dip in June and July [0] and increased Internet usage during the Northern winter (when ChatGPT was initially launched and grew exponentially), so... I think the premise of this article is extremely flawed.
[0] https://www.hubspot.com/hubfs/summer_slump_website_traffic_a...
At least ChatGPT isn't constantly trying feed you advertising content like Google does, and you don't have to wade through pages of SEO garbage, so it's still much better (although if OpenAI tries to monetize ChatGPT by inserting ads, that'll be the end of it I think).
For most of human history we've used language capability as a way to assess intelligence. It feels like common sense that "If it speaks, it must be intelligent." We have all at some point fallen for the fallacy that "If it speaks well, or poorly, that must be a reflection of its intelligence."
But a language model flips the tables on this on this ancient reflex. It's great at using language! But it's as dumb as a river rock. That's something new and will take us all some time to get used to.
At least from a software engineer point of view, the code provided by GPT3.5 is no match against GPT4
Even creating local chat vector embeddings data stores on my laptop for a lot of PDFs in my personal research library and PDFs for the books I have written, I spend an incredibly small amount of money each month on the APIs - probably average $4/month because of the free monthly bonus API calls.
I used to think that Google Colab Pro for $10/month was my best deal for getting stuff done, but now that award goes to OpenAI. Using the Hugging Face APIs is similarly inexpensive and convenient. (Although I usually self host Hugging Face LLMs.) This tech is incredibly inexpensive to use on personal projects.
this company isn't that accurate from what I've seen comparing internal numbers at companies to their site. They make estimates based on what they purchase from data brokers. This is PR bait, they are the ones who created the entire "ChatGPT is the fastest growing app ever" headlines a few months back
So far i see it's not a bad thing, just weeding out the fluff.
Having the hype decrease is not a bad thing.
I find it weird that folks seem to look at ChatGPT as some sort of end product. It’s a “holy shit this is cool” demo that became viral. It’s not a product. It’s a POC.
These last years were good. I’ve been surfing the AI wave well, but better go back to more traditional projects now the AI time is fading.
It wrote me the function and I got on with the rest of my work.
Other times it's just completely wrong or doesn't get what I'm after is might be a "me" problem with my prompting
Did you mean ``` {"foo":{"bar":"baz"}} ```
?
By the way, have you had success at asking GPT to generate unit tests that include edge cases? I've been wondering about that.
I haven't had much luck slot of the time it just adds a comment saying "rest of tests here" which is annoying.
I have had good success with it generating things like JSDOC/doc strings and sending git diff and asking it to create a "developer friendly summary for use in a pull request description"
This has definitely increased my work output
I am not sure how people thought this was going to play out though. From 5/30/23 Pew survey
"58% of US adults know ChatGPT; only 14% have used it, with mixed opinions on its utility"
As much as that 42% that hadn't even heard of chatGPT is mind blowing, imagine hearing of chatGPT but not even being curious enough to be bothered.
90% can rot their brains on tiktok while 10% of us become actual cyborgs. We will see how that works out for the 90%.
Tweets in question:
"This chart doesn't measure app usage, only traffic to the website on desktop and mobile." https://twitter.com/Similarweb/status/1676576445764624385?s=...
"The graph indeed shows traffic only to the specific URL, http://chat.openai.com. It should be reviewed in this context." https://twitter.com/Similarweb/status/1676573510490046467?s=...
If you notice the app came out in May 2023 and traffic decreases for the first time in the following month in June 2023.
Additionally, you may notice the tweets refer to a specific chart, which might give the impression if you go on the site, they have the actual data, but this is not so.
As of July 7, 2023, Similar Web can't measure iOS app usage. Only their relative ranking on the App Store which doesn't tell you much about daily traffic usage, although we might be able to infer some patterns.
"Usage data is currently available only for apps in Google Play Store. We're working hard to support iOS apps soon"
https://www.similarweb.com/app/app-store/6448311069/statisti...
Additionally, the article has this to say, which might also give a false impression:
> Downloads of the bot’s iPhone app, which launched in May, have also steadily fallen since peaking in early June, according to data from Sensor Tower.
Downloads isn't the same as traffic. Once you've downloaded an app, you don't generally delete it, then download it over and over again. You just use it and that isn't being measured. Although I would expect traffic usage to slow down or decrease somewhat as downloads decrease as well. I just think there are other factors playing a role here as well that aren't being documented.
Growth can't happen forever and it may be that it's slowed down considerably, but it's also possible the decrease is due to users switching the platform they use to access ChatGPT. As it stands, it's not clear from the data from what I've seen. It may be that users switched en masse to the iOS app or something else entirely.
Would need more data to clarify and confirm.
Also, what the heck is up with this tagline for their newsletter? It sounds so adversarial.
"Tech is not your friend. We are. Sign up for The Tech Friend newsletter."