Having the feature on a menu somewhere would be fine. The problem is the confluence of new features now becoming possible, and companies no longer building software for their users but as vehicles to push some agenda. Now we’re seeing this in action.
Part of the model of products like Adobe's Creative Suite [1] is that they are always adding new features -- and if you want people to keep renewing their subscription you want them to know about new features so they feel like they are getting more out of their product.
Trouble is using a product like that is like walking out of the Moscone Center and getting harassed by mentally ill people and addicts or like creating an account in Tumblr and getting five solicitations for pig butchering and NFT scams in DM in the first week -- you boot up the product, spend 20 seconds looking at the splash screen, then you have to clear five dialog boxes that you might not have time to deal with right now. Sometimes I open up a product because I have to do a task I have to do but don't really want to do and feeling a lot of stress and I just don't need to deal with any bullshit when I am under the gun.
I've seen Adobe trying gentler methods to point out new features in Lightroom, such as a filter that can automatically weed out photos where people have their eyes closed. It takes a lot of UX work to do that though.
Personally I'd like it a lot better if the nagging started after I finished a task, if I was feeling satisfied with the product and now relieved that the task is over that's a moment when I'd be receptive to learning more about the product.
[1] And also a lot of "free" software, it's not just money-grubbing, but the model of always rolling updates.
This is the fundamental problem and it has nothing to do with AI. Just look at the recent iOs 26 release. I am not convinced that any of the actual functional changes warranted a release or that they needed to be released at that point if a new release was needed. New software to justify new phones.
You get lots of features but performance takes a second seat. And sadly, I feel it works. I feel most would balk at paying a monthly subscription if only performance related improvements were made
Why do we need OS changes though? Well practically we don’t. But the platform owners all want to move new hardware so they need to shovel features in, which we could just completely ignore, except that they’ll abandon you to the wolves for security patches, which is about the only “new” thing we do need, if you’re not on the latest couple releases. And as for hardware, eventually you need new hardware and drivers only get created for current and future OS releases.
So the end result is we’re being led on a wild goose chase of trend-chasing shitty UI changes, adware, and performance-killing crap we don’t need, purely because we can’t run the old hardware forever, and even when we can keep the old hardware going, we can’t safely run old software for lack of patches.
And this is why the subscription model just doesn’t make sense for most businesses. I pay for a newspaper subscription because there is literally a brand new newspaper each day. A magazine subscription yields an entirely new set of articles every month. I pay for subscription access to data that is continuously updated. The subscription model makes sense for a product that is created anew on a regular basis. It doesn’t make sense for most software companies that are producing static software. What they are calling ‘subscriptions’ are really just rentals for their static products that get minimal surface changes to justify the ongoing rent charge. I’d much rather just pay a flat fee for the static software and upgrade it when I’m ready for the new features.
But now I’ve got several bugs (and I’m on last years flagship), liquid glass is ugly until you change a guy of settings, and I find myself accidentally triggering something (usually Siri) and being annoyed more.
This is especially irritating when, say, you set up a new phone and the app treats you as if you've never used it before.
Really my complaint is anything that covers up content; if instead of popping up a popover Firefox just took 75px above or below the page to show me something I’d complain about lot less — but if I had my way anything unwanted that covers unwanted content should bust down the whole c-suite to working in an Amazon warehouse. (I could trust those folks to deliver stuff with an e-bike but don’t want anybody with bad judgement like that driving a car or truck!)
All companies push an agenda all the time, and their agenda always is: market dominance, profitability, monopoly and rent extraction, rinse and repeat into other markets, power maximization for their owners and executives.
The freak stampede of all these tech giants to shove AI down everybody's throat just shows that they perceive the technology as having huge potential to advance the above agenda, for themselves, or for their competitors at their detriment.
2. AI could be the next technology revolution
3. If we get on the AI bandwagon now we're getting in on the ground floor
4. If we don't get on the AI bandwagon now we risk being left behind
5. Now that we've invested into AI we need to make sure we're seeing return on our investment
6. Our users don't seem to understand what AI could possibly do so we should remind them so that they use the feature
7. Our users aren't opting in to the features we're offering so we should opt them in automatically
Like any other 'big, unproven bet' everyone is rushing in. See also: 'stories' making their way into everything (Instagram, Facebook, Telegram, etc.), vertical short-form videos (TikTok, Reels, Shorts, etc). The difference here is that the companies put literally tens or hundreds of billions of dollars into it so, for many, if AI fails and the money is wasted it could be an existential threat for entire departments or companies. nvidia is such a huge percentage of the entire US economy that if the AI accelerator market collapses it's going to wipe out something like ten percent of GDP.
So yeah, I get why companies are doing this; it's an actual 'slippery slope' that they fell into where they don't see any way out but to keep going and hope that it works out for them somehow, for some reason.
Similar to how I read about a bar in the UK that has an intentional Faraday cage to encourage people to interact with people in the real world.
That's the core issue. No one wants to fail early or fail fast anymore. It's "lets stick to our guns and push this thing hard and far until it actually starts working for us."
Sometimes the time just isn't right for a particular technology. You put it out there, try for a little bit, and if it fails, it fails. Move on.
You don't keep investing in your failure while telling your users "You think you don't want this, but trust us, you actually do."
I think there are more mundane (and IMO realistic) explanations than assuming that this is some kind of weird power move by all of software. I have a hard time believing that Salesforce and Adobe want to advance an agenda other than selling product and giving their C-suite nice bonuses.
I think you can explain a lot of this as:
1. Executives (CEOs, CTOs, VPs, whatever) got convinced that AI is the new growth thing
2. AI costs a _lot_ of money relative to most product enhancements, so there's an inherent need to justify that expense.
3. All of the unwanted and pushy features are a way of creating metrics that justify the expense of AI for the C-suite.
4. It takes time for users to effectively say "We didn't want this," and in the meantime a whole host of engineers, engineering managers, and product managers have gotten promoted and/or better gigs because they could say "we added AI" to their product.
There's also a herd effect among competing products that tends to make these things go in waves.
If we didn't have pervasive telemetry, we also wouldn't have these obnoxious nudges; UX teams would get their feedback from QA testing and focus groups, and leave the end users in peace.
I'll bear that in mind the next time I'm getting a haircut. How do you think Bob's Barbers is going to achieve all of that?
some weeks if its slow he may struggle to make his rent for his apartment; he doesn't have time or capacity to engage in serious rent-seeking behavior.
but hair cut chains like Supercuts are absolutely engaging in shady behavior all the time, like games with how solons rent chairs or employing questionably legal trafficked workers.
and FYI turns out that Supercuts a wholly owned subsidiary of the Regis Corporation, who absolutely acquires other companies and plays all sorts of shady corporate games, including branching into other markets and monopoly efforts.
https://www.thebignewsletter.com/about
> The Problem: America is in a monopoly crisis. A monopoly is, at its core, a private government that sets the terms, services, and wages in a market, like how Mark Zuckerberg structures discourse in social networking. Every monopoly is a mini-dictatorship over a market. And today, there are monopolies everywhere. They are in big markets, like search engines, medicine, cable, and shipping. They are also in small ones, like mail sorting software and cheerleading. Over 75% of American industries are more consolidated today than they were decades ago.
> Unregulated monopolies cause a lot of problems. They raise prices, lower wages, and move money from rural areas to a few gilded cities. Dominant firms don’t focus on competing, they focus on corrupting our politics to protect their market power. Monopolies are also brittle, and tend to put all their eggs in one basket, which results in shortages. There is a reason everyone hates monopolies, and why we’ve hated them for hundreds of years.
https://blogs.cornell.edu/info2040/2021/09/17/graph-theory-o... (Food consolidation)
https://followthemoney.com/infographic-the-u-s-media-is-cont... (Media consolidation)
https://www.kearney.com/industry/energy/article/how-utilitie... (US electric utilities)
https://aglawjournal.wp.drake.edu/wp-content/uploads/sites/6... [pdf] (Agriculture consolidation)
https://www.visualcapitalist.com/interactive-major-tech-acqu... (Big Tech consolidation)
I think part of the Mozilla problem is that they are based in San Francisco which puts them in touch with people from Facebook and Google and OpenAI every frickin' day and they are just so seeped in the FOMO Dilemma [1] that they can't hear the objection to NFT and AI features that users, particularly Firefox users, hate. [2]
I'd really like to see Mozilla move anywhere but the bay area, whether that is Dublin or Denver. When you aren't hanging out with "big tech" people at lunch and after work and when you have to get in a frickin' airplane to meet with those people you might start to "think different" and get some empathy for users and produce a better product and be a viable business as opposed to another out-of-touch and unaccountable NGO.
[1] Clayton Christensen pointed out in The Innovator's Dilemma that companies like Kodak and Xerox die because they are focused on the needs of their current customers who could care less about the new shiny that can't satisfy their needs now but will be superior in say 15 years. Now we have The FOMO Dilemma which is best illustrated by Windows 8 which went in a bold direction (tabletization) that users were completely indifferent to: firms now introduce things that their existing customers hate because they read The Innovator's Dilemma and don't want to wind up like Xerox.
[2] we use Firefox because we hate that corporate garbage.
But if users really wanted agenda-free products and services, then those would win right? At least according to free market theory.
Not once in the history of tech “the free market” has succeeded in preventing big corps or investors with lots of money from doing something they want.
https://www.acquired.fm Acquired podcast does long (2 4-hour episodes on Google) episodes on various companies, mostly tech but recently Trader Joe's
I never used any of its collaboration features, just looked at them. I did use it as a friendly-for-non-geeks version of IRC for a group of people that lived in three separate cities as a virtual watch party for LOST. And for that, it was spectacular even if it was painfully slow on a netbook (so was everything else, but it was cheap and light and worked).
The thing is that a Google Mail early invitee could collaborate with everybody else via the pre-existing standard of SMTP email. They felt special because they got a new web UI, told their friends about it, generated hype, which then made the invites feel even more special, etc...
Google Wave had no existing standard to leverage, making it 100.00% useless if you couldn't invite EVERYBODY you needed to collaborate with. But you couldn't! You weren't allowed! They had to wait for an invite. Days? Weeks? Months? Years!? Who knows!
There was a snowball's chance in hell that this marketing approach could possibly work for a collaboration tool like Google Wave, but Google knew better. They knew better than every journalist that pointed this obvious flaw out. They knew better than every blog post, Slashdot commenter, etc...
It was one of the most spectacular failures caused by self-important hubris that I've ever seen in any industry.
Google Plus was 100% hubris. “If we build our version of Facebook, it course everyone will flock to it.”
I'm trying to remember all of the crap integrations with the likes of Youtube that were pushed. Just, screw that stuff. And quit trying to make yet another new messenger app!
Doing nothing while a competitor gains steam would've been hubris.
Maybe I’m wrong and internally they knew they had a major uphill battle, but I don’t think so. So many of the choices they made were needlessly user hostile (e.g. real name requirements) that it seems like they assumed it would be a given that people would want to use it. When they later realized their error they tried to cram it down everyone’s throats with stuff like YouTube comments only working from Google Plus accounts.
Firstly, that whole account-unification thing where YouTube accounts were getting merged with Google[+] logins. That rubbed me the wrong way.
Then the Google+ promotional stuff all talked about how you could use "Circles" to silo posts to different "circles" of friends. It sounded very complicated and I was worried that I'd publish something snarky to the wrong group of friends :)
I wonder how many others had the same concern? Given that Steve Yegge accidentally published one of his rants to the public that was meant purely for internal Google consumption (I think that was on G+ ...?) that might have been a legit thing to be wary of.
There was also the very minor annoyance of G+ taking over the + operator in Google search (previously you could say +keyword instead of "keyword" to force literal search), but I don't think that would have swayed me against joining.
If you can’t solve the chicken and egg problem of engagement then nothing else really matters.
I don't personally care if a product includes AI, it's the pushiness of it that's annoying.
That, and the inordinate amount of effort being devoted to it. It's just hilarious at this point that Microsoft, for example, is moving heaven and earth to put AI into everything office, and yet Excel still automatically converts random things into dates (the "ability" to turn it off they added a few years ago only works half the time, and only affects csv imports) with no ability to disable it.
Hopefully after the pop rather than shoving it in our face they can return to advertising at us to use the things, and the things needing to prove themselves to get to real sales, rather than corporations getting 10% stock pumps in a day based on statistics about how "used" their AI stuff is while they don't tell the market how few people actually chose to use their AI stuff rather than just becoming a metric when it was pushed on them.
I agree with you in principle, but in practice these two are currently inextricable; if there's AI in the product, then it will be pushed / impossible to turn off / take resources away from actual product improvement.
I mean, c'mon, its literally called the fucking windows key and it doesn't work. As per standard Microsoft it's a feature that worked perfectly on all versions before cortana (their last "ai assistant" type push), i wonder what new core functionalities of their product they're going to fuck up and never fix.
Windows as an OS really kind of peaked around Windows 7 IMO... though I do like the previews on the taskbar, that's about the only advancement since that I appreciate at all... besides WSL2(g) that is. I used to joke that Windows was my favorite Linux distro, now I just don't want it near me. Even my SO would rather be off of it.
Microsoft could have made Windows privacy respecting, continued investing in WSL, baked PowerToys into the OS, etc. and actually made one hell of a workhorse operating system that could rival the mac for developer mindshare. They could partner with Google and/or Samsung and make some deep Android integration to rival Apple's ecosystem of products. Make Windows+Android just as seamless and convenient as mac + iOS.
Instead they opted for forced online accounts, invasive telemetry, and ads in the OS instead of actually trying to keep and win over the very enthusiasts that help ensure their product gets chosen in the enterprise world where they make their cash.
Now they're going to scrap the concept of Windows as something you interact with directly all together and make it "Agentic" whatever the hell that means.
I don't think their bet is going to pay off, especially if the bubble crashes. I think it will be one of the biggest blunders and mistakes that Microsoft will have made.
Just to push their annoying google assistant
I can only hope they won't change it back at the next update (already happened once).
Although I never saw anybody reporting it was actually useful, it's tasteful, accessible, and completely out of your way until you need it.
You can disable AI in Google products.
E.g. in Gmail: go to Settings (the gear icon), click See all settings, navigate to the General tab, scroll down to find Smart features and personalization and uncheck the checkbox.
> Important: By default, smart feature settings are off if you live in: The European Economic Area, Japan, Switzerland, United Kingdom
(same source as in grandparent comment).
(I desperately want to disable the AI summaries of email threads, but I don't want to give up the extra spam filtering benefit of having the smart features enabled)
Google now "helpfully" decides that you must want a summary of literally every file you open in Drive, which is extra annoying because the summary box causes the UI to move around after the document is opened. The other day I was looking at my company's next year's benefits PDFs and Gemini decided that when I opened the medical benefits paperwork that the thing I would care about is that I can get an ID card with an online account... not the various plan deductibles or anything useful like that.
I turned off the "smart" features and the only thing that changed is that the nag box still pops up and shifts the UI around, but now there's a button that asks if you want a summary instead of generating it automatically.
And the worst thing is not only is it being pushed, it is being pushed at the expense of UI/UX. No, Google, I don't need 'help to write' or 'to summarize this document'. I can read and write just fine. And the worst thing of all is that you can't turn it off because they'll just move it around every other week.
I want to choose the extensions that go into my browser. I don't even use the browser's credential manager, and I've gotten to a point where I'm just not sure anything is actually getting better.
I will say that the Gemini answers at the top of Google searches are hit or miss, and I do appreciate that they're there. That said, I'm a bit mixed as the actual search results beyond that seem to be getting worse overall. I don't know if it's my own bias, but when the Gemini answer is insufficient, it feels like the search results are just plain off from what I'm looking for.
ai features in the right context are truly awesome, but the engagement hacking is getting old.
Maybe I'll ask Gemini to write one...
You're completely correct, that's fair criticism. The excitement made me skip the basics. Here's a quick breakdown:
What it does: It's a new optimization algorithm that finds exceptionally good solutions to the MAX-CUT problem (and others) very quickly.
What is MAX-CUT: It's a classic NP-hard problem where you split a graph's nodes into two groups to maximize the number of edges between the groups. It's fundamental in computer science and has applications in circuit design, statistical physics, and machine learning.
How it works (The "Grav" part): It treats parameters like particles in a gravitational field. The "loss" creates an attractive force, but I've added a quantum potential that creates a repulsive force, preventing collapse into local minima. The adaptive engine balances these forces dynamically.
Comparison: The script in the post beats the 0.878... approximation guarantee of the famous Goemans-Williamson algorithm on small, dense graphs. It's not just another gradient optimizer; it's designed for complex, noisy landscapes where Adam and others plateau.
I've updated the README with a "Technical Background" section. Thanks for the push—it's much better now.
LLM's are a product that want to data collect and get trained by a huge amount of inputs, with upvotes and downvotes to calibrate their quality of output, with the hope that they will eventually become good enough to replace the very people they trained them.
The best part is, we're conditioned to treat those products as if they are forces of nature. An inevitability that, like a tornado, is approaching us. As if they're not the byproduct of humans.
If we consider that, then we the users get the shorter end of the stick, and we only keep moving forward with it because we've been sold to the idea that whatever lies at the peak is a net positive for everyone.
That, or we just don't care about the end result. Both are bad in their own way.
Sounds like a return to "Clippy the paperclip" or the dog from the ill fated Microsoft Bob [1] that insisted on always popping up every five to ten minutes with something like: "I see you may be entering a ????, would you like to make it a ??? ???".
And I 'member that you could program it from VBA somehow. Think via OLE, but I was a kid back in the Clippy era.
Which meant you could use it in Internet Explorer but not anywhere else. But it did make for some interesting web pages. I built a custom one with the mascot of the university I was attending at the time. It was, let's say, some peak 1990s internet. (Never shipped it to anyone, just had it internally.)
That took some non-trivial web searching. "Microsoft" "Agent" and most of the other keywords are pretty well covered by a few million other web pages by now.
ActiveX and OLE... technologies ahead of their time, eh. VB, VBA, Internet Explorer, standalone VBScript, C/C++ - didn't matter, it all was (trivially) interoperable.
I don’t know what version of Gemini they’re stuffing into Google products, but sheets, docs, and colab/data science agent are all bad experiences.
If you aren’t putting something comparable to good paid models into your product then don’t bother putting that feature out.
Once you train your users that your ai is half baked junk they’re not coming back to waste their time with it. It’s 10x as frustrating than regular product failures.
As far as I can tell Gemini in gsuite can do nothing other than summarise text and regular LLM q&a (but with Gemini’s perennially sad, apologetic persona)
If the nagging didn't work would companies keep doing it? Someone's KPIs must be increasing for them to keep doing it.
It's fucking Clippy all over again
I want in Text to speech (TTS) engines, transliteration/translation and... routing tickets to correct teams/persons would also be awesome :) (Classification where mistakes can easily be corrected)
Anyways, we used TTS engine before openai - it was AI based. It HAD to be AI based as even for a niche language some people couldn't tell it was a computer. Well from some phrases you can tell it, but it is very high quality and correctly knows on which parts of the word to put emphasis on.
https://play.ht/ if anyone is wondering.
It's still AI, of course. But there is distinction between it and an LLM.
[0] https://github.com/openai/whisper/blob/main/model-card.md
Seems kinda weird for it not to meet the definition in a tautological way even if it’s not the typical sense or doesn’t tend to be used for autoregressive token generation?
Audio models tend to be based more on convolutional layers than Transformers in my experience.
Idk what the definition of an LLM is but it’s indisputable that the technology behind whisper is a close cousin to text decoders like gpt. Imo the more important question is how these things are used in the UX. Decoders don’t have to be annoying, that is a product choice.
On second thought this probably depends on the caption language.
Your point about the caption language is probably right though. It's worse with jargon or proper names, and worse with non-American English speakers. If we they don't even get right all the common accents of English, I have little hope for other languages.
The minimal grammatically correct sentence is simply a verb, and it's an exercise to the reader to know what the subject and object are expected to be. (Essentially, the more formal/polite you get, the more things are added. You could say "kore wa atsu desu" to mean "this is hot." But you could also just say "atsu," which could also be interpreted as a question instead of a statement.)
Chinese seems to have similar issues, but I know less about how it's structured.
Anyway, it's really nice when Japanese music on YouTube includes a human-provided translation as captions. Automated ones are useless, when it doesn't give up entirely.
It does seem to do a few clever things. For lyrics it seem to first look for existing transcribed lyrics before making their own guesses (Timing however can be quite bad when it does this). Outside of that, AI transcribed videos is like an alien who has read a book on a dead language and is transcribing based on what the book say that the word should sound like phonetically. At times that can be good enough.
(A note on sound quality. It not the perceived quality. Many low res videos has perfectly acceptable, if somewhat lossy sound quality, but the transcriber goes insane. It likes prefer 1080p videos with what I assume much higher bit-rate for the sound.)
and here's Jeff Geerling 15 months ago showing how to use Whisper to make dramatically better captions: https://www.youtube.com/watch?v=S1M9NOtusM8
I assume Google has finally put some of their multimodal LLM work to good use. Before that, they were embarrassingly bad.
For the most part, Whisper does much better than stuff I've tried in the past like Vosk. That said, it makes a somewhat annoying error that I never really experienced with others.
When the audio is low quality for a moment, it might misinterpret a word. That's fine, any speech recognition system will do that. The problem with Whisper is that the misinterpreted word can affect the next word, or several words. It's trying to align the next bits of audio syntactically with the mistaken word.
Older systems, you'd get a nonsense word where the noise was but the rest of the transcription would be unaffected. With Whisper, you may get a series of words that completely diverges from the audio. I can look at the start of the divergence and recognize the phonetic similarity that created the initial error. The following words may not be phonetically close to the audio at all.
You don't actually state whether you believe Parakeet is susceptible to the same class of mistakes...
What people want is something that is better than nothing, and in that sense I can see how automatic captions is transformative in terms of accessibility.
Subtitles are good zo
These days when the term "AI" is thrown around the person is usually talking about large language models, or generative adversarial neural networks for things like image generation etc.
Classification is a wonderful application of ML that long predates LLMs. And LLMs have their purpose and niche too, don't get me wrong. I use them all the time. But AI right now is a complete hype train with companies trying to shove LLMs into absolutely anything and everything. Although I use LLMs, I have zero interest in an "AI PC" or an "AI Web Browser" any more than I have a need for an AI toaster oven. Thank god companies have finally gotten the message about "smart appliances." I wish "dumb televisions" were more common, but for a while it was looking like you couldn't buy a freakin' dishwasher that didn't have WIFI and an app and a bunch of other complexity-adding "features" that are neither required or desired by most customers.
I very much do want what used to be just called ML that was invisible and actually beneficial. Autocorrect, smart touch screen keyboards, music recommendations, etc. But the problem is that all of that stuff is now also just being called "AI" left and right.
That being said I think what most people think of when they say "AI" is really not as beneficial as they are trying to push. It has some uses but I think most of those uses are not going to be in your face AI as we are pushing now and instead in the background.
FWIW, 10+ years ago I was arguing that your old pocket calculator is as much of an AI as anything ever could be. I only kinda stopped doing that because it's tiring to argue with silly buzzwords, not because anything has changed since. When "these things were called ML" ML was just a buzzword, same as AI and AGI are now. I'm kinda glad "ML" was relieved of that burden, because ultimately it means a very real thing (which is just "parametrizing your algorithm by non-hardcoded values"), and (unlike with basic autocorrect, which no end user even perceives as "AI" or "ML") when you use ChatGPT, you don't use "ML", you use a rigid algorithm not meaningfully different from what was running on your old pocket calculator, except a billion times bigger and no one actually knows what it does.
So, yes, AI is just a stupid marketing buzzword right now, but so was ML, so was blockchain, so was NoSQL and many more. Ultimately this one is more annoying only because of scale, of how detrimental to society the actions of the culpable people (mostly OpenAI, Altman, Musk) were this time.
And I hope no one gets started about how "AI" is an inaccurate term because it's not. That's exactly what we are doing: simulating intelligence. "ML" is closer to describing the implementation, and, honestly, what difference does it make for most people using it.
It is appropriate to discuss these things at a very high level in most contexts.
What I definitively don't want, yet it's what is currently happening, is a chatbot crammed into every single app and then shoved down your throat.
But we do have to acknowledge that AI is very much turned into an all encompassing term of everything ML. It is getting harder and harder to read an article about something being done with "AI" and to know if it was a custom purpose built model to do a specific task or is it throwing data into an LLM and hoping for the best.
They are purposefully making it harder and harder to just say "No AI" by obfuscating this so we have to be very specific about what we are talking about.
Wow, you are an optimist. I do feel "it's close", but I wouldn't bet this close. But I wouldn't argue either, I don't know. Also, when it really pops, the consequences will be more disastrous than the bubble itself feels right now. It's literally hundreds of billions in circular investing. It's absurd.
That could have been an amazing experience where the AI told me exactly how to use the product. That's what I want. It's not what I got.
Spoiler: you didn't.
Well, if you phrase it this way, then yes, people want this. AI can be useful, and integration is beneficial. But if we are talking about the momentary hype, then no, most people are against stupidly blindly shoving AI into something and getting annoyed with it the whole time.
Personally, I would prefer for apps to safely open up for any kind of integration, and AI being just one automation of many, whatever one prefers. It's so annoying for everything being either a walled garden, guarding every little bit they can grab; or having apps open, but so limited in what they actually can do, that you are basically forced to the walled gardens.
No? If anything, adding AI features to something is just driving away your user base. No one asked for a built-in AI. Why not provide an extension?
Have you seen usage statistics of AI integrations?
I personally don't like them, but I don't expect that I am a representative user. Nor are the people I know.
Also I believe some agentic tasking can make sense: scroll through all the Kindle unlimited books for critically acclaimed contemporary hard sci-fi.
But stapling on a chat sidebar or start page or something seems lacking in imagination.
I have it connected to a local Gemma model running in ollama and use it to quickly summarize webpages, nobody really wants to read 15 minutes worth of personal anecdotes before getting to that one paragraph that actually has relevant information, and for finding information within a page, kinda like ctrl-f on steroids.
The machine is sitting there anyway and the extra cost in electricity is buried in the hours of gaming that gpu is also used for, so i haven't noticed yet, and if you game, the graphics card is going to be obsolete long before the small amount of extra wear is obvious. YMMV if you dont already have a gaming rig laying around
literally googles first hit for me: https://www.reddit.com/r/Cooking/comments/jkw62b/i_developed...
I think its technically experiemntal, but ive been using this since day one with no issue
Openwebui is compatible with the firefox sidebar.
So grab ollama and your prefered model.
Install openwebui.
Connect openwebui to ollama
Then in firwdox open about:config
And set browser.ml.chat.provider to your local openwebui instance
Google suggests the you might also need to set browser.ml.chat.hideLocalhost to false. But i dont remember having to do that
So grab ollama and your prefered model, install openwebui.
Then open about:config
And set browser.ml.chat.provider to your local openwebui instance
Google suggests the you might also need to set browser.ml.chat.hideLocalhost to false. But i dont remember having to do that
I think there is a ton of potential for having an LLM bundled with the browser and working on behalf of the user to make the web a better place. Imagine being able to use natural language to tell the browser to always do things like "don't show me search engine results that are corporate SEO blogspam" or "Don't show me any social media content if its about politics".
But a more nuanced is: the term "AI" has become almost meaningless as everything is being marketed as AI, with startups and bigger companies doing it for different reasons. However, if you mean GenAI subset, then very few people want it, in very specific products, and with certain defined functionality. What is happening now though is that everybody and their mum try to slap it everywhere and see if anything sticks (spoiler: practically nothing does).
However, I think there is a demand of at least one (me) for a Linux system with no AI whatsoever. Firefox could make itself the browser of choice for the minority that don't want any AI. Sure, you can configure it to be AI free, but that is a bit like being able to be vegan at a meaty restaurant where you can always spit out the meat.
Firefox has been struggling of late and they don't do scoped CSS, which makes it as good as IE6 to me, but I think they could get their mojo back by being cheerleaders for the minority that have decided to go AI free. This doesn't mean AI is bad, but there is a healthy niche there.
Apart from anything else, there are new browsers like Atlas that are totally AI. I would say that an AI enabled Firefox is not going to compete with Atlas, but AI free is a market that could be dominated by them.
There is going to be a growing market for no AI. In my own case, my dad was 'pig butchered by an AI chatbot' to die penniless, so I have opinions on AI. Sam Altman would not want to meet me on a bad day, unless he has some AI that specialises in extreme ultraviolence.
Then there is an ever growing army of people that have lost their job to AI to get nothing but rejections from AI powered job boards.
Then there are those that have lost friends to AI psychosis, then there are those that have no water and massive utility bills due to AI data centers. The list goes on!
Sounds like I need to put together an AI free operating system with AI free browser for those that have their own reasons for resenting AI!
If the web doesn’t adapt, a lot of high-quality content will slowly disappear from the “AI layer” of discovery.
We’re trying to document this shift here: https://github.com/ai-first-guides/first.ai/blob/main/docs/i...
in general i agree with you, adding an AI chat window to an app that isn't an AI chat app is almost always a detriment. but i think it's shortsighted to assume there won't be other important use cases for AI, and we're in the experimentation phase right now where companies are trying to learn what that looks like. it's just unfortunate that there's so much incentive for apps to frame their AI chat as the best new thing ever and you should really use it, instead of introducing it more subtly.
It makes my life SO much easier (less time spent on editing config files, less chance to make a silly typo whilst writing scripts).
It definitely has its place.
I like to keep AI at arms length, it's there if I want it but can fuck off otherwise
Lots of people really do seem to want it in everything though
I am semi-confident that LLM backed interfaces will be the future of many UIs though. When it works it just is a way better UX. A smart chat instead of a <form> or crawling through pages of search results is just nicer.
It is bridging the gap between the hard data computers use and the generalized way humans communicate.
Building AI the Firefox way: Shaping what’s next together - <https://connect.mozilla.org/t5/discussions/building-ai-the-f...>
I want AI in my email to speed up (and avoid typos) in replying.
I want AI in my news feed to pull the topics that are interesting to me.
I want AI in online shopping to filter and recommend products by complex conditions.
I want AI in my car to make me safer.
I want AI in my calendar to schedule with a minimum of interruptions.
I want AI in my work chats to answer questions that people have already asked me.
I want AI to make clinical diagnoses more accurate.
I want AI for a thousand things and most people do, or will.
The AI to give you sources you check if you need the answer to be right. That's still better than a google search in many cases.
Most (all?) of it runs locally too
Mmm, summarized garbage.
>Also I imagine you frequently read summaries of books
This isn't what LLM summaries are being used for however. Also, I don't really do this unless you consider a movie trailer to be a summary. I certainly don't do this with books, again, unless you think any kind of commentary or review counts as a summary. I certainly would not use an LLM summary for a book or movie recommendation.
Managers think they want AI but they actually want their people to work faster or better. Higher managers think they want AI so they can save money, or at least not fall behind the competitors, if those were to use AI to get an advantage.
Companies making software think they want AI because their competitors are using it, and they think the users want AI so the software can be perceived as modern, not falling behind.
And so on, and so on... other than Nvidia, openAI, Anthropic, etc, no one really wants AI.
There are also a lot of subtle AI tools that aren’t in-your-face LLM prompts that flatter you with “Excellent question!”. It’s great having my photo library automatically annotated so I can search for things like “moose” and it will bring up that picture of the moose we saw, rather than me having to remember what year it happened and scroll through photos until I find it.
I like to have AI only when I specifically want it. Usually I just code in Emacs. If I specifically want help with something then for an IDE experience I will use the TRAE coding agent. For command line, I will use gemini-cli or codex. I like to use AI coding help 4 or 5 times a week. As an example, today I wanted some Python code that used a few libraries converted to Common Lisp (using several popular CL libraries). TRAE one-shotted this for me in two minutes. I think it would have taken me over 20 minutes to write it myself.
AI is OK for easy stuff you can do yourself, and save time.
The book AI Atlas tells a good narrative about natural resources used for AI, BTW.
Turns out that "if" part is fantastically difficult for some types to fathom, and what we're all experiencing now is just the same add-on tech-stench that has been typical of every digital era before us:
1970s: Calculators, calculators, calculators!
1980s: Miniaturised, digital quartz clocks anywhere they can fit.
1990s: Wouldn't this toaster be better.. WITH A LCD SCREEN?
2000s: MP3 players must outnumber the human population. No object or space should be without shitty, tinny music.
2010s: This easy-to-use device would be wonderfully enshittified by removing all of the buttons and switching to a touchscreen aka "Smart"-appliances.
2020s: AI, AI, AI!
I like LLM's, I've even build my own personal agent on our Enterprise GPT subscription to tune it for my professional needs, but I'd never use them to learn anything.
However 99% of the times i use this isn't because i need an accurate summary but because i come across some overly long article that i do not even know if i'm interested in reading, so i have Mistral Small generate a summary to give me a ballpark of what the article is even about and then judge if i want to spend the time reading the full thing or not.
For that use case i do not care if the summary is correct, just if it is in the ballpark of what the article is all about (from the few articles i did ended up reading, the summary was in the ballpark well enough to make me think it does a good enough work). However even if it is incorrect, the worst that can happen is that i end up not reading some article i might find interesting - but that'd be what i'd do without the summary anyway since because i need to run my Tcl/Tk script, select the appropriate prompt (i have a few saved ones), copy/paste the text and then wait for the thing to run and finish, i only use it for articles i'm in already biased against reading.
How do I know what I'd be reading is correct?
To your question: for the most part, I've found summaries to be mostly correct enough. The summaries are useful for deciding if I want to dig into this further (which means actually reading the full article). Is there danger in that method? Sure. But no more danger than the original article. And FAR less danger than just assuming I know what the article says from a headline.
So, how do you know its summaries are correct? They are correct enough for the purpose they serve.
Of course, as more and more pieces of writing out there become slop, does any of this matter?
Recipe pages full of fluff.
Review pages full of fluff.
Almost any web page full of fluff, which is a rapidly rising proportion.
> And how would I know the LLM has error bounds appropriate for my situation?
You consider whether you care if it is wrong, and then you try it a couple of times, and apply some common sense when reading the summaries, just the same as when considering if you trust any human-written summary. Is this a real question?
I was thinking more along the lines of asking an LLM for a recipe or review, rather than asking for it to restrict its result to a single web page.
For example - you summarize a YouTube link to decide if the content of it is something you're interested in watching. Even if summarizations like that are only 90% correct 90% of the times it is still really helpful, you get the info you need to make a decision to read/watch the long form content or not.
The opportunity cost of "missing out" on reading a page you're unsure enough about to want a summary of is not likely to be high, and similarly it doesn't matter much if you end up reading a few paragraphs before you realise you were misled.
There are very few tasks where we absolutely must have accurate information all the time.
If an app is a gateway to a bunch of data, it's cool to be able to "talk" to that data via any built-in LLM-based stuff, but typically the app is just a frontend anyway in that case, so the app isn't really needed.
I've vibe coded a few Godot games. It's all good fun.
But now everything is forcing it. Google is telling people what rocks are tasty, on Reddit bots are engaging with bots.
From what I can tell the only way to raise VC money is by saying AI 3 times. If the ritual is done correctly a magic seed round appears.
As they say, don't hate the player, hate the game.
Other than that, I don't think I'd be happy to see AI anywhere else. I pretty much don't want no AI in my operating system, browser.
There are certainly lots of great use cases, the problem is everyone is shoving it everywhere because they don’t want to feel behind the times and for every great use case there are several times where it accomplishes nothing but makes the UI worse.
Absolutely. I want a browser with AI -- just not the browser Mozilla wants to build. I want my browser to use AI-based adblocking and content filtering. I want my AI browser to notice when the site sends some stupid sticky high Z-index thing down the pipe and just quietly not show it to me at all. I want my AI browser to automatically detect cookie dialogs and click "Reject All" and if that option isn't available, I want it to parse the "Cookie Preferences" page and click all the buttons that equate to "Reject All".
I want an AI layer in my phone that spoofs my location and my contacts so that apps that insist on seeing those things see fake data that nevertheless looks plausible.
Best of all, I want the AI agents in my browser and my phone to do their work without leaving any trace of their activities so that the server on the other end cannot tell that I even have an AI agent at all.
Most of the above is possible now but it requires a plethora of different tools that are not cleanly integrated. And no VC is going to pay you to build such an integrated tool because it would not create a continuing revenue stream or a continuing stream of harvestable data compromising the user's privacy.
We are a very fucked-up industry.
I know there are tools where you can do it yourself but it is a hellish mess. I just move it drive to drive through the decades until it comes.
Unfortunately got to meet those KPIs.
I also wouldn’t want to go back to only web search for finding things out. Search engines are generally inferior.
AI ad blocking might be nice.
"How do I change the resource limits for CPU core count"
Beyond that I've never used Gemini for any actual purpose.
For example, translation can be considered AI, and I find it very useful, it is local too. Other AI features that could be nice would be speech-to-text, text-to-speech, advanced spellchecking, text autocomplete, etc... Bonus points if local models are used. I also see nothing wrong with having a "ask a LLM" entry in the right click menu like you have search, I think it is a common enough thing for people to do.
The problem with many AI features in software is that they serve no purpose besides "hey look, we have AI". Usually in the form of some button or text field that is always visible and does nothing more than prompt a poorly tuned LLM.
AI is basically only a shortcut to wikipedia, and i always anyway have to double check any AI response, making it kind of useless.
Its this constant fight that everyone must CAPTURE all revenue opportunities at the cost of complete overwhelming tsunami of bad forceful decisions on users, all JUST INCASE its an actual revenue stream that they could be missing out on, before even knowing if a single user gives the slightest shit about it
a sandboxed LLM ad block or filter could be handy, for instance
It's bad enough what Google did to search; a future where the only thing you get back is a) what the machine allows you to see or create (which may be determined by the built-in agent or by the programmers); b) what the machine wants you to see, & modified to be in line with its whims; & c) hallucinated slop where it is difficult to determine what is real, what is human-originated, & what is constructed out of whole cloth.
Well, yes. It's extremely useful. However, the hype bubble means it's getting added everywhere even when there's not a clear and vetted use case.
It works really well for navigating docs as a super-charged search--much better at mapping vague concepts and words back to the official terminology in the docs. For instance, library Z might have "widgets" and "cogs" as constructs, but I'm used to library A which has similar constructs "gadgets" and "gears". I can explain the library A concepts and LLMs will do a pretty good job of mapping that back to the library Z concepts--much better than traditional search engines can do.
This shit makes me want to stop interacting with tech altogether and live on a farm. I don't know how much more of this I can take.
Yeah, they do. Go talk to anyone who isn't in a super-online bubble such as HN or Bsky or a Firefox early-adopter program. They're all using it, all the time, for everything. I don't like it either, but that's the reality.
Not really. Go talk to anyone who uses the internet for Facebook, Whatsapp, and not much else. Lots of people have typed in chatgpt.com or had Google's AI shoved in their face, but the vast majority of "laypeople" I've talked to about AI (actually, they've talked to me about AI after learning I'm a tech guy -- "so what do you think about AI?") seem to be resigned to the fact that after the personal computer and the internet, whatever the rich guys in SF do is what is going to happen anyway. But I sense a feeling of powerlessness and a fear of being left behind, not anything approaching genuine interest in or excitement by the technology.
We can take principled stands against these things, and I do because I am an obnoxiously principled dork, but the reality is it's everywhere and everyone other than us is using it.
They are already preached at that they need a new phone or laptop every other year. Then there's a new social platform that changes its UI every 6 months or quarterly, and now similarly for their word processors and everything.
This is kinda like how if you ask everyone how often they eat McDonald's, everyone will say never or rarely. But they still sell a billion burgers each year :) Assuming you're not polling your Bsky buddies, I suspect these people are using AI tools a lot more than they admit or possibly even know. Auto-generated summaries, text generation, image editing, and conversation prompts all get a ton of use.
Do you know someone? Using Firefox nowadays is itself a "super-online bubble"
I guess they key is not in your face when you don't want them and actually useful.
Step 2. ???
Step 3. Profit
I most definitely do.
I want to be able to type into Finder on my Mac to rename all the files a certain way, without spending 10 minutes figuring out the right regex for it.
I want to be able to type into Firefox to go through 50 different versions of the current URL, using a different US state parameter for each, and download the table it shows into a single combined CSV with an added column for "state".
Every day there's 20 things like this. I absolutely want everything in my OS and browser to be exposed to an LLM that can do everything so much faster. Without the intermediate stage of having it write a script to do it. It would save so much time.
Unfortunately we're not quite there yet because the GUI programs we use haven't exposed all the views and actions. But hopefully soon!
So, yes, I want AI in "everything".
And it's not a waste of resources if it's not triggered automatically.
In fact, I'd say you're an edge case's edge case. There should be a word for that. Maybe "one-off."
The use-case, which generalised is "pull some information from a web page", is far less niche, and I'd argue extremely common.
I know a lot of people - including non-technical people - who spend a lot of time doing that in ways ranging from entirely manual to somewhat more sophisticated, and the more technically knowledgeable of those have started looking for AI tools to help them with that.
To the extent users "don't want" AI available for things like this, it is mostly because they don't know AI could help with this.
E.g. just a few days ago, I had someone show me how they painstakingly copied column by column from the exact same Notion site I mentioned into a Google sheet, without realising it was trivially automatable. Or rather: Trivially automatable to a technical user like me. But it could be trivially automatable to anyone with relatively little integration effort in the browsers.