Gemini "duck" demo was not done in realtime or with voice
twitter.com
twitter.com
My group (3 of us) bought a moisture sensor to plug into the pi, and had the idea to make a "flood detection system" that would be housed under a bridge, and would send an email to relevant people when the bridge home from work is about to flood.
So for our demonstration, we had a guy in the back of the class with gmail open ready to send an email saying some variation of "flood warning". Our script was literally just printing lines with wait statements in between. Running the script, it prints to the screen "awaiting moisture", and after 3 seconds it will print "moisture detected". In that 3 seconds I dip the sensor into the glass of water. Then the script would wait a few more seconds before printing "sending email to xxx@yyy.com". We then opened up our email, our mate at the back of the room hit send, and an email appeared saying flood warning, and we would get full marks.
We’d set up a dummy HMI and have someone pressing buttons on it for the demo, and someone in the next room manually driving outputs and inputs to make it seem like it was working. Very common.
Like you say it can be useful as a way to uncover overall process or UX issues before all the internals are coded.
I'm quite open about what I'm doing though, as most of our clients are reasonable.
To me it sounds like lying...
Context matters a ton though. Are you presenting the demo as if the events are being automatically triggered (as in the OP) or are you presenting as this is your plan? Explicitly. If it is implicit, it's deceptive. If you explicitly do not say what parts are faked, it is lying. Of course in a magic show this is totally okay because you're going to the show with the explicit intention to be lied to, but I'm not convinced the same is true for business but I'm sure someone could make a compelling argument.
And yet, you’ll still have people here acting like it’s totally fine.
As you said, it’s one thing to demonstrate a prototype, a “this is how we intend for it to work.” It’s a whole other thing to present it as the real deal.
It’s called prototyping, in this case it would be a hi-fi prototype, and it lets you communicate ideas, test the implementation and try out alternatives before committing to your final product.
Lo-fi prototyping usually precedes it and is done on pen and paper, Figma, or similar very basic approaches.
Google (unsurprisingly) uses it and has instructional videos about it: https://youtu.be/lusOgox4xMI
Relying on blackbox components first mocked then later swapped (or kept for testing) with the real implementation is also a thing, especially when dealing with hardware. It's a bit hard to put moisture sensors on CI!
That said...
I would have recommended at least being open about it being a mockup, but even when doing so I've had customers telling me "so why can't you release this tomorrow since it's basically done!? why is there still two months worth of work, you're trying to rip us off!!!"
Garbage in, garbage out.
Here’s a whole list of projects intended for kids.
https://all3dp.com/2/best-raspberry-pi-projects-for-kids/
It includes building out a whole weather station which includes a humidity sensor as one of the many things it can do.
yes for whomever organized such a curse and didn't give such guidance.
And besides curse asked for project to do something. It did. It printed lines. We can call the email gimmick, the marketeering strategy, making a turd look good.
Don't blame students for failure of whomever designed the curse.
The problem with the email isn’t it’s a gimmick etc. it’s that it appears quite clear that the students created the impression that it was the Pi doing it.
Your excuse that it’s difficult for first year college students with no coding experience to do something useful with the Rapberry Pi is disproven by the fact that there exist many extremely useful projects that kids with no coding experience can do, so college students almost certainly should be able to do without needing to resort to gimmicks.
So I don’t understand your complaints about the course. It’s clearly not too hard which is what you’re implying. And if you’re suggesting that the wording for the project wasn’t clear enough then that’s a huge claim to make considering you don’t know what the wording was.
Also, college (at least in the U.S.) was never about playing funny word games with the professor. There’s a level of maturity, reasonableness, and respect that is expected of the students. None of which is indicated in the response here.
Given that the general teaching style of colleges isn't unique to the US, and based on my experience throughout my degree at a similar institution, I somehow doubt that statement.
> It’s clearly not too hard which is what you’re implying.
It sounded like the students received literally no guidance, in the way the course is described. These types of assignments usually result in those with previous programming experience showing off their skills, while the actual rookie students are left in the mud. I.e. an assignment that targets the top 20% of the class.
Regardless, to my knowledge I never cheated during my college degree, but I can't hold it against people that do. Criticism such as yours disregards the reality that students face, pressure to graduate with good marks and whatnot. Not cheating will put you at an disadvantage, because your competition is actively doing so and they are already skewing the marks that way. If the intention of the assignment was to identify honest work, it was certainly structured wrong (as a required submission would have eliminated the cheaters).
Looking at this a different way, they gave first-year students, likely with no established pre-requistites, an open-ended project with fixed hardware but no expectation to submit the final project for review. If they wanted to verify the students actually developed a working program, they could have easily asked for the Pi's to be returned along with the source code.
A project like this was likely intended to get the students to think about the "what" and not worry so much about the "how." Faking it entirely may have gone a bit further than intended, but would still meet the goal of getting the students to think about what they could do with this computer (if they knew how)
While university instructors can vastly underestimate student's creativity, they are, generally speaking, not stupid. At the very least, they know if you don't tell students to submit their work, you can often count on them doing as little as possible.
Wait, is your argument honestly "it's not cheating because they just trusted the students"?
There's a huge difference between demoing something as "this is what we did" vs "we didn't quite get there, but this is what we're envisioning."
Edit: You all are responding very weirdly. The cheating is because you're presenting "something" that is not that thing. Put a dog in a dress and call it a pretty woman and I'll call you a conman.
In US colleges at least (only because that’s where I have personal experience…not because I believe standards are any higher or lower here), this is cheating if they led their professor to believe that it was indeed the raspberry pi sending out an email rather than someone at the back of the class.
But what if the Pi was only ever meant as a story-telling device to get the students thinking about the kinds of things computer programs can do?
Sure, some of students would be able tell a story by building a functioning program, but dvsfish simply found another way to tell theirs.
>> It was our first comp sci class ever
>> asked to create "something"
It doesn't matter what that "something" was if you are claiming that you made "something" else. The lie is not necessarily the final result, it is the telling of the final result. The context of the whole thread is also about deceit because rtfa.
They did use the hardware provided, and did use software to accomplish a goal. If the teacher just wanted to test what problem solving skills the students walked in, I’d say that’s a fair result.
For anyone not getting why this is cheating, try this. Pretend you are the teacher. You saw students pull this stunt and say that it was a live demo. Would you feel like you were deceived if later you found out it was not in fact how they presented it? (I'm honestly unsure how you can answer this in the negative)
Or how about you watch an ad of someone taking to a box and the box responding, then you buy the box and it's a pile of rubber bands and paper clips. Would you want your money back?
Well if you're the TA and you're unwilling/too lazy to call out the conman, I call you an accomplice! Also, since when was the ideal scientific rigour ever build on interpersonal trust?
Literally all that matters is that they passed.
Are you trying to bully me or something? Not going to work with me. You've revealed your poor character with that comment.
The Wiimote connection was the star of the show by a long shot :P
The amount of fumbles is monumental.
I know that MS Teams is a more full-featured product suite, but even at companies that used it, we still used Zoom for meetings.
In addition meet insists I click on the same about 4 or 5 different "Got it" feature popups every single call, and every call also insists on asking me if I want to use Duet AI to make my background look shit which just adds to annoyance.
Sure, Teams is a steaming pile of crap to use day-to-day as a chat app, the search is slow and vague - and depending on policy, probably links you to messages that no longer exist in the archive lol. Oh you want to download message history? Nah gotta get an admin to do that bruh.
The only saving grace is that members who can't deal with it are using local MS Office, which has some integrations with 365, thus making it kinda viable. But I feel like it's still a net negative.
I base about 50% of my choice of employer on what they choose in that area.
If you have a problem there’s no one available to help you.
On the MS side they will literally pull an engineer who is writing the code for the product you have a problem for to help resolve the issue if you’re large enough.
The part you see in your browser isn’t the only part of the product a company has to buy. In fact, it’s not even the most expensive bit. If you see the most expensive plans for most SAAS products (ie the enterprise plans) almost the entire difference in costs is driven by support illustrating the importance and value of support.
Google unfortunately is awful at this.
Google has the problem that it's typically the first to encounter a problem, and it has the resources to approach it (from search), but the incentive to monetize it (to get away from depending entirely on search revenue). And, management.
An AI product that makes search irrelevant is an existential threat, but I don’t think Google has the product DNA to pull off a replacement product for search themselves. I heard Google has been taken over by more business / management types, but it is still missing product as a core pillar.
Under Eric Schmidt they were engineer-driven, during the golden era of the 2000s. Nowadays they're MBA driven, which is why they had 4 different messaging apps from different product managers.
Also, "had." Google cleaned things up. They still sometimes do stuff just cause, but it's a lot less now. I still feel like Meet using laggy VP9 (vs H.264 like everyone else) is entirely due to engineer stubbornness.
Procuring Azure is a good option for lots of companies because most companies' IT staff know AD and Microsoft in general, and Microsoft's cloud offers them a way to use the same (well, not the same, but it's too late by then) tools to manage their company IT.
I'm not disagreeing with its success, but I do think they had a much simpler journey, as to my understanding a lot of it involved cloudifying their locked-in enterprise customers, rather than diversifying into new markets.
I also don't see anything big Google has leveraged Android for, besides Pixel, which is actually more to cement Android cause they know they don't have enough control with software alone. At least I have decent amount of faith in them pulling that off.
And I hate Teams personally, but lots of teams use it.
I bet most team members who switched from Slack to Microsoft Teams do not feel like they consented or were asked for their opinions beforehand.
And this business is so totally different to Google in every way imaginable.
Senior Managers love customer support, SLAs - Google loves automation. Two worlds collide.
Good if you want someone else to google the error message for you though.
Garbage in = garbage out.
If Google cannot deign to assign internal resources and staffing towards providing first-party support for paid products, it's not a good choice over the competition. You're not going to beat the incumbent (Office 365) by skimping on customer service.
Even then usage of AJAX declined rather slowly as it was so established, and indeed even now it's still used by many websites!
Before the invention of the xmlhttprequest there was so little you could do with JS most dynamic content was some version of shifty tricks with iframes or img tags or anything that could trigger the browser to make a server request to a url that you could generate dynamically.
Fetch was the formalization of the xmlhttprequest (hence the use of xhr as the name of the request type ). Jquery wrapped it really nicely and essentially popularized (they may have invented async js leveraging callbacks and the like), the creation of promises was basically the formalization and standardization of this.
So AJAX itself is in fact used almost in the entire totality of the web, the term has become irrelevant given the absolute domination of the technology.
"It was probably Office Web Apps. It was a web-based office suite that was introduced in 2008. It included Word Web App, Excel Web App, Powerpoint Web App, and OneNote Web App. It was not SharePoint, but it was based on SharePoint technology."
Wild that we have to ask these questions.
Plus it's point in time - I'm sending you the document as it is now, and I might start cutting it about or changing it after to send to someone else, but this the version I want to send you.
Google Analytics - acquired 2004, renamed from Urchin Analytics
Google Docs - acquired 2004, renamed from Writely
Youtube - acquired 2005
Android - acquired, 2005 (Samsung have done more to advance the OS than Google themselves)
2008 is late to the party for a docs competitor! Microsoft got the runaround by Google and after Google launched docs they could have clobbered Microsoft which kind of failed to respond properly in kind, but they didn’t push the platform hard enough to eat the corporate market share, and didn’t follow up with a share point alternative that would appeal to the enterprise, and kind of blew the opportunity imo.
I mean to this day Google docs is free but it still hasn’t unseated Word in the marketplace, but the real killer app that keeps office on top is Excel, which some companies built their entire tooling around.
It’s crazy interesting to look back and realize how many twists there were leading us to where we are today.
Btw it was Office Server or Sharepoint Portal earlier (this is like Frontpage days so like 2001?) and Microsoft called it Tahoe internally. I don’t think it became Sharepoint until Office 365 launched.
The XMLHTTP object launched in 2001 and was part of the dhtml wave. That gave a LOT of the capabilities to browsers that we currently see as browser-based word processing, but there were efforts with proprietary extensions going back from there they just didn’t get broad support or become standards. I saw some crazy stuff at SGI in the late 90s when I was working on their visual workstation series launch.
1. Poor Google Drive interface makes managing documents difficult.
2. You cannot just get a first class Google Doc file which you can then share with others over email, etc. Very often you don’t want to just share a link to a document online.
3. Lack of desktop apps.
So why did they never release that and went with Office 365 instead?
https://www.zdnet.com/article/netdocs-microsofts-net-poster-...
tech was based on an acquired company, Google just abused their search monopoly to make it more popular(same thing they did with YT). This has been the strategy for every service they've ever made, Google really hasn't launched a decent in-house product since Gmail and even that was grown using their search monopoly as free advertising
>Google Docs originated from Writely, a web-based word processor created by the software company Upstartle and launched in August 2005
What about Chrome? And Chromebooks?
https://www.chromium.org/developers/design-documents/display...
The process model was the novel selling point at the time from my memory [1].
Which uses the WebKit engine and is kindof a showcase for Safari, granted, but it still exists as distinct browser under that name.
The amount of data that is collected by these cars is massive.
A product requires commitment, it requires grind. That 10% is the most critical one, and Google persistently refuses to push products across the finish line, just giving up on them and adding to the infamous Google Product Graveyard.
Honestly, what is the point? They could just maintain the core search/ads and not pay billions of dollars for tens of thousands of expensive engineers who have to go through a bullshit interview process and achieve nothing.
It's a way to strangle the competition. But also not good for the industry in general.
This is what Google has always cared about. Bring application to the billions of users.
People are forgetting Google is the most profitable AI company in the world right now. All of their products use ML and AI.
So who is losing?
The goal of Gemini isn't to build a chatbot like ChatGPT despite Google having Bard.
The goal for Gemini is to integrate it into those 10 products they have with a billion users.
Is that supposed to be a vote of confidence for the current state of Google search?
> So who is losing?
The people who use their products, which are worse than they’ve been in decades? The people who make the content Google now displays without attribution on search results?
Well, that is trully shocking.
We'll get a definitive answer in a few years. Til then, OpenAI benefits from the $ value from their end of products, MSFT eats the compute costs, but also gets a stock bump.
What would the ideal setup in your opinion?
Word and Excel have been dominant since the early 1980s. Google has never had a real shot in the space.
[0] yes, not literally nobody. I know about the Windows 2.0 Excel or whatever, but the user base compared to WordPerfect or 1-2-3 was tiny up until MS was able to start driving them out by leveraging Windows in the early-mid 90s.
Of all the things, this.
I use both Google and Microsoft office products. One thing that strikes you is just how feature rich Microsoft products are.
Google doesn't look like is serious about making money.
I squarely blame rockstar product managers and OKRs for this. Not everything can be a 1000% profitable product built in the next quarter. A lot of things require small continuous improvement and care over years.
Shopped it around VCs. Got laughed out of all the meetings. "Companies storing their documents on the Internet?! You're out of your mind!"
Globe dot com was basically Facebook, but the critical mass wasn't there. Nor were the smartphones.
Nadella is an all time great CEO. Pichai is an uninspired MBA-type.
The difference is Nadella is a good CEO and Pichai isn’t.
Part of it could also be a result of circumstance. Nadella came at a time when MS was foundering and he had to make what appeared to be fairly obvious decisions (pivot to cloud…he was literally picked because of this, and reducing dependence on Windows…which was an obvious necessary step for the pivot to cloud). Pichai OTOH was selected to run Google when it was already doing pretty well. His biggest mandate was likely to not upset the Apple cart.
If roles were reversed, I suspect Nadella would still have been more successful than Pichai, but you never know. I’d Nadella introduction to the CEO job was to keep things going as they were, and Pichai’s was to change the entire direction of the company, maybe a decade later Pichai would have been the aggressive decision maker whereas Nadella would have been the overly cautious guy making canned demos.
AI is a threat there, but it'd require an AI company to transform the culture of Internet use to stop people 'Googling', and that will require two things: something significantly better than Google Search that's worth switching to, and a company that is willing to reject whatever offer Google makes to buy it. Neither is very likely.
Also "interesting" to see the if results being SEO spam generated using AI will keep seo search viable.
They can't do anything that threatens their main income. They are tied to ads and ads technology, and can't do anything about it.
Microsoft had a crisis and that drives focus. Google... they probably mistreat their good employees if they don't work on ads.
Besides collaborative reuniting is a no feature and there is much more important stuff than this for a word processor to be useful.
I agree with your general post but I disagree with this. Tesla's FSD is so far behind Google it's almost negligent on the part of Tesla despite having so much more data.
What does that mean and why is it bad?
Diversity in marketing is used because, well, your desired market is diverse.
I don't know what it means for it to be surgically precise, though.
Now, I don't know how diverse the AI workforce is at Google, but the YT thumbnails show precisely 50% of white men. Maybe that's what the parent meant by "surgically precise".
I don't mean to imply that companies should avoid displays of diversity, I just mean that it's obvious when it's inauthentic. Virtue signaling in exchange for business is not progress.
But for diversity.
Imagine how cringe it would be if only white guys were allowed to work at Google and they displayed in all their marketing a fully diverse group of non-white girls. That would be... inauthentic.
Just the fact girls are less than guys in IT is something we should demonstrate, understand, change if needed. Not hide behind a facade of 50/50 display everywhere as if the problem was already solved or that it was even a problem in the first place.
Google Keynote feels like it's emulating the Apple keynote from 5 years ago.
And the Apple keynote looks like robots just out of an uncanny valley pretending to be humans - just like keynotes might look in 5 years, but actually made by AI. Apple is always ahead of the curve in keynote trends.
It’s the stillness between “beats” that does it, I think, and the very-constrained and repetitive motion.
Where humans behave so awkwardly that they seem artificial but are just not quite close enough…
If so, Apple have totally nailed the reverse uncanny valley!
That's what Apple keynotes feel like now. It seems like each year, they're trying to make their presentations even more essentially 'Apple.' They crossed the uncanny valley a long time ago.
I used to look forward to watching them live but now I just go back after the event and skip through the videos to see the most relevant bits (or just read the various reports instead).
They oversold its capabilities, but it does still seem that multi-modal models are going to be a requirement for AI to converge on a consistent idea of what kinds of phenomena are truly likely to be observed across modalities. So it's a good step forward. Now if they can just show us convincingly that a given architecture is actually modeling causality.
"do you believe that a pocket of hot air would lead to lower air pressure causing my plane to stall?"
he could barely even phrase the question correctly because it was so awkward. just embarrassing.
I believe it is Hinton that prefers “confabulation” to “hallucination” because it’s more accurate. The example in the discussion about hallucination/confabulation was that of someone who had been present in the room during Nixon’s Watergate conversations. Interviewed about what he heard, he provided a narrative that got many facts wrong (who said what, and what exactly was said). Later, when audio tapes surfaced, the inaccuracies in his testimony became known. However, he had “confabulated truthfully”. That is, he had made up a narrative that fit his recall as best as he was able, and the gist of it was true.
Without the ability to confabulate, he would have been unable to tell his story.
(Incidentally, because I did not check the facts of what I just recounted, I just did the same thing…)
You can tell a story without making up fiction. Just say you don’t know when you don’t know.
Inaccurate information is worse than no information.
The point is that humans can't in general, because we don't actually know which parts of what we "remember" are real and which parts are our brain filling in the blanks. And maybe it's the same for nonhuman intelligences too.
Often we don’t quote people and instead provide a high level description of what they said, e.g. “Harry described the problems with his car.”, where detail is omitted.
After a few months AI developers refined the process to just replicate images so they looked like a human being made them, in effect killing what was the real AI art.
It simply spits out whatever output sequence it feels is most likely to occur after your input sequence. How it defines “most likely” is the subject of much research, but to optimize for factual correctness is a completely different endeavor. In certain cases (like coding problems) it can sound smart enough because for certain prompts, the approximate consensus of all available text on the internet is pretty much true and is unpolluted by garbage content from laypeople. It is also good at generating generic fluffy “content” although the value of this feature escapes me.
In the end the quality of the information it will get back to you is no better than the quality of a thorough google search.. it will just get you a more concise and well-formatted answer faster.
What if the input sequence says "the following is truth:", assuming it skillfully predicts following text, it would mean telling the most likely truth according to its training data.
I would say it’s worse than Google search. Google tells you when it can’t find what you are looking for. LLMs “guess” a bullshit answer.
https://chat.openai.com/share/ca733a4a-7cdb-4515-abd0-0444a4...
https://chat.openai.com/share/dced0cb7-b6c3-4c85-bc16-cdbf22...
Hallucinations are definitely a problem, but they are certainly less than they used to be - They will often say that they aren't sure but can speculate, or "it might be because..." etc.
I think you're slightly mischaracterising things here. It has potential to be at least slightly and possibly much better than that. This is evidenced by the fact it is much better than chance at answering "novel" questions that don't have a direct source in the training data. Why it can do it is because at a certain point, to solve the optimisation problem of "what word comes next" the least complex strategy actually becomes to start modeling principles of logic and facts connecting them. It is not in any systematic or reliable way so you can't ever guarantee when or how well it is going to apply these, but it is absolutely learning higher order patterns than simple text / pattern matching, and it is absolutely able to generalise these across topics.
Seems like there are quite a few obvious possibilities here off the top of my head. Ground truth for correct facts could be:
1) Wikidata
2) Mathematical ground truth (can be both generated and results validated automatically) including physics
3) Programming ground truth (can be validated by running the code and defining inputs/outputs)
4) Chess
5) Human labelled images and video
6) Map data
7) Dependent on your viewpoint, peer reviewed journals, as long as cited with sources.
If it was possible to solve LLM hallucinations with simple Chain-of-Thought style agents, someone would have done that and released a product by now.
The fact that nobody has released such a product, is pretty strong evidence that you can't fix hallucinations via Chain-of-Thought or Retrieval-Augmented Generation, or any other band-aid approaches.
For example, generating json.
You can explicitly follow a defined grammar to get what will always be a valid json output.
Similarly, structured output such as code can be passed to other tools such as compilers, type checkers and test suites to ensure that at a minimum the output you selected passes some minimum threshold of “isn’t total rubbish”.
For unstructured output this a much harder problem, and bluntly, it doesn’t seem like there’s any kind of meaningful solution to it.
…but the current generation of LLMs are driven by probabilistic sampling functions.
Over the probability curve you’ll always get some rubbish, but if you sample many times for structure and verifiable output you can, to a reasonable degree, mitigate the impact that hallucinations have.
Currently that’s computationally expensive, to drive the chance of error down to a useful level, but compute scales.
We may seem some quite reasonable outputs from similar architectures wrapped in validation frameworks in the future, I guess.
…for, a very specific subset of types of output.
But it's unrelated to the problem of LLM hallucinations. A hallucination that's been validated as correct json is still a hallucination.
And if your problem space is simple enough that you can validate the output of an LLM well enough to prove it's free of hallucinations, then your problem space doesn't need an LLM to solve it.
Hmmm… kinda opinion right?
I’m saying; in specific situations, you can validate the output and aggregate solutions based on deterministic criteria to mitigate hallucinations.
You can use statistical methods (eg. There’s a project out there that generates tests and uses “on average tests pass” as a validation criteria) to reduce the chance of an output hallucination to probability threshold that you’re prepared to accept… for certain types of problems.
That the problem space is trivial or not … that’s your opinion, right?
It has no bearing on the correctness of what I said.
There’s no specific reason to expect that just like you can validate output against a grammar to require output that is structurally correct, you can’t validate output against some logical criteria (eg. unit tests) to require output that is logically correct against the specified criteria.
It’s not particularly controversial.
Maybe the output isn’t perfectly correct if you don’t have good verification steps for your task, maybe the effort required to build those validators is high, I’m just saying: it is possible.
I expect we’ll see more of this; for example, this article about decision trees —> https://www.understandingai.org/p/how-to-think-about-the-ope..., requires no specific change in the architecture.
It’s just using validators or search the solution space.
This was all fake. You are taking a collection of cherry picked prompt engineered examples, then dramatizing them for maximum shareholder hype. The music example was just outputting a description of a song, not the generated music we heard in the video.
It’s one thing to release a hype video with what-ifs and quite another to claim that your new multi-modal model is king of the hill then game all the benchmarks and fake all the demos.
Google seems to be in an evil phase. OpenAI and MS must be quite pleased with themselves.
1) Forward looking demoes that demonstrate the future of your product, where it’s clear that you’re not there yet but working in that direction
or
2) Demoes that show off current capabilities, but are scripted and edited to do so in the best light possible.
Both of those are standard practice and acceptable. What Google did was just wrong. They deserve to face backlash for this.
And iirc Tesla is also being investigated for fraudulent claims for faking the safety of their self driving cars.
You do understand the concept of reputation, right?
If lying helps, which can happen if there aren't large legal costs or social repercussions on brand equity, or if the lie goes undetected, then they'll lie. This is what we necessarily get from the upstream incentives. Fortunately, lying in a marketing video is fairly low on the list of ethical violations that have happened in the recent past.
We've effectively got a governance alignment problem that we've been trying to solve with regulations, taxes and social norms. How can you structure guardrails in the form of an incentive system to align companies with ethical outcomes? That's the question and it's a difficult one. This question also applies to any form of human organization, including governments.
Doing all these hype videos just for the sake of satisfying shareholders or whatever is just making me loose trust in their research division. I don't think they did anything like this when they released Bert.
Given the massive data volume in videos, I assumed it processed video into pictures by extracting a frame per second or something along those lines, while still taking the entire video as the initial input.
Turns out, it wasn't even doing that!
My friend, all these large corporations are going to get away with exactly as much as they can, for as long as they can. You're implying there's nothing to do but wait until they grace us with a "not evil phase", when in reality we need to be working on restoring our anti-monopoly regulation that was systematically torn down over the last 30 years.
If I demoed swype texting as it functions in my day to day life to someone used to a querty keyboard they would never adopt it
The rate at which it makes wrong assumptions about the word, or I have to fix it is probably 10% to 20% of the time
However because it’s so easy to fix this is not an issue and it doesn’t slow me down at all. So within the context of the different types of text Systems out there, I t’s the best thing going for me personally, but it takes some time to learn how to use it.
This is every product.
If you demonstrated to people how something will actually work after 100 hours of habituation and compensation for edge cases, nobody would ever adopt anything.
I’m not sure how to solve this because both are bad.
(Edit: I’m keeping all my typos as meta-comment on this given that I’m posting via swype on my phone :))
Unfortunately iOS text editing is also completely worthless. It forces strange selections and inserts edited text in awkward ways.
I’m a QWERTY texter but text entry on iOS is a complete disaster that has only gotten worse over time.
My only issue is that no keyboard implementation really supports more than two languages which makes me switch back to plain qwerty with autocomplete all the time.
It's even worse on the watch somehow. I take care to hit every key exactly, the correct word is there, I hit space, boom replaced with a completely different word. On the watch it seems to replace almost every word with bullshit, not just technical terms.
Sort of related, it also doesn't let you cuss. It will insist on replacing fuck with pretty much anything else. I had to add fuck to the custom replacement dictionary so it would let me be. What language I choose to use is mine and mine alone, I don't want Nanny to clean it up.
No, that is how you're told to type. You have to be told to type that way precisely because QWERTY is not designed to keep your fingers on the home row. If you type in a layout that is designed to do that, you don't need to be told to keep your fingers on the home row, because you naturally will.
Nobody really knows what the designers were thinking, which I do not mean as sarcasm, I mean it straight. History lost that information. But whatever they were thinking that is clearly not it because it is plainly obvious just by looking at it how bad it is at that. Nobody trying to design a layout for "keeping your fingers on the home row" would leave hjkl(semicolon) under the resting position of the dominant hand for ~90% of the people.
This, perhaps in one of technical history's great ironies, makes it a fairly good keyboard for swype-like technologies! A keyboard layout like Dvorak that has "aoeui" all right next to each other and "dhtns" on the other would be constantly having trouble figuring out which one you meant between "hat" and "ten" to name just one example. "uio" on qwerty could probably stand a bit more separation, but "a" and "e" are generally far enough apart that at least for me they don't end up confused, and pushing the most common consonants towards the outer part of the keyboard rather than clustering them next to each other in the center (on the home row) helps them be distinguishable too. "fghjkl" is almost a probability dead zone, and the "asd" on the left are generally reasonably distinct even if you kinda miss one of them badly.
I don't know what an optimal swype keyboard would be, and there's probably still a good 10% gain to be made if someone tried to make one, but it wouldn't be enough to justify learning a new layout.
My understanding of QWERTY layout is that it was designed so that characters frequently used in succession should not be able to be typed in rapid succession, so that typewriter hammers had less chance of colliding. Or is this an urban myth?
People don't hunt and peck after years of keyboard use because of the keyboard; they do it because of the keyboard layout.
If you want to prove I'm wrong, go learn Dvorak or Colemak and show me that once you're comfortable you still hunt and peck. You won't be, because it wouldn't even make sense. Or, less effort, find a hunt & peck Dvorak or Colemak user who is definitely at the "comfortable" phase.
I understand how dvorak is designed. I am still not convinced people will be using all their fingers especially their pinkys in a consistent manner that without learning that this is what you should work towards.
I did it. It happened. Theorizing about why it didn't happen is not terribly productive.
The design was to spread out the hammers of the most frequently used letters to reduce the frequency of hammer jamming back when people actually used typewriters and not computers.
The problem it attempted to improve upon, and which is was pretty effective at, is just a problem that no longer exists.
However, I use a Dvorak layout and my hands feel like they alternate better on that due to the vowels being all on one hand. The letters are also in more sensical locations, at least for English writing.
It can get annoying when G and C are next to each other, and M and W, but most of the time I type faster on Dvorak than I ever did on Qwerty. It helps that I learned during a time where I used qwerty at work and Dvorak at home, so the mental switch only takes a few seconds now.
And it does a bad job at it, which is further evidence that it was not the design consideration. People may not have been able to run a quick perl script over a few gigabytes of English text, but they would have gotten much closer if that was the desire. I don't believe that was their goal but they were just too stupid to get it even close to right.
That's a folk myth that's mostly debunked.
https://www.smithsonianmag.com/arts-culture/fact-of-fiction-...
After all, with TVs I've had the same experience as you with the annoying alphabetical keyboard, but we type into they maybe a couple of times a year, or maybe once in 5 years, whereas if we changed our phone keyboard layout we'd likely get used to it quite quickly.
Even if not going so far as to push it as a new default for all users (I'm willing to accept the possibility that I'm speaking for myself as the kind of geeky person who wouldn't mind the initial inconvenience of a new kb layout if it meant saving time in the long run, and that maybe a large majority of people would just hate it too much to be willing to give it a chance), they could at least figure out what the best layout is (maybe this has been studied and decided already, by somebody?) and offer that as an option for us geeks.
This is the sort of demo that 1) gives people a misleading idea of what the product can actually do; and 2) ultimately contributes to the inevitable cynical backlash.
If the product is really great, people can see it in a realistic demo of its capabilities.
I would agree completely that it’s not ready for consumers the way it was displayed, which is my point.
I do want to add that I believe that the right way to do these types of new product rollout is not with these giant public announcements.
In fact, I think generally speaking the “right” way to do something like this demonstrates only things that are possible robustly. However that’s not the market that Google lives in. They’re capitalists trying to make as much money as possible. I’m simply evaluating that what they’re showing I think is absolutely technically possible and I think Google can deliver it even if its not ready today.
Do I think it’s supremely ethical the way that they did it? No I don’t.
And it's dangerous to assume they can just "deliver later". It's not that simple. If it is why not bake it in right now instead of committing fraud?
This is damaging to companies that walk the walk and then people have literally said to me "but what about that Gemini"? and dismiss our work.
That was basically what magic leap did to the whole AR development market. Everyone deep in it knew they couldn’t do it but they messed up so badly that it basically killed the entire industry
But that's a different issue than LLM hallucinations.
With Swype, you already know what the correct output looks like. If the output doesn't match what you wanted, you immediately understand and fix it.
When you ask an LLM a question, you don't necessarily know the right answer. If the output looks confident enough, people take it as the truth. Outside of experimenting and testing, people aren't using LLMs to ask questions for which they already know the correct answer.
It is the main reason that handwriting recognition did not displace keyboards. Once the handwriting is converted to text, it’s easier to fix errors with a pointer and keyboard. So after a few rounds of this most people start thinking: might as well just start with the pointer and keyboard and save some time.
So the question is, how easy is it to detect and correct errors in generative AI output? And the unfortunate answer is that unless you already know the answer you’re asking for, it can be very difficult to pick out the errors.
Yeah the feedback loop with consumers has a higher likelihood of being detrimental, so even if the iteration rate is high, it’s potentially high cost at each step.
I think the current trend is to nerf the models or otherwise put bumpers on them so people can’t hurt themselves. That’s one approach that is brittle at best and someone with more risk tolerance (OpenAI) will exploit that risk gap.
It’s a contradiction then at best and depending on the level of unearned trust from the misleading marketing, will certainly lead to some really odd externalities
Think “man follows google maps directions into pond” but for vastly more things.
I really hated marketing before but yeah this really proves the warning I make in the AI addendum to my scarcity theory (in my bio).
Maybe you can spice up a demo, but misleading to the point of implying things are generated when they're not (like the audio example) is pretty bad.
Except actual good ones, like ChatGPT or Gmail (by their time).
In your Swype analogy, it would be as if Swype works by having to write out on a piece of paper the general goal of what you're trying to convey, then having to write each individual letter on a Post-it, only for you to then organize these Post-its in the correct order yourself.
This process would then be translated into a slick promo video of someone swiping away on their keyboard.
This is not a matter of “eh, it doesn't 100% work as smooth as advertised.”
0: https://techcrunch.com/2023/12/07/googles-best-gemini-demo-w...
[1] https://www.bloomberg.com/opinion/articles/2023-12-07/google...
[2] https://www.bloomberg.com/opinion/articles/2023-12-07/google...
I am curious, if the voice, email, chat, and shortly video can all be entirely generated in real or near real time, how can we be sure that remote employee is actually not a full or partially generated entity?
Shared secrets are great when verifying but when the bodies are fully remote - what is the solution?
I am traveling at the moment. How can my family validate that it is ME claiming lost luggage and requesting a Venmo request?
PGP
(I say this in jest, as a PGP user)
Hopefully realtime interaction will be part of an app soon. Doesn’t seem like there would be too many technical hurdles there.
A local model could send relevant still images from the camera feed to Gemini, along with the text transcript of the user’s speech. Then Gemini’s output could be read aloud with text-to-speech. Seems doable within the present cost and performance constraints.
:%s/Google/the team
:%s/people/the promotion board
Conway's law applied to the corporate-public interface :)I'm not saying there have been no improvements in AI. There is and this includes Google. But the reason why ChatGPT has really taken over the world is that the demo is in your own hands and it does quite well there.
Thinking back to the firm's early days, it strikes me that some HN users and perhaps even some Googlers have no memory of a time before Google Maps and simply can't imagine how disruptive and innovative things like that were at the time. Being able to browse satellite imagery for the whole world was something previously confined to the upper echelons of the military-industrial complex.
That's one reason I wish the firm (along with several other tech giants) were broken up; it's full of talented innovative people, but the advertising economics at the core of their business model warp everything else.
That's different from "Gemini was shown selected still images and not video".
They do disclose most of the details elsewhere, but the video itself is produced and edited in such a way that it's extremely misleading. They really want you to think that it's responding in complex ways to simple voice prompts and a video feed, and it's just not.
The video fooled many people, including myself. This was not your typical super optimized and scripted demo.
This was blatant false advertising. Showing capabilities that do not exist. It’s shameful behavior from Google, to be perfectly honest.
In my estimation, given the context around AI-generated content and general fakery, this video was deceptive. The only impressive thing about the video (to me) was how snappy and fluid it seemed to be, presumably processing video in real time. None of that was real. It's borderline fraudulent.
The video has a disclaimer that it was edited for latency.
And good speech-to-text and text-to-speech already exists, so building that part is trivial. There's no deception.
So then it seems like somebody is pressing a button to submit stills from a video feed, rather than live video. It's still just as useful.
My main question then is about the cup game, because that absolutely requires video. Does that mean the model takes short video inputs as well? I'm assuming so, and that it generates audio outputs for the music sections as well. If those things are not real, then I think there's a problem here. The Bloomberg article doesn't mention those, though.
> The video has a disclaimer that it was edited for latency.
There was no disclaimer that the prompts were different from what's shown.
> And good speech-to-text and text-to-speech already exists, so building that part is trivial. There's no deception.
Look at how many people thought it can react to voice in real-time - the net result is that a lot of people (maybe most?) were deceived. And the text prompts were actually longer and more specific than what was said in the video!
> somebody is pressing a button to submit stills from a video feed, rather than live video.
Somebody hand-picked images to convey exactly the right amount of information to Gemini.
> Does that mean the model takes short video inputs as well? I'm assuming so
It was given a hand-picked series of still images with the hands still on the cups so that it was easier to understand what cup moved where.
Source for the above: https://developers.googleblog.com/2023/12/how-its-made-gemin...
> My main question then is about the cup game, because that absolutely requires video.
They did it with carefully timed images, and provided a few examples first.
> I'm assuming so, and that it generates audio outputs for the music sections as well
No, it was given the ability to search for music and so it was just generating search terms.
Here's more details:
https://developers.googleblog.com/2023/12/how-its-made-gemin...
But the most impressive part of the demo, was the way the LLM just seemed to know when to jump in with a response. It appeared to be able to wait until the user had finished the drawing, or even jumping in slightly before the drawing finished. At one point the LLM was halfway though a response and then saw the user was now colouring the duck in blue, and started talking about how the duck appearing to be blue.
The LLM also appeared to know when a response wasn't needed because the user was just agreeing with the LLM.
I'm not sure how many people noticed that on a conscious level, but I positive everyone noticed it subconsciously, and felt the interaction was much more natural.
As you said, good speed-to-text and speech-to-text has already been done, along with multi-model image/video/audio LLMs and image/music generation. The only novel thing google appeared to be demonstrating and what was most impressive was this apparent natural interaction. But that part was all fake.
...OpenAI solved this by generating LLM text for you to wade through?
LLMs have been very useful, but regular search is still a big part of everyday life for me.
Today, LLMs excel in providing concise responses, addressing simple, curious questions like, "Do all bees live in colonies?"
For technical questions, ChatGPT has almost completely replaced Google & Stack Overflow for me.
Though because you don’t see the answers it doesn’t show you, it’s hard to really validate the quality, so I’m still wary, but when I look for specific stuff it tends to find it.
Turns out it was all staged.
Lost a lot of trust after that.
Google is stuck in innovators dilemma.
They make 300B of revenue which ~90% is ads revenue.
Their actual mission that management chain optimizes for is their $ growth.
A superior AI model that gives the user exactly what they want would crash their market cap.
Microsoft has tons of products with Billion+ profit, Google has only a handful and other than cloud they all tie to Ads.
Google is addicted to ads. If chrome adds a feature that decreases ad revenue, that team gets a stick.
Nothing at Google should jeopardize their ad revenue.
AI is directly a threat to Google’s core business model - ads. It’s obvious they’re gonna half ass it.
For OpenAI, AI is existential for them. If they don’t deliver, they’ll die.
It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctness is important. Even when the model has the knowledge it would need to answer correctly, as in this case, it will still lie.
The false statement is after it says the duck floats, it continues "It is made of a material that is less dense than water." This is false; "rubber" ducks are made of vinyl polymers which are more dense than water. It floats because the hollow shape contains air, of course.
If I pushed the model further on the composition of a rubber duck, and it failed to mention its construction, then it’d be lying.
However there is this disgusting part of language where a statement can be misleading, technically true, not the whole truth, missing caveats etc.
Very challenging problem. Obviously Google decided to mislead the audience and basically cover up the shortcomings. Terrible behaviour.
What does matter, however, is the overall density of the space that was water and became displaced by the canoe. That space can be populated with dense water, or with a less dense canoe+air (or canoe+vacuum) combination. That's what a rubber duck also does: the duck+air (or duck+vacuum) combination is less dense than the displaced water.
Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.
Maybe I've only ever seen terrible doctors but I always cross reference what doctors say with reputable sources like WebMD (which I understand likely contain errors). Sometimes I'll go straight to WebMD.
This isn't a knock on doctors - they're humans and prone to errors. Lawyers, engineers, product managers, teachers too.
What I'm saying is that the tolerance for mistakes is strongly correlated to the value ChatGPT creates. I think both will need to be improved but there's probably more opportunity in creating higher value.
I don't have a horse in the race.
You don't?
https://fortune.com/2023/06/23/lawyers-fined-filing-chatgpt-...
I generally agree with you, but it's funny that you use this as an example when it already happened. https://arstechnica.com/tech-policy/2023/06/lawyers-have-rea...
I really don’t recommend using ChatGPT (even GPT-4) for legal research or analysis. It’s simply terrible at it if you’re examining anything remotely novel. I suspect there is a valuable RAG application to be built for searching and summarizing case law, but the “reasoning” ability and stored knowledge of these models is worse than useless.
Oh dear.
For mainstream stuff on the other hand ChatGPT is great. And I'm sure that Gemini will be even better.
"Flat out wrong" implies determinism. For answers which are deterministic such as "syntax checking" and "correctness of code" - this already happens.
ChatGPT, for example, will write and execute code. If the code has an error or returns the wrong result it will try a different approach. This is in production today (I use the paid version).
With more mainstream libraries it's pretty good though
If I ever worried about being quoted then I’ll verify the information
otherwise I’m conversational, have taken an abstract idea into a concrete one and can build on top of it
But I’m quickly migrating over to mistral and if that starts going off the rails I get an answer from chatgpt4 instead
With a chatbot that's largely impossible, or at least impractical. I don't know where it's getting anything from - maybe it trained on a shitty Reddit post that's 100% wrong, but I have no way to tell.
There has been some work (see: Bard, Bing) where the LLM attempts to cite its sources, but even then that's of limited use. If I get a paragraph of text as an answer, is the expectation really that I crawl through each substring to determine their individual provenances and trustworthiness?
The shape of a product matters. Google as a linker introduces the ability to adapt to imperfect information quality, whereas a chatbot does not.
As an exemplar of this point - I don't trust when Google simply pulls answers from other sites and shows it in-line in the search results. I don't know if I should trust the source! At least there I can find out the source from a single click - with a chatbot that's largely impossible.
It's a computer. That's why. Change the concept slightly: would you use a calculator if you had to wonder if the answer was correct or maybe it just made it up? Most people feel the same way about any computer based anything. I personally feel these inaccuracies/hallucinations/whatevs are only allowing them to be one rung up from practical jokes. Like I honestly feel the devs are fucking with us.
I don’t necessarily disagree with your interpretation, but there’s a revealed preference thing going on.
The number of non-tech ppl I’ve heard directly reference ChatGPT now is absolutely shocking.
The problem is that a lot of those people will take ChatGPT output at face value. They are wholly unaware that of its inaccuracies or that it hallucinates. I've seen it too many times in the relatively short amount of time that ChatGPT has been around.
To avoid the darker topics to keep the conversation on the rails, if there were a misinformation campaign that was trying to state that the Earth's sky is red, then the fine tuning should be able to inform that this is clearly fake so when quoting this it should be stated as incorrect information that is out there. This kind of development should be how we can clean up the fake, but nope, we're seemingly quite happy at accepting it. At least that's how your question comes off to me.
So the labels thing is something that obviously will never work. But the system has all of the information it needs to know if the question is definitively answerable. If it is not, do not phrase the response definitively. At this point, I'd be happy if it responded to "Is 1+1 = 2?" with a wish washy answer like, "Most people would agree that 1+1 = 2", and if it wanted to say "in base 10, that is the correct answer. however, in base 2, the 1+1 = 10" would also be acceptable. Fake it till you make it is not the solution here.
You can kill people with a fork, it doesn't mean you should legally be allowed to own a nuclear bomb "because it's just the same". The problem always come from scale and accessibility
How to enforce it is the real sticky wicket though, so it's only something best discussed at places like this or while sitting around chatting while consuming
let me show you this "genius"/"wrong-thinking" person as to say about AL(artificial life) and deterministic computing.
https://www.cs.unm.edu/~ackley/
https://www.youtube.com/user/DaveAckley
To sum up a bunch of their content: You can make intractable problems solvable/crunchable if you allow just a little error into the result (which is reduced the longer the calculation calculates). And this is acceptable for a number of use cases where initial accuracy is less important that instant feedback.
It is radically different from a Von Neumann model of a computer - where there is a deterministic 'totalitarian finger pointer' pointing to some registry (and only one registry at a time) is an inherently limited factor. In this model - each computational resource (a unit of ram, and a processing unit) fights for and coordinates reality with it's neighbors without any central coordination.
Really interesting stuff. still in its infancy...
The problem is a matter of degree. These models are substantially less reliable than humans and far below the threshold of acceptability in most tasks.
Also, it seems to me that AI can and will surpass the reliability of humans by a lot. Probably not by simply scaling up further or by clever prompting, although those will help, but by new architectures and training techniques. Gemini represents no progress in that direction as far as I can see.
Of course, I agree that if we want computers to “think on their own“ or otherwise “be more human“ (whatever that means) we should expect a downgrade in correctness, because humans are wrong all the time.
Computer engineers maybe. I think the general population is quite tolerant of mistakes as long as the general value is high.
People generally assign very high value to things computers do. To test this hypothesis all you have to do is ask folks to go a few days without their computer or phone.
The opposite. Far too tolerant of the excuse "sorry, computer mistake." (But yeah, just at the same time as "the computer says so".)
With the rush of investment in dollars and to use these in places like healthcare, government, security, etc. there should be absolute precision.
First, we know they are imperfect. People seem to put more faith into machines, though I do sometimes see people being too trusting of other people.
Second, we have methods for measuring their imperfection. Many people develop ways to tell when someone is answering with false or unjustified confidence, at least in fields they spend significant time in. Talk to a scientist about cutting edge science and you'll get a lot of 'the data shows', 'this indicates', or 'current theories suggest'.
Third, we have methods to handle false information that causes harm. Not always perfect methods, but there are systems of remedies available when experts get things wrong, and these even include some level of judging reasonable errors from unreasonable errors. When a machine gets it wrong, who do we blame?
At least we won’t have to worry about it obtaining god-like powers over our society…
We all know someone who's better at self promotion than at whatever they're supposed to be doing. Those people often get far more power than they should have, or can handle—and ChatGPT is those people distilled.
How much inaccuraciy would that be ?
Just like the hype with AI and the billions of dollars going into it. There’s something there but it’s a big fat unknown right now whether any part of the investment will actually pay off - everyone needs it to work to justify any amount of the growth of the tech industry right now. When everyone needs a thing to work, it starts to really lose the fundamentals of being an actual product. I’m not saying it’s not useful, but is it as useful as the valuations and investments need it to be? Time will tell.
This insistence on comparing machines to humans to excuse the machine is as tiring as it is fallacious.
As others hinted at, there's some bias because it's coming from a computer, but I think it's far more nuanced than that.
I've worked with many experts and professionals through my career ranging across medicine, various types of engineers, scientists, academics, researchers and so on and the pattern I often see is the level of certainty presented that always bothers me and the same is often embedded in LLM responses.
While humans don't typically quantify the certainty of their statements, the best SMEs I've ever worked with make it very clear what level of certainty they have when making professional statements. The SMEs who seem to be more often wrong than not speak in certainty quite often (some of this is due to cultural pressures and expectations surrounding being an "expert").
In this case, I would expect a seasoned scientist to say something in response to the duck question that: "many rubber ducks exist and are designed to float, this one very well might, we'd really need to test it or have far more information about the composition of the duck, the design, the medium we want it in (Water? Mecury? Helium?)" and so on. It's not an exact answer but you understand there's uncertainty there and we need to better clarify our question and the information surrounding that question. The fact is, it's really complex to know if it'll float or not from visual information alone.
It could have an osmimum ball inside that overcomes most the assumed buoyancy the material contains, including the air demonstrated to make it squeak. It's not transparent. You don't know for sure and the easiest way to alleviate uncertainty in this case is simply to test it.
There's so much uncertainty in the world, around what seem like the most certain and obvious things. LLMs seem to have grabbed some of this bad behavior from human language and culture where projecting confidence is often better (for humans) than being correct.
Deception isn't always outright lying. This video was deceitful in form and content and presentation. Their product can't do what they're implying it can, and it was put together specifically to mislead people into thinking it was comparable in capabilities to gpt-4v and other competitor's tech.
Working for Google AI has to be infuriating. They're doing some of the most cutting edge research with some of the best and brightest minds in the field, but their shitty middle management and marketing people are doing things that undermine their credibility and make them look like untrustworthy fools. They're a year or more behind OpenAI and Anthropic, barely competitive with Meta, and they've spent billions of dollars more than any other two companies, with a trashcan fire for a tech demo.
It remains to be seen whether they can even outperform Mistral 7b or some of the smaller open source models, or if their benchmark numbers are all marketing hype.
LLMs do none of this. They pose as a confident expert on almost everything, and are just as likely to spit out BS as a true answer. They don't cite their sources, and if you ask for the source sometimes they provide ones that don't contain the information cited or don't even exist. If you hired a researcher and they did that you wouldn't hire them again.
If a tool is broken, you seek to fix it. You don't just say "ah yeah it's a broken tool, but it's better than nothing!"
All these LLM releases are amazing pieces of technology and the progress lately is incredible. But don't rag on people critiquing it, how else will it get better? Certainly not by accepting its failings and overlooking them.
Overpowered LLMs like GPT-4 are both broken (according to how you are defining it) and useful -- they're just not the idealized version of the tool.
I don't know why "critiquing the tool" is being equated to "refusing to use the tool."
I don't like calling something a strawman, because I think it's an overused argument, but...I mean...
My point is that the attempt to critique it was a failure. It provided no critique.
It was incomplete at the very least -- it assigned it the label of broken, but didn't explain the implications of that. It didn't define at what level of failure it would need to be to valuable.
Additionally, I didn't indicate whether or not he would refuse to use it -- specifically because I didn't know, because he didn't say.
We all use broken tools built on a fragile foundation of imperfect precision.
If I don't use it, then the tool is not used and provided no benefit...
Your CPU, right now, has known defects. It will produce the wrong outputs for some inputs. It seems to meet your definition of broken.
Do you agree with that premise ?
If we do somehow try to apply your analogy, it would indicate that the LLM output is flawed in a way we cannot scrutinize -- the hidden failures that we aren't detecting (why? It's not specified, I am assuming because we didn't check to see if the tool was "broken" and not meeting some unspecified quality-level; that is, it's an unknown unknown failure mode).
This doesn't really comport with the LLM scenario, where the output is fully viewed, and the outputs are widely understood (that is it is a known failure mode).
This is more closely related to a computing service -- of which you are an active user of. You are using a "broken" computer right now, according to your definition of broken correct ?
Yes, it does. People are regularly using LLMs to brief themselves on topics they are ignorant about. Are you actually serious?
>This is more closely related to a computing service -- of which you are an active user of. You are using a "broken" computer right now, according to your definition of broken correct ?
Doesn't really matter, because I'm not relying upon the computer for the verity of its functions
Is a drug “broken” because it only cures a disease 80% of the time?
The framing most critics seem to have is “it must be perfect”.
It’s ok though, their negativity just means they’ll miss out on using a transformative technology. No skin off the rest of us.
I think the folks making these tools tend to oversell their capabilities because they want us to imagine the applications we can come up with for them. They aren’t selling the tool, they are selling the ability to make tools based on their platform, which means they need to be speculative about the types of things their platform might enable.
They might have been trained/prompted with misinformation, but then it's the people doing the training/prompting who are lying, still not the LLM.
I think this is a different usage of the word, and we're pretty used to making the distinction, but it gets confusing with LLMs.
The model consistently and purposefully withheld knowledge it was directly aware of. This is lying under any useful definition of the word. You're veering off into meaningless philosophy that has no bearing on outcomes and results.
Perhaps some humans never lie, but should the LLM be trained only on that tiny slice of people? It's part of life, even non-human life! Evolution works based on things lying: natural camouflage, for example. Do octopuses and chameleons "lie" when they change color to fake out predators? They have intent to deceive!
The ones that do are people I do my best to avoid interacting with.
LLMs act more like the latter, than the former.
LLMs right now are most practical for generating templated text and images, which when paired with an experienced worker, can make them orders of magnitude more productive.
Oh, DALL-E created graphic images with a person with 6 fingers? How long would it have taken a pro graphic artist to come up with all the same detail but with perfect fingers? Nothing there they couldn't fix in a few minutes and then SHIP.
If by ship, you mean put directly into the public domain then yes.
https://www.goodwinlaw.com/en/insights/publications/2023/08/...
and for more interesting takes: https://www.youtube.com/watch?v=5WXvfeTPujU&
I suppose there’s two possible solutions: one is a new training or inference architecture that somehow understand “facts”. I’m not an expert so I’m not sure how that would work, but from what I understand about how a model generates text, “truth” can’t really be a element in the training or inference that affects the output.
the second would be a technology built on top of the inference to check correctness, some sort of complex RAG. Again not sure how that would work in a real world way.
I say it might be fundamental to how the model works because as someone pointed out below, the meaning of the word “material” could be interpreted as the air inside the duck. The model’s answer was correct in a human sort of way, or to be more specific in a way that is consistent with how a model actually produces an answer- it outputs in the context of the input. If you asked it if PVC is heavier than water it would answer correctly.
Because language itself is inherently ambiguous and the model doesn’t actually understand anything about the world, it might turn out that there’s no universal way for a model to know what’s true or not.
I could also see a version of a model that is “locked down” but can verify the correctness of its statements, but in a way that limits its capabilities.
Is there some sense in which this isn't obvious to the point of triviality? I keep getting confused because other people seem to keep being surprised that LLMs don't have correctness as a property. Even the most cursory understanding of what they're doing understands that it is, fundamentally, predicting words from other words. I am also capable of predicting words from other words, so I can guess how well that works. It doesn't seem to include correctness even as a concept.
Right? I am actually genuinely confused by this. How is that people think it could be correct in a systematic way?
The issue is that “correctness” isn’t a differentiable concept. So there’s no gradient to descend. In general, there’s no way to say that a sentence is more or less correct. Some things are just wrong. If I say that human blood is orange that’s not more incorrect than saying it’s purple.
This is maybe a pedantic "yes", but is also extremely relevant to the outstanding performance we see in tasks like programming. The issue is primarily the size of the correct output space (that is, the output space we are trying to model) and how that relates to the number of parameters. Basically, there is a fixed upper bound on the amount of complexity that can be encoded by a given number of parameters (obvious in principle, but we're starting to get some theory about how this works). Simple systems or rather systems with simple rules may be below that upper bound, and correctness is achievable. For more complex systems (relative to parameters) it will still learn an approximation, but error is guaranteed.
I am speculating now, but I seriously suspect the size of the space of not only one or more human language but also every fact that we would want to encode into one of these models is far too big a space for correctness to ever be possible without RAG. At least without some massive pooling of compute, which long term may not be out of the question but likely never intended for individual use.
If you're interested, I highly recommend checking out some of the recent work around monosemanticity for what fleshing out the relationship between model-size and complexity looks like in the near term.
Modern machine learning models contain a lot of inscrutable inner layers, with far too many billions of parameters for any human to comprehend, so we can only speculate about what's going on. A lot of people think that, in order to be so good at generating text, there must be a bunch of understanding of the world in those inner layers.
If a model can write convincingly about a soccer game, producing output that's consistent with the rules, the normal flow of the game and the passage of time - to a lot of people, that implies the inner layers 'understand' soccer.
And anyone who noodled around with the text prediction models of a few decades ago, like Markov chains, Bayesian text processing, sentiment detection and things like that can see that LLMs are massively, massively better than the output from the traditional ways of predicting the next word.
So if the chatbot is used to talking, knows what you'd expect, and listens to your feedback, why wouldn't it also want to tell the truth like you would instinctively, even best effort only ?
Sadly, the chatbots doesn't yet really care about the game it's playing, it doesn't want to make it interesting, it's just like a slave producing minimal low-effort outputs. I've talked to people exploited for money in dark places, and when they "seduce" you, they talk like a chatbot: most of it is lie, it just has to convince you a little bit to go their way, they pretend to understand or care about what you say, but end of the day, the goal is for you to pay. Like the chatbot.
That's a very interesting point, both technically and philosophically.
Where Gemini is "multi-modal" from training, how close do you think that gets? Do we know enough about neurology to identical a native language in which we think? (not rhetorical questions, I'm really wondering)
We don’t use neural networks because they’re similar to brains. We use them because they are arbitrary function approximators and we have an efficient algorithm (backprop) coupled with hardware (GPUs) to optimize them quickly.
> And frankly I don’t think it matters that humans think non verbally
You're right, that's not the reason non-verbal is better, but it is evidence that non-verbal is probably better. I think the reason it's better is that language is extremely lossy and ambiguous, which makes a poor medium for reasoning and precise thinking. It would clearly be better to think without having to translate to language and back all the time.
Imagine you had to solve a complicated multi-step physics problem, but after every step of the solution process your short term memory was wiped and you had to read your entire notes so far as if they were someone else's before you could attempt the next step, like the guy from Memento. That's what I imagine being an LLM using CoT is like.
Insofar as we can say that models think at all between the input and the stream of tokens output, they do it nonverbally. Forcing the structure of reduce some of it to verbal form short of the actual response-of-concern does not change that, just as the fact that humans reduce some of their thought to verbal form to work through problems doesn't change that human thought is mostly nonverbal.
(And if you don't consider what goes on between input and output thought, than chain of thought doesn't force all LLM thought to be verbal, because only the part that comes out in words is "thought" to start with in that case -- you are then saying that the basic architecture, not chain of thought prompting, forces all thought to be verbal.)
The models are too useful to say, "don't use them at all." Hopefully people will heed the warnings of how they can hallucinate, but further than that I'm not sure what more you can expect.
I suppose they could've deliberately found a hallucination and showcased it in the demo. In which case, pretty much every company's promo material is guilty of not showcasing negative aspects of their product. It's nothing new or unique to this case.
If you go to such technical routes your definition is wrong too. It doesn't float because it contains air. If you poke in the head of the duck it will sink. Even though at all times it contains air.
I mean apply this logic to a boat, right? Is the entire atmosphere part of the boat? Are we all on this boat as well? Is it a cruise boat? If so, where is my drink?
Not if there is vacuum outside too. In a vacuum it remains a duck and still floats.
Maybe AI correctness will be similar to automobile safety. It didn’t take long for both to be recognized as fundamental issues with new transformative technologies.
In both cases there seems to be no silver bullet. Mitigations and precautions will continue to evolve, with varying degrees of effectiveness. Public opinion and legislation will play some role.
Tragically accidents will happen and there will be a cost to pay, which so far has been much higher and more grave for transportation.
Preserving the original comment so the replies make sense:
---
I think it's a stretch to say that's false.
In a conversational human context, saying it's made of rubber implies it's a rubber shell with air inside.
It floats because it's rubber [with air] as opposed to being a ceramic figurine or painted metal.
I can imagine most non-physicist humans saying it floats because it's rubber.
By analogy, we talk about houses being "made of wood" when everybody knows they're made of plenty of other materials too. But the context is instead of brick or stone or concrete. It's not false to say a house is made of wood.
> Oh, it it's squeaking then it's definitely going to float.
> It is a rubber duck.
> It is made of a material that is less dense than water.
Full points for saying if it's squeaking then it's going to float.
Full points for saying it's a rubber duck, with the implication that rubber ducks float.
Even with all that context though, I don't see how "it is made of a material that is less dense than water" scores any points at all.
When we're grading the "correctness" of these answers, we're really just judging the average correctness of Google's training data.
Maybe the next step in making LLM's more "correct" is not to give them more training data, but to find a way to remove the bad training data from the set?
Disagree. It could easily be solid rubber. Also, it's not made of rubber, and the model didn't claim it was made of rubber either, so it's irrelevant.
> It floats because it's rubber [with air] as opposed to being a ceramic figurine or painted metal.
A ceramic figurine or painted metal in the same shape would float too. The claim that it floats because of the density of the material is false. It floats because the shape is hollow.
> It's not false to say a house is made of wood.
It's false to say a house is made of air simply because its shape contains air.
I loved it when the lawyers got busted for using a hallucinating LLM to write their briefs.
As an example, when most people say a balloon's lighter than air, they mean an inflated balloon with hot air or helium, but you catch their meaning and don't rush to correct them.
Also, lighter-than-air balloons are intentionally filled with helium and sealed; rubber ducks are not sealed and contain air only incidentally. A balloon in a vacuum would still contain helium (if strong enough) but would not rise, while a rubber duck in a vacuum would not contain air but would still easily float on a liquid of similar density to water.
The reason why it seems like a nitpick is that this is such an inconsequential thing. Yeah, it's a false statement but it doesn't really matter in this case, nobody is relying on this answer for anything important. But the point is, in cases where it does matter these models cannot be trusted. A human would realize when the context is serious and requires accuracy; these models don't.
Am I missing something?
I asked both Bard (Gemini at this point I think?) and GPT-4 why ducks float, and they both seemed accurate: they talked about the density of the material plus the increased buoyancy from air pockets and went into depth on the principles behind buoyancy. When pressed they went into the fact that "rubber"'s density varies by the process and what it was adulterated with, and if it was foamed.
I think this was a matter of the video being a brief summary rather than a falsehood. But please do point out if I'm wrong on the rubber bit, I'm genuinely interested.
I agree that hallucinations are the biggest problems with LLMs, I'm just seeing them get less commonplace and clumsy. Though, to your point, that can make them harder to detect!
Of course the ultimate skeptic would say one test doesn't prove that all rubber ducks are the same. I'm sure someone at some point in history has made a rubber duck out of material that is less dense than water. But I invite you to try it yourself and I expect you will see the same result unless your rubber duck is quite atypical.
Yes, the models will frequently give accurate answers if you ask them this question. That's kind of the point. Despite knowing that they know the answer, you still can't trust them to be correct.
I guess for me the question of whether or not the model is lying or hallucinating is if it's correctly summarizing its source material. I find very conflicting materials on the density of rubber, and most of the sources that Google surfaces claim a lower density than water. So it makes sense to me that the model would make the inference.
I'm splitting hairs though, I largely agree with your comment above and above that.
To illustrate my agreement: I like testing AIs with this kind of thing... a few months ago I asked GPT for advice as to how to restart my gas powered water heater. It told me the first step was to make sure the gas was off, then to light the pilot light. I then asked it how the pilot light was supposed to stay lit with the gas off and it backpedaled. My imagining here is that because so many instructional materials about gas powered devices emphasize to start by turning off the gas, that weighted it as the first instruction.
Interesting, the above shows progress though. I realized I asked GPT 3.5 back then, I just re-asked 3.5 and then asked 4 for the first time. 3.5 was still wrong. 4 told me to initially turn off the gas to disappate it, then to ensure gas was flowing to the pilot before sparking it.
But that said I am quite familiar with the AI being confidently wrong, so your point is taken, I only really responded because I was wondering if I was misunderstanding something quite fundamental about the question of density.
It certainly isn't how I would phrase it, and I wouldn't count air as what something is made of, but...
Soda pop is chocked full of air, it's part of it! And I'd say carbon dioxide is a part of the recipe, of pop.
So it's a confusing world for a young LLM.
(I realise it may have referenced rubber prior, but it may have meant air... again, Devil's advocate)
If you remove the air from the duck, and stop it so it won't refill, you have a flat rubber duck, which is useless for its design.
Much as flat pop is useless for its design.
And this nuance is even more nuance-ish than this devil's advocate post.
But pedantic correctness isn't even what matters here. The model made a statement where the straightforward interpretation is false and misleading. A person who didn't know better would be misled. Whether you can possibly come up with a tortured alternative interpretation that is technically not incorrect is irrelevant.
- Just after that it doesn't translate the "rubber" part
- It states there's no land nearby for it to rest or find food in the middle of the ocean: if it's a rubber duck it doesn't need to rest nor feed. (That's a missed opportunity to mention the infamous "Friendly Floatees spill"[1] in 1992 as some rubber ducks floated to that map position). Although it seems to recognize geographical features of the map, it fails to mention Easter Island is relatively nearby. And if it were recognized as a simple duck — which it described as a bird swimming in the water — it seems oblivious to the fact that the duck might feed itself in the water. It doesn't mention either that the size of the duck seems abnormally big in that map context.
- The concept of friends and foes doesn't apply to a rubber duck either. Btw labeling the duck picture as a friend and the bear picture as a foe seems arbitrary (e.g. a real duck can be very aggressive even with other ducks.)
Among other things, the astronomical riddle seems also flawed to me: it answered "The correct order is Sun, Earth, Saturn".
I'd like for it to state :
- the premises it used, like "Assuming it depicts the Sun, Saturn and the Earth" (there are other stars, other ringed-planets, and the Earth similarity seems debatable)
- the sorting criteria it used (e.g. using another sorting key like the average distance from us "Earth, Sun, Saturn" can be a correct order)
AI is in the "hold it together with hope and duct tape" phase, and marketing videos claiming otherwise are easy to spot and debunk.
While I agree, I wouldn't call proofs or concepts and demos pointless. They often illustrate a goal or target functionality you're working towards. In some cases it's really just a matter of allotting some time and resources to go from a concept to a product, no real engineering is needed, it all exists, but there's capital needed to get there.
Meanwhile some proof of concepts skip steps and show higher level function that needs some serious breakthrough work to get to, maybe multiple steps of that. Even this is useful because it illustrates a vision that may be possible so people can understand and internalize things you're trying to do or the real potential impact of something. That wasn't done here, it was embedded in a side note. That information needs to be before the demo to some degree without throwing a wet blanket on everything and needs to be in the same medium as the demo itself so it's very clear what you're seeing.
I have no problem with any of that. I have a lot of problems when people don't make it explicitly clear beforehand that it's a demo and explain earnestly what's needed. Is it really something that exists today in working systems someone just needs to invest money and wire it up without new research needed? Or is it missing some breakthroughs, how many/what are they, how long have these things been pursued, how many people are working on them... what does recent progress look like and so on (in a nice summarized fashion).
Any demo/poc should come up front with an earnest general feasibility assessment. When a breakthrough or two are needed then that should skyrocket. If it's just a lot of expensive engineering then that's also a challenge but tractable.
I've given a lot of scientific tech demonstrations over the years and the businesses behind me obviously want me to be as vague as possible to pull money in. I of course have some of those same incentives (I need to eat and pay my mortgage like everyone else). None-the-less the draw of science to me has always been pulling the veil from deception and mystery and I'm a firm believer in being as upfront as possible. If you don't lead with disclaimers, imaginations run wild into what can be done today. Adding disclaimers helps imaginations run wild about what can be done tomorrow, which I think is great.
Ultra is not yet available.
"There's no such thing as bad publicity."
Reading the comments of all these disillusioned developers, it’s already damaged them because now smart people will be extra dubious when Google starts making claims.
They just made it harder for themselves to convince developers to even try their APIs, let alone bet on them.
This was stupid.
Literally everyone I know who works at Google hates their job and are completely checked out.
If it had been real, that is.
Implying they’ve solved single token latency, however, is very distasteful.
Very misleading and sad Google would so obviously fake a demo like this. Mentioning in the description that it's edited is not really in the realm of doing enough to make clear the fakery.
mea cupla i should have looked at the bottom of the description box on youtube where it probably says "this demonstration is based on an actual interaction with an LLM"
All they've done is completely destroy my trust in anything they present.
I wish people stop posting Twitter messages to HN and provide a link directly to the original article. What's next, post on HN on an Instagram post?
And another: https://techcrunch.com/2023/12/07/googles-best-gemini-demo-w...
edit: s/stuck/stock
This is common in all industries. Take gaming, for example. Game publishers love this kind of publicity, as it creates hype, which leads to sales. There have been numerous examples of this over the years: Watch Dogs, No Man's Sky, Cyberpunk 2077, etc. There's a period of controversy once consumers realize they've been duped, the company releases some fake apology and promises or doubles down, but they still walk out of it richer, and ready to do it again next time.
It's absolutely insidious, and should be heavily fined and regulated.
But I recognize we're all different programmers in different circumstances. But at a minimum, I'd like to be honest with my work. My bosses seem to agree with me and I've never been pressured into hosting a fake demo or lie about the features.
In most cases, demos are needed because there's that dogfood problem. Its just not possible for me to know how my (prospective) customers will use my code. So I need to show off what has been coded, my progress, and my intentions for the feature set. In response, the (prospective) customer may walk away, they may have some comments that increases the odds of adoption, or they think its cool and amazing and take it on the spot. We can go back and forth with regards to feature changes or what is possible, but that's how things should work.
------------
I've done a few "I could do it like this" demos, where everyone in the room knew that I didn't finish the code yet and its just me projecting into the future of how code would work and/or how it'd be used. But everyone knew the code wasn't done yet (despite that, I've always delivered on what I've promised).
There is a degree of professional ethics I'd expect from my peers. Hosting honest demos is one of them, especially with technical audience members.
[Tech CEO read instructions to "show empathy" from his assistant via Slack]
CEO: I hear you. It must be very hard for you.
Shooter: Of course you fucking hear me, we're on the phone! Talk like a normal person!
But then I soon noticed some things that were too smooth, so seemed at best to be cherry-picked interactions occasionally leaning on hand-crafted situation handlers. Or, it turns out, faked.
Regardless of disclaimers, this video seems misleading to be releasing right now, in the context of OpenAI eating Google's lunch.
Everyone is expecting Google to try to show they can do better. This isn't that. This isn't even an mocked-up interaction future of HCI concept video, because it's not showing a vision of what people want to do --- it's only showing a demo of technical capabilities.
It's saying "This is what a contrived tech demo (not application vision concept) could look like, but we can't do it yet, so we faked it. Hopefully, the viewer will get the message that we're competitive with OpenAI."
(This fake demo could just be an isolated oops of a small group, not representative of Google's ability to rise to the current disruption challenge, I don't know.)
BUT
The fact this wasn’t realtime or with voice is not the issue. Voice to text could absolutely capture this conversation easily. And from what I’ve seen Gemini seems quicker than GPT4
Being able to respond quicker and via voice chat is not actually a big deal.
The underlying performance of the model is what we should be focussing on
Corporate tech demo exaggerates actual capabilities and smoothes over rough edges? Impossible, this is unprecedented!!
The Apple vs Google brand war is so tiresome. Let's focus on the tech.
One of the major issues in LLMs is the economics; a lot of people suspect ChatGPT loses money on every user, or at least every heavy user, because they've got a big model and A100 GPUs are expensive and in short supply.
They're kinda reluctant to have customers, with API rate limits galore, and I've heard people claiming ChatGPT has lost the performance crown having switched to a cheaper-to-run model.
If google had a model that operated on video in realtime, that would imply they've got a model that performs well, and is also very fast or that their 'TPUs' outperform the A100 quite a bit, either of which would be a big step forward.
It goes from a model being able to infer a great deal at the level of human intelligence to a model that needs to be fed essential details, and that doesn't do much inferring.
I get the feeling that many here on HN who are just shrugging it off don't realize how much of the “demo” was faked. Here’s a decent article that goes into it some more: https://techcrunch.com/2023/12/07/googles-best-gemini-demo-w...
As much as I hate it, this is absolutely fine by our society's standards. https://www.theregister.com/AMP/2014/02/14/apple_prevails_in...
The former is considered Puffery[1] and is completely legal, and the latter is straight up lying.
0: https://techcrunch.com/2023/12/07/googles-best-gemini-demo-w...
The entire launch felt like a concentrated effort to “appear” competitive to OpenAI. Google was splitting hairs talking about low single digit percentage improvement in benchmarks. Against a model that has been out for over 6 months.
I have never been so unimpressed with them. Not only has OpenAI managed to snag this one from under Google’s nose, IMO - they seem to be defending their lead quite well. Now that is something unmistakably remarkable. Color me impressed!
The more these models improve the more we will want less friction and faster interactions, this means that in the long term having to open an app and ask a question is not gonna fly compared to just pointing your phone camera to something, asking a question and getting an answer that's tailored to everything Google knows about you in real time.
Apple will most likely also roll their own in house solution for Siri instead of relying on an external company. This leaves OpenAI and the other small companies not just competing for the best models but also on how to put them in front of people in the first place and how to get access to their personal information.
I think you have too much information to form a reasonable opinion on the situation. Google is using editing techniques and specific scripting to try to demonstrate they have a sufficiently powerful general AI. The magnitude of this claim is huge, and the fact that they're faking it should be a likewise enormous scandal.
To sum this up "well I guess they're doing better than XYZ" discounts the absurd context of all this.
But right now, all the bits needed to do this already exist (just need to be assembled and -to be fair- given a LOT of polish), so it would be somewhat reasonable to think that someone had actually Put In The Work already.
The thing is that for all practical intents and purposes, if I can't try it, it doesn't exist. If they claim it exists, they should show it, a video or a few cherry-picked examples don't prove anything. It's easy to make a demo making even Eliza look like AGI by asking the right questions.
Problem showed in the video was reused in a recent competition (so could have been available in the dataset).
Gemini Ultra is barely beating ChatGPT in their manufactured benchmarks, and this is all that they got.
What this means is that those who are saying, including people at Google, that they have better models but are not releasing them in the name of AI safety, imply that, at least in the realm of LLMs, Google DeepMind had nothing all along.
Shame on them :(
The eventual removal of that tier and anything even close speaks to Google's general issues with cancelling services, but that doesn't mean it was less real while it existed.
I don't remember the "increasing forever" ever being particularly fast. I found some results from 2007 and 2012 both saying it was 4 bytes per second, <130MB per year.
So it's true that the number hasn't increased in ten years, but that last increase was +5GB all by itself. They've done a reasonable job of keeping up.
Arguably they should have kept adding a gigabyte each year, based on the intermittent boosts they were giving, but by that metric they're only about 5GB behind.
But neither AI not climate are there yet…
All companies are just yelling that they're "in" the AI/LLM game.. If they don't, share prices will drop.
"Absolutely mindblowing. The amount of understanding the model exhibits here is way way beyond anything else."
Can we change the URL to the real article if it still exists?
I wasn't at all under the impression they were showcasing TTS or low latencies as product features. I don't find the marketing misleading at all, and find these criticisms don't hit the mark.
[0] - https://www.bloomberg.com/opinion/articles/2023-12-07/google...
Brilliant idea!
I for one, welcome our new quack overlords.
Still I would distinguish between genuine optimism and cynical exaggeration. If you go back and read what was being said in and around the 70s about technology and watch the demos they gave…it always makes me feel a little sad for them. People really thought so many things were right around the corner, e.g. if a computer can beat a human star chess player, surely understanding natural language or driving a car will be easy. There were so many things being made like primitive robots that could follow a line painted on the floor, that in a controlled environment seemed so promising but the realization of the vision would take decades longer than predicted. Without their efforts we would not be where we are today.
This AI stuff is becoming super cool and the last thing we need is the usual charlatans that surround technology stuff. They should be ashamed of themselves, but I know who called for it - not the engineers, but the "management" layer and they're incapable of feeling shame.
As soon as I saw that opening paragraph from Sundar and how it was written I knew that Gemini is going to be a steaming pile of shit.
They should have watched the GPT-4 announcement from OpenAI again. That demo Greg Brockman did with converting a sketch on a piece of paper to a CodePen from a Discord channel, with all the error correcting and whatnot, is how you launch a product that's appealing to users.
TechCrunch, Twitter and some other sites (including HN i guess) are already piling on to this and by Monday things will go back to how they were and Google will have to go back to the drawing board to figure out another way to relaunch Gemini in the future.
Anyone with half a brain could see through this "demo." It was vastly too uncanny to be real, to the point that it was poorly setup. Google should be ashamed.
So what?
The fact that they did not fact check the videos again makes me not particularly confident in the quality of Google's work. The bit where the model misinterpreted music notation (the circled area does not mean "piano"), and the "less dense than water" rubber duck are beyond the pale. The SVG demo where they generate a South Park looking tree looks like a parody.