Gemini 1.5 outshines GPT-4-Turbo-128K on long code prompts, HVM author
old.reddit.com
old.reddit.com
OpenAI is gradually turning into the Fisher-Price of AI. I cancelled my ChatGPT Pro subscription after 12mo due to increasingly low quality answers since ~Nov 2023. I occasionally try the API but the quality there is only slightly better than ChatGPT.
I asked Gemini v1 2 days ago about programming and the first response was some nonsense about elections being a complex topic.
if gpt 4 is Fisher Price then it's competitors are ill fitting kinder egg toys with missing pieces.
Edit: And in your screenshot you did get real responses, you see gemini generates several look at response 2, looks like what you wanted.
Not once in my thousands of questions to gpt did it go so far off base. Maybe in the gpt3 days sometimes it would get caught in a repeating character loop.
Yet my fourth or so time back to Bard and it shits the bed on the first try makes me think, yeah... Still kinder surprise
Also, the suggested changes from the remaining responses are incredibly naive. It suggests pulling numbers out and appending them back in, which breaks if the numbers aren't simply at the end of the strings, as well as not going with the pronunciation spirit of metaphone.
"Google search for election info" https://g.co/gemini/share/7afd63de0079
Custom GPTs exist for a reason. Grimoire (custom GPT) outperforms raw ChatGPT by miles.
My workflow is usually Grimoire to get a broad overview of things, then a model that can use my repo via tools like Cursor.
I built a smart contract decompiler (eveem.org), and got deep into formal verification / code analysis path during that time.
There was a ton more I could do back then if I had access to gpt-4. Some analysis methods were extremely slow using solvers, and it was a pita to write code for dedicated code pattern recognition. Gpt4 improves that, and it's easy to verify if it's guesses are correct using formal methods.
That being said, I recently read a paper whose authors made use of traditional statistical analysis to find unusual implementation patterns in an IR decompilation of smart contracts. I wonder if custom LLMs can be used to perform that analysis.
I'm very much not on the LLM hype-train, but I've been saying from day one that I'd really like something like this. Code has strong local similarities, so it's not far fetched to imagine that future LLMs -- if not current ones -- could be used as a stochastic linter of sorts.
Programming is hard, it's not insane to make use of all the help we can get.
For example, on a side project I used to explain it shortly and the issue I am facing, mostly looking to discover packages or libraries that can help me build it.
Over 3 months ago GPT4 used to jump straight to the point and recommend alternative scenarios and packages. Now it just gives me a giant text with considerations for the whole project, but rarely actually tackles the issue I am having. Even if I ask it to actually go in-depth into whatever I am interested in it usually doesn't and repeats the same surface level information.
I started feeling similar recently, but then I added the "Custom Instructions" [0] feature in ChatGPT, and the results are much better. The instruction that I've put there is that I'm a highly information-aware person and that I require detailed answers. The reason that I feel this works is that GPT is essentially trained (caveat: the RLHF part could partially falsify my "theory" here) on internet textual data whereby the usual method of answering on internet forums (and other sites) has a relatively low-to-moderate detailed way of answering things; but given that we know that GPT has in its latent space pretty much has all the useful knowledge on the internet, that if one asks for very detailed (information diverse/inclusive) answers by default that it can perform really well.
Which includes:
Conversations, Location, Feedback, Usage information
And specifically this clause:
"Please don't enter confidential information in your conversations or any data you wouldn't want a reviewer to see or Google to use to improve our products, services, and machine-learning technologies."
Update: *There is a Gemini setting here https://myactivity.google.com/u/1/product/bard but it does not explicitly state that it disables training, just activity history.
Presumably Gemini 1.5 will be available there soon.
Allows training on input:
1. Gemini (formerly Bard): https://gemini.google.com/. Policy: https://support.google.com/gemini/answer/13594961. Can opt out of training by disabling history.
2. AI Studio free plan: https://ai.google.dev/.
Does not allow training on input:
1. AI Studio paid plan: https://ai.google.dev/.
2. Google Cloud Vertex AI: https://cloud.google.com/vertex-ai. Policy: https://cloud.google.com/vertex-ai/docs/generative-ai/data-g...
3. Gemini for Workspace (formerly Duet AI for Workspace).
Also (On manual deletion):
"Even when Gemini Apps Activity is off, your conversations will be saved with your account for up to 72 hours to allow Google to provide the service and process any feedback."
And in any case:
"Conversations that have been reviewed or annotated by human reviewers (and related data like your language, device type, location info, or feedback) are not deleted when you delete your Gemini Apps activity because they are kept separately and are not connected to your Google Account. Instead, they are retained for up to three years."
The setting is explicitly about the activity history (which is 3-mo auto-delete, 72h minimum).
"Who has access to my Gemini Apps conversations?
How you can control what’s shared with reviewers
If you turn off Gemini Apps Activity, future conversations won’t be sent for human review or used to improve our generative machine-learning models."
The way I interpret the policy is that if Gemini Apps Activity is off:
1. Conversations after activity is turned off that are linked to your account are retained for up to 72 hours for "safety and security", but are not sent to human reviewers or used for model training.
2. In the event that you submit feedback (eg good/bad response), then the conversation (disassociated from your account) can be used for training. This is also flagged when attempting to submit feedback: "Even when Gemini Apps activity is off, feedback submitted will also include up to the last 24 hours of your conversations to help improve Gemini."
1. If you are an enterprise customer - the waitlist page is a blank black page as of 2/18/24
2. If you are non-enterprise customer - you can signup for waitlist.
1.5 is not available in the Gemini ChatGPT like interface. It is only available through AI Studio. But you have to signup for a waitlist.
The free tier of Gemini (formerly Bard) has been on Gemini Pro 1.0 for a bit now.
This has some serious implications for OpenAI as Gemini Pro 1.5 that is tested here indeed seems to beat their premium tier, assuming Google will keep using these tiers for what's paid and not.
Yup. If a french company, Mistral, can do it you know the cat is out of the bag (ok, ok, started by an ex- OpenAI person I think).
I tried mistral-next and it is impressive.
The "moat" at this point seems to be: "Invest and train models using $$$ of hardware and electricity".
It's both exciting (nobody has an edge) and depressing (the edge consist in buying billions worth of NVidia AI chips).
The issue is that it's only democratized if you have the money.
What's with the France slander? France is not some technological backwater - it's given us BeOS[1], VLC and Fabrice Bellard.
1. I know Be Inc was an American company, but check the nationalities of its alumni.
Say I design a language, and instead to writing an implementation, I train an AI with tons of input/output examples, the kinda of errors to emit for violations of the grammar. Will be cool to try this out.
I imagine it might be more feasible to transpilation rsther than real compilation.
It seems like a very interesting project with quite a steep learning curve, given interaction nets look like a whole new way of thinking.
At some point they run out of training data and others can catch up.
Honestly suspicious if it even shares much with 4 as it disappoints more times than it impresses.
Et viola, your application is now free from LLM lock-in.
ChatGPT 4 doesn't work as well as ChatGPT 3.5. The little brother is smarter han the big brother.
ChatGPT 3.5 works better than Eliza.
No-one has heard of any other AIs.
That's pretty much it - the entire state of AI early 2024.