Unfortunately I can't change the original link. Perhaps dang can.
31,843 karma · joined October 9, 2012
Socials: - x.com/dvyio - linkedin.com/in/dvyio
Interests: AI/ML, Gaming, Marketing, Programming, Science, Startups, Technology, UI/UX Design, Web Development
---
Unfortunately I can't change the original link. Perhaps dang can.
1. On-device AI
2. AI using Apple's servers
3. AI using ChatGPT/OpenAI's services (and others in the future)
Number 1 will pass to number 2 if it thinks it requires the extra processing power, but number 3 will only be invoked with explicit user permission.
[Edit: As pointed out below, other providers will be coming eventually.]
You could use something like Little Snitch (on Mac) to check if it makes any calls to their servers.
They also allow you to override the URL for the OpenAI models, so although I haven't tried, perhaps you can use local models on your own machine.
It's a fork of VS Code with some AI features sprinkled in. It writes around 80% of my code, these days.
It also has a few useful features:
- a chat interface where you can @-mention files, folders, and even documentation
- if you edit a line of code, it suggests edits around that line that are useful (e.g. you change a variable name and it will suggest updating the other uses, which you accept just by pressing Tab)
- as you're writing/editing code, it will suggest where your cursor might go next — press Tab and your cursor jumps there
https://replicate.com/philz1337x/clarity-upscaler https://replicate.com/philz1337x/multidiffusion-upscaler
However, this weekend someone released an open-source version which has a similar output. (https://replicate.com/philipp1337x/clarity-upscaler)
I'd recommend trying it. It takes a few tries to get the correct input parameters, and I've noticed anything approaching 4× scale tends to add unwanted hallucinations.
For example, I had a picture of a bear I made with Midjourney. At a scale of 2×, it looked great. At a scale of 4×, it adds bear faces into the fur. It also tends to turn human faces into completely different people if they start too small.
When it works, though, it really works. The detail it adds can be incredibly realistic.
Example bear images:
1. The original from Midjourney: https://i.imgur.com/HNlofCw.jpeg
2. Upscaled 2×: https://i.imgur.com/wvcG6j3.jpeg
3. Upscaled 4×: https://i.imgur.com/Et9Gfgj.jpeg
----------
The same person also released a lower-level version with more parameters to tinker with. (https://replicate.com/philipp1337x/multidiffusion-upscaler)
Disclosure: Professor Malan is a friend of mine, but I was a fan of CS50 long before that!
In particular, looking at the video titled "Borneo wildlife on the Kinabatangan River" (number 7 in the third group), the accurate parallax of the tree stood out to me. I'm so curious to learn how this is working.
[Direct link to the video: https://player.vimeo.com/video/913130937?h=469b1c8a45]
It uses GPT (I believe they're running their own fine-tuned version), and allows you to '@' items like your files, folders, and (most impressively) any documentation.
So, if I'm working on something that uses React Hook Form, I can paste the URL to the RHF documentation. Cursor will index it (presumably creating embeddings for each page), and then I just '@' the RHF documentation in the chat and it'll find the relevant pages and include it them the context for GPT to answer.
It's very useful. I already subscribe to ChatGPT Plus but Cursor adds something special that I'm more than happy to subscribe to Cursor too.
Then, over time, it started to feel like the people higher up were less and less interested in the site, and no amount of moderation could fix some of the issues like the amount of spam.
I had a brief moment a couple of years ago where I went back to the site, cleaned up the spam and low quality posts I could find in the top few pages in the hope that it might kick off some kind of resurrection. It didn't, mainly because hardly anyone except spammers were posting to the site at that point.
The first that comes to mind: https://chrome.google.com/webstore/detail/superpower-chatgpt...
The docs say:
> By default, images are generated at standard quality, but when using DALL·E 3 you can set quality: "hd" for enhanced detail. Square, standard quality images are the fastest to generate.
If you ask it for 4 images with an exact prompt, it'll generate 4 identical (or almost identical) images. Then if you ask it to re-run it with a different seed, it will say that it's done it but it'll still generate the same images.
I hope that's going to be changed in the future.
Last year I generated around 7,000 images using DALL·E 2 and uploaded them to https://generrated.com/
I've been re-running the same prompts using DALL·E 3, although haven't updated the site yet (although I'm planning to). So far I've created 2,000 like-for-like using those prompts.
----------
In the meantime, here are some things I've noticed with DALL·E 3 vs. DALL·E 2:
- the quality is astounding, especially illustrations (vs. photographs) — as I've been looking at the DALL·E 2 images I've constantly felt like the old images look like potatoes now we have DALL·E 3 and Midjourney (even though at the time they seemed stunning)
- you will struggle to get an output that references a specific artist, but it will sometimes offer to make images in the general style of the artist as a compromise
- it can get quite repetitive when you ask it for concepts — if you look at the 'representation of anxiety' images on Generrated, you'll see that there's a huge variety in them, but as I've been running them with DALL·E 3 it seems to prefer certain imagery (in this case, a human heart under stress appears a lot), and the 'discovery of gravity' will include a tree and an apple 80% of the time
- some of the prompts need guidance to get the output you desire — 'iconic logo symbol' works well on DALL·E 2 to create a logo, but with DALL·E 3 will often produce a general image/painting with a logo somewhere in the image (e.g. a NASA logo on an astronaut's suit rather than a logo of an astronaut)
Those are some I can remember off of the top of my head. But it's so much fun to play with!
----------
Edit: I quickly put together 3 comparisons between v2 and v3: https://imgur.com/a/L9DYCSA
The closest you can get (which is still very far) is to upload your image to GPT-4 Vision and ask it to give you a prompt to describe the image, and then put that prompt into DALL-E 3.
I guess it'll just have to be comparisons of the general concepts. It'll be good to see the change in understanding of the prompt and the change in image detail.
If anyone at OpenAI wants to give me early access to give me a head-start… smiles
But otherwise you're right — some might not work.
Because there's no way to control the seed, a direct comparison (using a before/after slider, for example) probably wouldn't make sense. But I could put the group of 4 images from each version above/below each other as a general comparison, perhaps?