The younger generations think it’s a hoax but are more uncertain because they know that any information about the event is likely to be propaganda
2,438 karma · joined July 13, 2012
https://www.producthunt.com/@jak_9994
http://jack.ventures
The younger generations think it’s a hoax but are more uncertain because they know that any information about the event is likely to be propaganda
Among my relatives I would say all are anti-US. About 5-6 of them vehemently so and want the war to start immediately.
I grew up in China and if you think there’s less propaganda in China compared to the US I don’t know what to tell you.
That said I still use my AVP regularly, it's a great home theater system for 4k HDR content that's portable for travel. I've owned other headsets like the valve index but the novelty of 3d gaming wore off pretty quickly.
I'm not certain about the specific models tested, but some VLMs just embed the image modality into a single vector, making these tasks literally impossible to solve.
An LLM could play chess though, all it needs is grounding (by feeding it the current board state) and agency (RF to reward the model for winning games)
The real problem that an LLM is trying to solve is to create a model that can enumerate all meaningful sequences of words. This is just an insane way of approaching the problem of intelligence on the face of it. There's a huge difference between a model of language and an intelligent agent that uses language to communicate.
What LLMs show is that the hardest problem - of how to get emergent capabilities at scale from huge quantities of data - is solved. To get more human-like thinking, all that is needed is to find the right pre-training task that more closely aligns with agentic behavior. This is still a huge problem but it's an engineering problem and not one of linguistic theory or philosophy.
The explanation is perfectly sensical, just too complex for humans to understand as the model scales up.
The thing you're looking for - a reductive explanation of the weights of a ANN that's easy to fit in your head, does not exist. If it were simple enough to satisfy your demands, it wouldn't work at all.
https://chat.openai.com/share/e630c5d4-d492-43cb-b7e5-214ff8...
I finished the app in 2 days, with a third day for css/visual styling. I previously might have hired someone to do this or tried to figure it out myself, which would have taken about a month.
At the end of this thread chatgpt kind of goes off the rails a bit and fails to center the uploaded image. I think this is because it can't execute the code and see its results, and can only take blind guesses at what the problem might be.
I still had to write about 10% of the code myself, but it's about 10x faster for me to use chatgpt. I think I'd use chatgpt even if it were slower, because I prefer "thinking in natural language" vs "thinking in code"
here is the final app in production: https://tinyurl.com/368w3a9y
You could start off with a random polygon and the reverse diffusion process would slowly turn it into a text glyph.
This kind of reminds me of dalle-1 where the image is represented as 256 image tokens then generated one token at a time. That approach is the most direct way to adapt a causal-LM architecture but it clearly didn't make a lot of sense because images don't have a natural top-down-left-right order.
For vector graphics, the closest analogous concept to pixel-wise convolution would be the Minkowski sum. I wonder if a Minkowski sum-based diffusion model would work for svg images.
Hallucinations are an engineering problem and can be solved. Compute per dollar is still growing exponentially. Eventually this technology will be widely proliferated and cheap to operate.
It’s trivial to come up with a prompt that doesn’t exist in the dataset. To generalize, the model cannot memorize.
The solution is to build more and more quickly.
washluv.com is still available though
> the model may encounter challenges when synthesizing intricate structures, such as human hands
I think there's two main reasons for poor hands/text
- Humans care about certain areas of the image more than others, giving high saliency to faces, hands, body shape etc and lower saliency to backgrounds and textures. Due to the way the unet is trained it cares about all areas of the image equally. This means model capacity per area is uniform, leading to capacity problems for objects with a large number of configurations that humans care more about.
- The sampling procedure implicitly assumes a uniform amount of variance over the entire image. Text glyphs basically never change, which means we should basically have infinite CFG in the parts of the image that contain text.
I'm not sure if there's any point in working on this though, since both can be fixed by simply making a bigger model.
Why would a man eat a shark? Shouldn't it be the other way around?
Which cell is the “real” one when it divides? Which branch of a tree is the true tree?
Both versions of you have equal claim to the past, as long as the copy is truly identical.
I'm not sure why the share page says "Model: Default" but it's "Model: GPT-4" in my webui. Seems like a bug in the share feature
svg editor:
early april: https://chat.openai.com/share/c235b48e-5a0e-4a89-af1c-0a3e7c...
now: https://chat.openai.com/share/e4362a56-4bc7-45dc-8d1b-5e3842...
originally it correctly inferred that I wanted a framework for svg editors, the latest version assumes I want a js framework (I tried several times) until I clarify. It also insists that the framework cannot do editable text until I nudge it in the right direction.
Overall slightly worse but the code generated is still fine.
word embeddings:
early april: https://chat.openai.com/share/f6bde43a-2fce-47dc-b23c-cc5af3...
now: https://chat.openai.com/share/25c2703e-d89d-465c-9808-4df1b3...
in the latest version it imported "from sklearn.preprocessing import normalize" without using it later. It also erroneously uses pytorch_cos_sim, which expects a pytorch tensor whereas we're putting in a numpy array.
overall I think the quality has degraded slightly, but not by enough that I would stop using it. Still miles ahead of Bard imo.
Me: I want to make an svg editor, give me some suggestions. I mainly want something with mobile support.
GPT4: gives me some options
I look over the options and choose fabricjs
Me: Start by loading an svg at a predefined url
GPT4: <code>
Me: Ok now implement a save feature, send the json to this url ...
GPT4: <code>
Me: The text is loaded as an image, I want the text to be editable
GPT4: <code>
Me: I'd like to add some google fonts to the text editor
GPT4: <code>
Me: The fonts aren't loading, I think we need to load the fonts first before initializing the canvas.
GPT4: <code>
Me: Ok add a undo/redo feature
GPT4: <code>
Me: Let's add some clickable buttons instead of hotkeys, here is the html..
GPT4: <code>
I probably could have done this myself, but frankly it would have taken me a long time to figure out the fabricjs api. It probably saved me at least a week making this thing.
here's the live app: https://tinyurl.com/2dhh58cn
and the code: https://tinyurl.com/2tu4xrtn
you can tell the GPT generated sections by the (overly) verbose comments
I would add that the above comparison is misleading, because humans have a massive advantage in that they have prior knowledge of what words mean. A more apples-to-apples comparison would have the human do next word prediction on a language they don't know.
This would be akin to me giving you a few GBs of Chinese text, with no grounding or translation, then try to communicate with you in Chinese after you've read the whole thing.
Language models do not emulate human minds - they are models of language. The emergent behavior from these models are only a side effect of their main training task, which is to build a model of all meaningful sequences of words. We then use RFHL to bias the model toward a small area of the language latent space which conforms to our idea of intelligent behavior.
Humans (a GI) have zero ability to do language modeling. Human equivalent AGI would similarly fail at this task.
The technology behind language models is more important than general intelligence - it is a universal induction engine that can model (and truly understand) the latent structure of any signal.
GPT4 is trained with PPO+RLHF. The web text that is produced by the LLM then fed back in will be more proximal to the original token distribution.
In other words, by selectively publishing LLM output you’re effectively performing the same action as clicking the thumbs up/thumbs down button on the chatgpt webui.
I agree with openai that this will not be a problem at all, since you would need a process to gauge the quality of the data anyways, even for human text.
The real solution to these problems is to train transformers on a more human-like information context rather than pure text. Hallucinations should naturally decrease as LLMs become more "agentic"
- every image on the site fits the minimalist theme (and there are hundreds). The photos are always dark in the appropriate areas so the text is readable, and they're dynamically cropped/resized with good positioning for all breakpoints.
- the front page appears to be simple, but if you click through the site there are actually a ton of dynamic illustrations and charts like this: https://www.spacex.com/human-spaceflight/
The key thing is that the aesthetic is executed consistently and with restraint. This video is a bit old now but it's as true now as it's ever been: https://www.youtube.com/watch?v=EUXnJraKM3k
Minimalism is the most difficult aesthetic to pull off, because it's easy for the end result to look unfinished, cheap or lazy.
I think the problem is not that coders can't be designers, but that coders generally don't really care about design, and I say that as a coder. If you gave the spacex project to the average coder, you'll probably end up with something like sap.com or fortinet.com
Web design has converged a lot in the last 20 years because people don't want to learn a new UI for every website they visit. This makes good web design harder, because now if you want to create something memorable and distinctive, you'll need to do it in spite of the navigation and functional parts looking identical to every other site.