Something that I have realized recently is that it has become so easy to get an answer to almost any question with the help of chatbots that its almost unnecessary to spend any effort thinking about the problem or the solution. I feel like before when I had to spend time researching a problem to find an answer I learned so many things around the topic itself which helped me understand the problem itself better and gained a deeper understanding. Today it feels like you can have an answer to the most complex questions you might have, yet you gain a superficial understanding of the topic and might forget about it quickly.
IME this works until it does not. This approach works well at the beginning of a greenfield project, but at the same time because it is so easy to add features, you will likely ship something that is way too over engineered. And that complexity will not amortize over next increments and will more likely lead to the entire project being a black box only fully understood by AI. However a more careful use of AI for targeted surgical changes is far more ”productive” in the long term IMO.
One advantage that I see in models that are implemented as code is that they can quickly and cheaply be modified using LoRAs. What would the equivalent be in hardware? Another piece of hardware you would attach like adding a graphics card to a computer?
While fundamentals are important to learn, there is also a huge benefit in learning specific tools and frameworks. There is no ”one size fits all” when it comes to software and more often than not, you need very customized solution for optimal performance. Moreover, learning to master a tool often means you are also able to improve it, which is part of the reason open-source tools usually improve with the loyal contributers!
> This ‘goal drift’ means that agents, or tasks done in a sequence with iteration, get less reliable. It ‘forgets’ where to focus, because its attention is not selective nor dynamic.
I don't know if I agree with this. The attention module is specifically designed to be selective and dynamic, otherwise it would not be much different than a word embedding (look up "soft" weights vs "hard" weights [1]).
I think deep learning should not be confused with deep RL. LLMs are autoregressive models which means that they are trained to predict the next token and that is all they do. The next token is not necessarily the most reasonable (this is why datasets are super important for better performance). Deep RL models on the other hand, seem to be excellent at agency and decision making (although in restricted environment), because they are trained to do so.
> The facilities will reportedly consume as much as 13 percent of the plant's output.
Why are AI products being shipped so aggressively despite being so inefficient? Is code autocompletion and generating random images really worth so much electricity? Shouldn’t we wait until the research has created an efficient architecture that is easily scalable first?
Isn’t that what the softmax layer is doing? The token with highest probability among all the available tokens in the model dictionary is chosen as the next token!
I think in the long run, it is in the interest of AI companies to incentivize creators to create high quality data! Not paying them their fair share will likely decrease the volume of high quality data available (or make it much less accessible). Unless these companies already have developed another architecture that can learn much more from the same dataset, the lack of new high quality data will be a problem for future larger models!
Merry Christmas everyone! Thanks everyone for you great contributions to the community and thanks to the dear moderators for the great job they are doing!
It is the first time, to my knowledge, that there is strong evidence that indicates it could actually happen! (The evidence being the rapidly changing climate and our inability to adapt quickly enough)
I believe RAG is more appropriate for this! While you can certainly fine tune on a pdf, you are essentially fine turning with batch size == 1, so you should not expect good results! Also you need a label (for example summary) in order to fine-tune!
I honestly don't think there is any incentive for large corporation to care too much about things like this and, on the contrary, they are incentivized to do more! The fine they get is basically cost of doing business and since our attention is engineered to be short, the negative PR won't last long enough to have a huge impact. Unless we see a "too big to fail" company actually fail because something like this, we won't see any change in their behvaiour!
I think this is really interesting from a meta (in-context) learning pov, but I think at some point prompt engineering stops being prompt engineering and instead becomes in-context training. This is problematic for applications such as search or completion, unless some other software/model does this automatically for you!
I don’t see the contradiction. Humans can be emotional and at the same use science to make all humans life better! In fact why would we ever develop any technology that makes life better for others if we don’t have any feelings for them?