Artificial Intelligence, Scientific Discovery, and Product Innovation [pdf]
aidantr.github.io
aidantr.github.io
> I find that AI substantially boosts materials discovery, leading to an increase in patent filing and a rise in downstream product innovation. However, the technology is effective only when paired with sufficiently skilled scientists.
I can see the point here. Today I was exploring the possibility of some new algorithm. I asked Claude to generate some part which is well know (but there are not a lot of examples on the internet) and it hallucinated some function. In spite of being bad, it was sufficiently close to the solution that I could myself "rehallucinate it" from my side, and turn it into a creative solution. Of course, the hallucination would have been useless if I was not already an expert in the field.
Also, I think it would be faster to write my own than try to fully understand others (LLM) code. I have developed my own ways of ensuring certain aspects of the code, like security, organization, and speed. Trying to knead out how those things are addressed in code I didn't write takes me longer.
Edit; spelling
Maybe that's an argument for simpler chat modalities over shared codepads, as forcing the human to assemble bits of code provided by the LLM helps keep the human in the driver's seat.
That's how progress works. Clever people will still be clever, but maybe about slightly different things.
I see a parallel in how web search replaced other skills like finding information in physical libraries. We might not do research the old way, but we learned new tricks for the new tools. We know when to rely on them and how much, how to tell useful from garbage. We don't write by hand much, do computation in our heads much, but we type and compute more.
If a model is right 99.99% of the time (which nobody has come close to), we still need something that understands what it's doing enough to observe and catch that 0.01% where it's wrong.
Because wrong at that level is often dangerously wrong.
This is explored (in an earlier context) in the 1983 paper "Ironies of Automation".
Which means the people who catch these mistakes have to be operating at a very high level.
This means we need to resist getting lulled into a false sense of security with these systems, and we need to make sure we can still get people to a high level of experience and education.
Nobody has figured out how to get a confidence metric out of the innards of a neural net. This is why chatbots seldom say "I don't know", but, instead, hallucinate something plausible.
Most of the attempts to fix this are hacks outside the LLM. Run several copies and compare. Ask for citations and check them. Throw in more training data. Punish for wrong answers. None of those hacks work very well. The black box part is still not understood.
This is the elephant in the room of LLMs. If someone doesn't crack this soon, AI Winter #3 will begin. There's a lot of startup valuation which assumes this problem gets solved.
Not just solved, but solved soon. I think this is an extremely difficult problem to solve to the point it'd involve new aspects of computer science to even approach correctly, but we seem to just think throwing more CPU and $$$ at the problem will work itself out. I myself am skeptical.
As for "solved soon", the market can remain irrational longer than you can stay solvent. Look at Uber and Tesla, both counting on some kind of miracle to justify their market cap.
Uber seems to have become sustainable thid year.
There's little reason to expect a correction any soon on any of those.
I'm just an outside observer, though...
Just treat the hallucinations as the non-linear distortion and harmonics phenomena that come from amplification process. You can just filter the unwanted signals and noises judiciously if you're well informed.
Taking this analogy further you need to have an appropriate and proper impedance matching to maximize the accuracy, and impedance matching source or load-pull (close-loop or open-loop) and for LLM it can be in the form of RAG for example.
When i saw how Alphazero played chess back in 2017, different than other engines, that's what i described it usually, as a habit forming machine.
Or maybe in general we can say that to do something really hard and complex you must and should put a lot of effort into getting all the not-hard not-complex pieces in place, making yourself comfortable with them so they don't distract, and setting the stage for that hard part. And when you look back you'll find it odd how the hard part wasn't where you spent most of the time, and yet that's how we actually do hard stuff. Like we have to spend time knolling our code to be ready for the creative part.
What an interesting finding and not what I was expecting. Is this an issue with the UX/tooling? Could we alleviate this with an interface that still incorporates the joy of problem solving.
I haven’t seen any research that Copilot and similar tools for programmers have a similar reduction in satisfaction. Likely with how much the tools feel like an extension of traditional auto complete, and you still spend a lot of time “programming”. You haven’t abandoned your core skill.
Related: I often find myself disabling copilot when I have a fun problem I want the satisfaction of solving myself.
- Reduced creativity and ideation work (dropping from 39% to 16% of time)
- Increased focus on evaluating AI suggestions (rising to 40% of time)
- Feelings of skill underutilization
The way things seem to be going, I'd be worried management will find a way to monitor and try cut out this "security risk" in the coming months and years.
Half statement, half question… I have personally stopped using AI assistance in programming as I felt it was making my mind lazy, and I stopped learning.
So the AI is in charge, and mostly needs a bunch of lab assistants.
"Machines should think. People should work." - not a joke any more.
Biological neurons have many features like active dendritic compartmentalization that perceptrons cannot duplicate.
They are different with different advantages and limitations.
We have also known about the specification and frame problems for a long time also.
Note that part of the reason for the split between the symbolic camp and statistical camp in the 90s was due to more practical models being possible with existential quantification.
There have been several papers on HN talking about a shift to universal quantification to get around limitations lately.
Unfortunately discussions about the limits of first order logic have historical challenges and adding in the limits of fragments of first order logic like grounding are compounded upon those challenges with cognitive dissonance.
While understanding the abilities of multi level perceptrons is challenging, there is a path of realizing the implications of an individual perceptron as a choice function that is useful for me.
The same limits that have been known for decades still hold in the general case for those who can figure a way to control their own cognitive dissonance, but they are just lenses.
As an industry we need to find ways to avoid the traps of the Brouwer–Hilbert controversy and unsettled questions and opaque definitions about the nature of intelligence to fully exploit the advantages.
Hopefully experience will tempor the fear and enthusiasm for AGI that has made it challenging to discuss the power and constraints of ML.
I know that even discussing dropping the a priori assumption of LEM with my brother who has a PhD in complex analysis is challenging.
But the platonic ideals simply don't hold for non-trivial properties, and no matter if we are using ML or BoG Sat, the hard problems are too high in the polynomial hierarchy to make that assumption.
- Examine how the human-AI relationship evolved as the AI system improved during the study period
- Theorize more explicitly about which aspects of human judgment might be more vs less persistent
- Consider how their findings might change with more capable AI systems
https://pubs.acs.org/doi/10.1021/acs.chemmater.4c00643
were considered in the analysis?