There are computer programs that do the kind of thing you're thinking about, for example, for protein structure analysis. They're incredibly complicated and generally require a lot of processing power.
https://www.frontiersin.org/articles/10.3389/fpls.2023.11283...
?
That's a simple application of machine learning algorithms you might find in scikit-learn. Here is a special issue of another alleged "predatory journal" that is full of papers on the subject
https://www.mdpi.com/journal/agronomy/special_issues/E18K759...
The OP did not suggest machine learning in generally, they suggested LLMs specifically, which time and time again have been shown to be incapable of this task as a matter of fundamental design. Worse, because LLMs can't understand the their training data, the output of an LLM must be verified, which in situations like this would probably take more time than simply conducting novel research in the form of random experimentation.
Also, you really need to read your citations. The first one found that machine learning was unsuited for the task of agricultural prediction...
Do they? You're talking the agricultural equivalent of something like "devise a new sorting algorithm with sota performance on x, y, z", not "write me some crud boilerplate".
After reading the article I'm sure there are plenty of low hanging fruits to uncover in yield optimization by trying different schedules for flooding and soil enrichment with different kinds of fertilizers. A neural network doesn't have to understand anything to point out useful statistical correlations just like it doesn't have to understand code semantics for incomplete code fragments to suggest potential completions which are then verified by the programmer/compiler/type system.
https://www.frontiersin.org/articles/10.3389/fsufs.2022.1053...
It's a big problem that ChatGPT has seduced a large number of people into thinking chatbots = AI and those people have convinced most other people that it is a scam.
I find 77,000 or so articles on "rice" in PubAg
https://search.nal.usda.gov/discovery/search?query=any,conta...
Just like many other areas, agriculture responds to knowledge and is a highly competitive international business. For instance, rice is cultivated by very different methods in Louisiana and Bangladesh and rice from either place could make it to your table.
See
https://en.wikipedia.org/wiki/System_of_Rice_Intensification
for a method which is heavy on labor input and light on fossil fuel input.
Analyzing this data set with an LLM would be a very good research project.
Research is by definition not repetitive, the text is free form and the data is never formatted in a way that makes comparison between different papers straight forward.
https://www.technologyreview.com/2022/11/18/1063487/meta-lar...
"A fundamental problem with Galactica is that it is not able to distinguish truth from falsehood, a basic requirement for a language model designed to generate scientific text. People found that it made up fake papers (sometimes attributing them to real authors), and generated wiki articles about the history of bears in space "
Having an AI shout random ideas is very easy for software people to grok, but isn't going to help. If you want AI to assist with this, you'd need to build an 'AI' that can run the real-world experiments, and that's a few orders of magnitude harder than feeding a text corpus to an LLM.
'Thinking' about this problem isn't the hard part, the hard part is doing it. Even using an LLM for something like a meta-analysis of existing research is unlikely to find many profitable avenues of exploration.
[1] Experimental research is incredibly difficult, which is a fact that's highly underappreciated by people working in abstract and theoretical disciplines.
AI interpolates across it's parameter space, but typically performs poorly in extrapolation exercises.