On the paper “Exploring the MIT Mathematics and EECS Curriculum Using LLMs” [pdf]
people.csail.mit.edu
people.csail.mit.edu
I appreciate the clear stance that MIT has taken regarding where the responsibility lies in this situation. I think some people are missing the context that many of the authors were undergraduate/early career. Research is an iterative process and every paper has to start somewhere. I don't agree that the paper should be withdrawn because Arxiv is not technically a publication, but I also wouldn't consider the paper properly peer reviewed. Teachers own the copyright the exam material. I was taught copyrighted material can't and shouldn't be used as part of an eval dataset.
The followup by three other MIT ('24) seniors is a great peer review.
https://flower-nutria-41d.notion.site/No-GPT4-can-t-ace-MIT-...
No, GPT4 Can’t Ace MIT - https://news.ycombinator.com/item?id=36370685 - June 2023 (120 comments)
“arXiv will not consider removal for reasons such as journal similarity detection, nor failure to obtain consent from co-authors, as these do not invalidate the license applied by the submitter”
The submitter can mark the paper as “withdrawn” but it will remain available
> Iddo did not have permission from all the instructors to collect the assignment and exam questions that made up the dataset that was the subject of the paper.
Legally, you are allowed to refuse to comply with DMCA take down notices. If you do, you are increasing your risk of being sued for copyright infringement, and increasing the potential damages if you lose – but, if you decide (in any individual case) that is a risk worth taking, you are free to take that risk. If MIT tried to issue a DMCA takedown to arXiv over this, arXiv might decide that defending their own policies is worth the risk of being sued by MIT.
> Tables 17 and 18 in the Appendix could probably be removed as they seem to verbatim copy course descriptions, as well as, maybe, Figure 4.
Probably a sufficiently small extract from the source material, that it would fall under fair use? (Lack of acknowledgement of the specific source may be an issue; but that can be remedied by adding an acknowledgement, rather than removal.)
For the tables there's very little transformation, and a huge chunk of verbatim text. I don't see how there is any gain versus just publishing the course numbers and titles.
For figure 4 this might fall under "unpublished material" protections, which are: https://www2.archivists.org/publications/brochures/copyright...
> Generally, material is considered unpublished if it was not intended for public distribution or if only a few copies were created and distribution was limited.
> The law distinguishes between published and unpublished material and the courts often afford more copyright protection to unpublished material when an asserted fair use is challenged.
> Rather, courts evaluate fair use cases based on four factors, no one of which is determinative in and of itself:
2) > Courts give more protection to works that are “closer to the core of copyright protection,” such as unpublished
4) > The effect of the use upon the potential market for, or value of, the copyrighted work: This factor assesses how, and to what extent, the use damages the existing and potential market for the original.
Publication of the (possibly) previously unpublished copyrighted work in figure 4 fully and completely destroys its value. I don't know if a fair use claim can overcome such an impact, though that is up to a court to determine.
IANAL either–but how is the copyright owner (MIT presumably) harmed by the reproduction of these course descriptions? It isn't like they harm the commercial value of the courses in any way; the course is the actual product here, the description is just sales and marketing collateral, and has minimal value apart from the product it is selling.
Furthermore, given the fact the paper was coauthored by MIT employees – arXiv could argue that MIT (through its employees acting as its agents) had granted them an implied license to reproduce it. Which is the other issue – even if this isn't fair use, MIT may have agreed to license it through its agents. You can still be bound by the actions of your employees, even if those actions violated your own internal policies–especially in dealings with third parties who had no reason to suspect there was any such violation.
> I don't see how there is any gain versus just publishing the course numbers and titles.
"Algebra I" and "Algebra II" don't mean much – what topics do they actually cover? A one sentence/paragraph course description adds a lot, because they tell you what topics are actually covered. Yes, someone could probably look it up on the MIT website – but it saves the reader a lot of effort doing that. Especially if someone is reading this 20 years from now, by which time the content of MIT courses may have changed a lot (despite having the same title), and finding what their content was 20 years ago may require a lot of research effort (if the reader even thinks to do that).
> Publication of the (possibly) previously unpublished copyrighted work in figure 4 fully and completely destroys its value
Figure 4 is likely not the "work", rather a small quote from a much larger work. How does a small quote from a work (even if allegedly unpublished) "fully and completely destroys its value"?
> IANAL either, but figure 4 is likely not the "work", rather a small quote from a much larger work. How does a small quote from a work (even if allegedly unpublished) "fully and completely destroys its value"?
Exams are often composites of multiple independent works. Said exams being recomposited periodically (i.e. using a database of questions to create an exam). The argument here is that the individual question is itself a complete work (equivalent to an independent chapter in a book of works on a topic). And here it is not just on its lonesome, but with its answer, too.
If figure 4 came from an exam. For all we know, figure 4 actually came from course notes, assignments, etc. Whether or not issuing those to students counts as "publication", they are easily available to future students in a way that past exam questions are often not, hence their publication does far less damage to their value.
Also, MIT says that "Iddo did not have permission from all the instructors" – for all we know, figure 4 is from one of those instructors for which he did have that permission.
Based on a quick search it seems the figure 4 question and answer have to do with https://en.wikipedia.org/wiki/Markov_decision_process , which seem to be used in computer science. Iddo Drori is an associate professor of CS, so it seems quite likely it's his own question.
No, GPT4 Can’t Ace MIT - https://news.ycombinator.com/item?id=36370685 - June 2023 (120 comments)
This is the conclusion of the memo. The problems with methodology are clearly a secondary concern. They also seem to imply that issues with consent are solely or mainly the cause of the methodology problems. This could be true, but idk.
This is my opinion as an MIT grad who still views OCW content and similar content from other universities from time-to-time.
(Complete speculation: I almost feel as if certain professors or lecturers are “embarrassed” about their content.)
(Also an mit grad).
This thread was about whether the response focused on the consent or methodology problems and it’s my opinion that they are focusing on the consent issues more.
I only copied a quote from the memo.
I think they are point out consent issues in multiple cases, but also talk about "problems that should be corrected before publication".
I can't really see how there would problems to be corrected before publication.
If it's only a consent issue but the rest of the methodology is sound, they should just say "Drori should have waited for the consent to go ahead with publication, but the paper is fine"?
I didn’t mean that it was the be-all and end-all of those CSAIL profs’ perspective(s).
The wording is still very suspect to me and seems like the Writing Center would have some constructive feedback for them (unless the meaning of words doesn’t matter).
Apparently some of the other profs were not in the loop about the arxiv submission though
However, I generally agree with your take, and the response seems to me to kind of dodge the responsibility for methodology issues. that responsibility is not really compatible with the “hey we’re submitting to arxiv in a few days, last chance for comments” approach to collaboration that seems to be the minimum expected bar for signing off on the submission form that all authors are aware and agree to publication
I know it's standard on HN to accuse senior academics of exploiting their PhD students and I'm pretty sure there is plenty of that to go around but it is a very rare PhD student that can write a publishable paper without advise and guidance from their advisor. The name gives a hint, even.
I'm speaking in this as a recently graduated PhD student btw. I don't have any conflicts of interest (well, not yet, hopefully). It sucks that many students have absolutely rotten relations with their advisors but that's exactly because the student depends so much on the advisor for direction that it's easy for the power differential to be exploited by unscrupulous individuals.
That is certainly an issue that would be discovered well before anyone is sitting down to write the actual paper.
Sounds more like they're throwing the one senior prof under the bus to protect their own faces; saying 'I didn't agree to preprint submission' would be the first excuse I'd come up with.
At that point, they had been majorly involved in the research, if they were not involved in the preprinting they must have seen several drafts at least, at which point they could've asked to get taken out. But they didn't do that, and that's telling.
The takeaway IMO seems to be to prepend the abstract with a clear disclaimer sentence conveying the uncertainty of the research in question. For instance, adding a clear "WORKING DRAFT: ..." in the abstract section.
However, there is no shortage of projects with sketchy data collection methodologies on arXiv that haven't received this amount of attention. The point of putting stuff on arXiv _is_ that the paper will not pass / has not passed peer review in its current form! I might even call arXiv a safe space to publish ideas. We all benefit from this: a lot of interesting papers are only available on arxiv v.s. being shared between specific labs.
I'm concerned that this fiasco was enabled by this new paradigm in AI social media reporting, where a project's findings are amplified and all the degrees of uncertainty are repressed. And I'm honestly not sure how to best deal with this other than either amplifying the uncertainty and jankyness in the paper itself to an annoyingly noticeable level, or just going back to the old way of privately sharing ideas.
Maybe this is the best case scenario for these sorts of papers? They pushed a paper on a public journal, and got a public "peer review" of the paper. Turns out the community voted "strong reject;" and it also turns out that the stakes for public rejection are (uncomfortably, IMO) higher than for a normal rejection. Maybe this causes the researchers to only publically release better research, or (more likely) this causes the researchers to privately release all future papers.
The other side is flag planting with half-baked ideas and results.
in a practical sense, most people don't think of it this way though. putting something on arXiv means "i want people to be able to cite this, and for it to show up on Google Scholar", which might include works in progress, short notes, or lots of other things that don't fit the official criteria.
TL;DW - Their approach is to sequentially test methods, moving on to the next if one fails. However, this strategy is flawed as it requires ground truth, particularly for multi-choice answers. The analogy could be made to continually rolling a dice until landing on six. Similarly, if a question has four potential answers, the model merely has to attempt four times to stumble upon the correct response. And then they report 100% success rate.
The bigger BS is that several questions weren’t questions and some didn’t have enough detail to be answered and it somehow got 100% on those. This IMO out the entire paper into the trash.
In addition, for me there is a much wider ethics issue here. How many papers can really be checked properly by someone claiming authorship even if that author isn't really contributing? I can read a paper like this in about a fortnight because I have a job and a life. A faculty member also has a job and a life - they are doing admin and teaching as well as research. So to me it's impossible for someone to check 25 papers or more a year.
I am seeing far higher counts than this by many academics.
But, this is an extremely conservative threshold in my opinion. When I have contributed to scientific papers it has taken me at least three months of solid work each time. Often these papers get rejected (rightly) and then have to be substantially amended (or occasionally just abandoned) I am not that talented for sure, but I really find it hard to credit that anyone with an actual job (so not a post-doc or a student) can contribute to more than one academic paper a year. Potentially two or three if there is a confluence of papers getting ready for print... but not on a sustained basis.
There are two solutions. Every university and research institute needs to investigate all the publications of academics with high paper counts per year. This is a red flag. I definitely think that if folks are in the top quartile in a department it needs to be looked at carefully.
The other solution is that no academic publishing venue (conference or journal) should accept more than one paper per year from any author.
I assume the paper itself is long gone.
Yes, it shows fortitude to pass MIT’s exam, but beyond that, a lot of stress for a test that will soon be forgotten after their undergraduate degree.
Perhaps they should crack down on that, but I think the relative openness of MIT is overall a good university culture to have, and only a small minority of those with tenuous affiliation are actually grifters.
Sadly, this is an extreme outlier for now.
So what did the student authors work "really hard on" if not the same data that was collected without consent? Either all student authors are at fault for working on a paper based on data collected without consent or none are. It's not publishing the paper that is the problem here.
This is seperate from the issues related to the methodology not agreeing with conclusion that GPT4 could "get an MIT degree".
The second conclusion and methodology could have been fixed potentially.
Copyrighted data would make it harder to share and use, to reproduce results. I don't remember coming across a benchmark dataset that was copyrighted.
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
1. Professors and MIT students attempting to gain fleeting fame riding on the bandwagon without trying to do deep work
2. Professors at MIT getting upset and lash out that GPT-4 can now get a MIT degree devaluing said degree.
What has become of academia these days ?