Sincerely, someone who reads a lot of research but contributes none because I’m an amateur.
Edit: when I say data, I mean the raw data.
But the code is secondary to the idea. The idea and the discussion around how it was arrived at and what it means is the key thing. The code is just there to implement it. You could code the same idea ten different ways.
“But the proof is secondary to theorem. The theorem and the discussion around how it was arrived at and what it means is the key thing. The proof is just there to show it’s true. You could write a proof for the same theorem ten different ways.”
Which is all true. But man it wastes so much time having to re-prove everything. Also some lemmas/theorems are so hard to prove. It’s much easier when you see some incredible statement and can’t believe it’s true to look at the proof and see where the mistake / contentious part is.
If anything, I'd like to see more focus on giving evidence of generality of a result, vs just sharing everything needed to get back the same specific result
You're saying that the majority of CS papers only appear to work because the analysis code has bugs?
And that checking the code (presumably also the analysis code) is easier than "understanding the idea"?
Neither of those ring true to me, but your mileage may vary.
I don't think any specific field was mentioned
Likely the tech stack you use is built on a tower of ‘just idea’ papers.
You are right. But that is why we are having this discussions, so we can improve situation.
Having even bad code (and corresponding data) available is always better than not. You can always just ignore it, and read the papers like today.
Honestly I am ok with just zip file of project directory that you have anyway, with hopefully list of versions of os, libs and programs used.
We could do a lot better than just a zip file, but that would be a nice start.
https://en.wikipedia.org/wiki/Growth_in_a_Time_of_Debt#Alleg...
For example the paper on polymorphic inline caching, which is the key idea for the performance of many programming languages today, just described the idea, and didn't present any code. How was it evaluated? People sat and thought about it. Holds up today.
You can reason about an idea through other things than concrete code. Code is transient and incidental. Ideas persist.
It's not like it takes a whole lot of time to just dump your code in a github repo once you're done and link it somewhere on the paper (if you wrote code at all while working on the paper).
Sometimes I did just want to run my own experiments with different datasets, and those algorithms aren't always trivial to implement :|
Yep but we should still show we did actually simulate our idea, and the methodology that gave rise to the simulation. Not because of the code but to test at all a simulation we describe actually outputs what we propose
Not everyone is a programmer but they could find one to confirm, or better yet, invalidate code my team relied on
> but contributes none because I’m an amateur
I don't mean to be rude but it seems relevant to point out. Papers aren't written for the benefit of amateurs. They're written for experts who actively work in that specific field. I don't think there's anything wrong with that.
And I don’t think it’s rude, that’s why I included that statement!
This is because the datasets were subscriber logs from mobile operators. They are both highly privacy sensitive and contain sensitive business knowledge. There is no way they will ever get published, even in some anonymized form.
Ultimately it always comes down to trust. You need to convince your peer reviewers to trust you that you have correctly done what you have claimed to have done. Of course, even when you publish datasets, you need to convince the peer reviewers to trust you that you didn't fake the data.
I like the idea of publishing the data with the paper but it's not feasible in every case.
In my case I'm currently finishing up a paper where the raw data it's derived from comes to 1.5 PB. It is not impossible to share that, but it costs time and money (which academia is rarely flush with), and even if it was easy at our end, very few groups that could reproduce it have the spare capacity to ingest that. We do plan to publicly release it, but those plans have a lot of questions.
Alternatively we could try to share summary statistics (as suggested by a post above), but then we need to figure out at what level is appropriate. In our case we have a relevant summary statistic of our data that comes to about 1 TB that is now far easier to share (1 TB really isn't a problem these days, though you're not embedding it in a notebook). But a large amount of data processing was applied to produce that, and if I give you that summary I'm implicitly telling you to trust me that what we did at that stage was exactly what we said we'd done and was done correctly. Is that reproducibility?
You could also argue this the other way. What we've called "raw data" is just the first thing we're able to archive, but our acquisition system that generates it is a large pile of FPGAs and GPUs running 50k lines of custom C++. Without the input voltage streams you could never reproduce exactly what it did, so do you trust that? Then you're into the realm of is our test suite correct, and does it have good enough coverage?
I think we have a pretty good handle on one aspect of this, is our analysis internally reproducible? i.e. with access to the raw data can I reproduce everything you see in the paper? That's a mixture of systems (e.g. configs and git repo hashes being automatically embedded into output files), and culture (e.g. making sure no one things it's a good idea to insert some derived data into our analysis pipeline that doesn't have that description embedded; data naming and versioning).
But the external reproducibility question is still challenging, and I think it's better to think about it as being more of a spectrum with some optimal point balancing practicality and how much an external person could reasonably reproduce. Probably with some weighting for how likely is it that someone will actually want to attempt a reproduction from that level. This seems like the question that could do with useful debate in the field.
certainly sharing apparatus is hard but you could release the schematics, board designs and BOMs of the electronics involved.
The problem now is that 1) very few even try to reproduce 2) very little money is available for reproduction
fixing those incentives would help alot.