332 karma · joined July 14, 2016
The training times obviously vary on the network architecture, software and hardware. I can safely say you can process 7200+ protein sequence with average sequence length of 120 amino acids in 2h on 2 x NVIDIA Titan XP
Coming back to your comments about the canonical secondary structures; I couldn't agree more with you. The problem is quite simple, how are we going to convince the >90% of structural biochemistry society to simply accept the fact proteins are bloody dynamic and X-ray / eye candy structures may have quite little to do with the "real" picture at room temperature?
We have in "stock" a network (obviously another paper) that will aim at propensity prediction, still in trivial alpha/coil/beta phase space.
In my company, we are ridiculously pedantic about unit testing, but even with proper level of attention to detail we sometimes fail with getting 100% code coverage.
The biggest pain are log(x)/ln(x) issues in numerical optimization.
We are essentially at the crossroad and obviously, programming, which develops nowadays a bit faster than theoretical physics or mathematics, pushes in one direction.
The problem is the actual math/physics formalism is not!
Start reading here: https://en.wikipedia.org/wiki/Protein_pKa_calculations
This is just the tip of the "log(x) vs ln(x)" iceberg.
I have copy pasted the ln(10) straight from Apple Calc. Thanks for pointing out.
It's not my area of expertise, though. My personal experience at Cambridge and Groningen University was always very good with students. Yet, these are peculiar places.
Basic mathematical knowledge + ability to think logically = success