The one-line command to go from EMAN2 coordinates to Unreal Engine 5 is kind of crazy.
As usual on these (rare) threads, I'm happy to answer any questions about structural biology or cryo-EM.
The one-line command to go from EMAN2 coordinates to Unreal Engine 5 is kind of crazy.
As usual on these (rare) threads, I'm happy to answer any questions about structural biology or cryo-EM.
I haven't used VMD for about 30 years, but even in the 1990s I was using it to visualize the full poliovirus structure (4 proteins in 2PLV * 60 copies, as I recall).
It took about 6-10 seconds per update on our SGI Onyx, but again, that was 25 years ago.
Loading a 60-fold icosahedral virus has used > 100 GB memory on my workstation, and resulted in a 0fps experience. It might render OK from the command line, but now imagine a few dozen of those, plus a cell, plus all the proteins in the cell...
Each atom record has ~60 bytes (x, y, z, occupancy, bond list, resid, segid, atom name, plus higher-level structure information about secondary structure, connected fragments, etc.) We had our own display list, so another (x, y, z, r, color-index) per atom, giving 20 more bytes. We probably used a GL/OpenGL display list for the sphere, and immediate mode to render that display list for each point, so all-in-all about 100 bytes per atom, which just barely fits in 128 MB.
That was also all single-threaded, with a ~0.1 Hz frame rate. Again, in the 1990s.
I wanted to see what more recent projects have done. Google Scholar found "cellVIEW: a Tool for Illustrative and Multi-Scale Rendering of Large Biomolecular Datasets" (2017) at https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5747374/ which says
> The most widely known visualization softwares are: VMD [HDS96], Chimera [PGH04], Pymol [DeL02], PMV [S99], ePMV [JAG11]. These tools, however, are not designed to render a large number of atoms at interactive frame-rates and with full-atomic details (Van der Walls or CPK spherical representation). Megamol [GKM15] is a state-of-the-art prototyping and visualization framework designed for particle-based data and which currently outperforms any other molecular visualisation software or generic visualization frameworks such VTK/Paraview [SLM04]. The system is able to render up to 100 million atoms at 10 fps on commodity hardware, which represents, in terms of size, a large virus or a small bacterium.
Following that is a section on Related Work:
> With their new improvement they managed to obtain 3.6 fps in full HD resolution for 25 billion atoms on a NVidia GTX 580, while Lindow et al. managed to get around 3 fps for 10 billions atoms in HD resolution on a NVIDIA GTX 285. Le Muzic et al. [LMPSV14], introduced another technique for fast rendering of large particle-based datasets using the GPU rasterization pipeline instead. They were able to render up to 30 billions of atoms at 10 fps in full HD resolution on a NVidia GTX Titan
Checking up on VMD, in "Atomic detail visualization of photosynthetic membranes with GPU-accelerated ray tracing" from 2016:
> VMD has achieved direct-to-HMD rendering rates limited by the HMD display hardware (75 frames per second on Oculus Rift DK2) for moderate complexity scenes containing on the order of one million atoms, with direct lighting and a small number of ambient occlusion lighting samples.
Those citations are 6-7 years ago, which make me scratch my head wondering why ChimeraX can't handle a picornavirus.
For example, in this case I loaded 2PLV with "open 2PLV" and on the right side, there's an option to select one of the mmcif assemblies, with select 1 being "complete icosahedral assembly, 60 copies of chains 1-4". With the default ribbon rendering, rotating is completely smooth; with all atoms displayed (wireframe or sphere), it's still smooth. Computing a surface for the entire capsid takes well under a second(!) and still renders smoothly. Rotating shows my GPU (nvidia RTX 3080 Ti) at about 50% utilization, and if I exit Chimera, my GPU's releases ~200MB of memory.
Chimera was never intended to do high quality rendering of cellular environments with many hundreds of proteins. It was intended for a combination of nice rendering and scripting directly in python. VMD definitely handled some extremely large scenarios faster. A dedicated small C++ using modern OpenGL would be able to do far, far more than Chimera when it comes to simple rendering without any scripting control.
EDIT - that's also for atoms. Going to 3D maps is significantly more computationally intensive. A typical, sub-tomogram average, or annotation will be in MRC file format, which is horrendously slow with a box size > 1024 pixels or so.
I've worked with 3d maps and the same thing has been true for over 20 years: if you want to work with extremely large systems, you need an expensive graphics card (I was the first person to port Chimera to Linux some ~20 years ago, and we always compared the performance of my gaming cards to faster professional cards. for volumes like 3d maps, the largest volumes always were quite close to the largest graphics cards)
Lots of current visualization software is focused on visualizing a single protein structure (for example, ChimeraX). New visualizing and modeling systems are being developed to go up in scale to cellular scenes and even whole cells. For example, systems like le Muzic et al.'s CellView (2015) [1] are capable of rendering atomic resolution whole cell datasets like this in realtime: https://ccsb.scripps.edu/gallery/mycoplasma_model/
[1] : https://www.cg.tuwien.ac.at/research/publications/2015/cellV...
I still think "few" is the wrong word. I usually think of "few" as meaning up to around 6, while Chimera and VMD can easily handle hundreds of proteins at the atomic level.
For minor differences, there's something called the guassian mixture model (lots of software packages have similar, but GMM is EMAN2's version).
https://blake.bcm.edu/emanwiki/EMAN2/e2gmm
What you can get out of the other end is a volume series, that shows reconstructed 3D volumes along various axes of conformational variability. This quickly turns into a multi-dimensional problem, but it has been very successful in, for instance, seeing all the different states of an active ribosome.
The problem is that many of the interesting or urgent pathologies have no obvious (or weak) associations. Or maybe the noise level is too high. So there's got to be a piece of the puzzle we're missing, or something is getting lost in the noise. Whether a neural network can pull something out of the noise remains to be seen, but if it can't, I'm not super optimistic about our chances. Overall I'd say trying to tackle the problem at the outcome level is probably more promising right now. Even if we can find good associations, we're still lacking therapies.