Are there any best practices for the construction of similar databases?
Are there any best practices for the construction of similar databases?
Best practices are incredibly difficult. We're trying to establish a common API currently (https://github.com/Materials-Consortia/optimade) that can be adopted by all database providers. How the data is stored behind the scenes is something that ends up being very specific to how the data is generated and what its applications are. We're definitely better as a community than we were ten years ago, but there's a lot of work to be done here.
In terms of scientific databases outside crystallography/materials science, Nature's Scientific Data is a good open-access journal to peruse: https://www.nature.com/sdata/
https://www.ncbi.nlm.nih.gov/geo/
The Cancer Genome Atlas (TCGA) for a wide-array of data from cancer patients in the US (Treatment history, outcomes, histology, expression, etc)
https://portal.gdc.cancer.gov/
InterPro from EMBL harmonizes multiple protein family databases
https://www.ebi.ac.uk/interpro/
STRING harmonizes protein-protein interaction databases