https://gitlab.com/machine-biology-group-public/pancleave
>This package implements a scikit-learn-based random forest classifier to predict the location of proteoylytic cleavage sites in amino acid sequences. The panCleave model is trained and tested on all human protease substrates in the MEROPS Peptidase Database as of June 2020. This pan-protease approach is designed to facilitate protease-agnostic cleavage site recognition and proteome-scale searches. When presented with an 8-residue input, panCleave returns a binary classification indicating that the sequence is predicted to be a cleavage site or non-cleavage site. Additionally, panCleave returns the estimated probability of class membership. Through probability reporting, this classifier allows the user to filter by probability threshold, e.g. to bias toward predictions of high probability.
In the face of current hype around LLMs and 'fear of AI', calling a Random Forest Classifier 'AI' is a bit... far