What problem is this method solving?

Animal venoms are mixtures of proteins and peptides that have produced real medicines, and around twenty venom derived medications are in approved use. New species keep being sequenced, and for each one somebody has to decide which of the resulting protein sequences are toxins. Doing that from sequence alone, quickly and cheaply, is what a toxin classifier is for.

A note on terms. A toxin is a single compound with a specific binding site, usually encountered passively or ingested. A venom is a blend of compounds, stored in a specialised structure and actively delivered, such as through a fang. The distinction is the delivery mechanism, not the chemistry. In the dataset, toxic and atoxic are simply the two class labels.

What data was used, and why that data?

Sequences come from UniProtKB, the central collection of functional protein information, obtained through the dataset published with the TOXIFY project. That choice was deliberate: using the same data as TOXIFY, ToxClassifier and ClanTox means the resulting numbers can be compared with theirs rather than standing alone.

Three preparation decisions shaped the training set.

  • Length cut off. The raw training set held 6,001 toxic and 49,764 atoxic sequences. Distribution plots showed that a cut off of 500 amino acids captured the bulk of both classes, so longer sequences were dropped. This also kept training feasible on the hardware available.
  • Class balancing. With roughly eight times more atoxic than toxic sequences, a classifier that simply favoured the majority class would have scored well on accuracy while learning very little. The atoxic sequences were randomly downsampled without replacement to match the toxic count, giving 5,896 in each class.
  • Cleaning. FASTA records occasionally contain letters that are not one of the twenty standard amino acids and so have no Atchley values. Those were removed before parsing.

The benchmark set was left untouched in terms of class balance, and was not inspected during modelling decisions, so that it could not quietly inform the design. It was put through exactly the same preprocessing as the training data, because the network requires inputs of a fixed shape.

How do Atchley factors encode amino acids?

The underlying difficulty has a name: the sequence metric problem. Writing a protein as letters is convenient but tells you nothing useful about chemical similarity. Leucine (L) is more similar to valine (V) than to alanine (A) in its physicochemistry, yet the alphabet quietly suggests the opposite ordering. Letters give no measure of distance that reflects reality.

The Atchley factor analysis addressed this by starting from 494 measured amino acid attributes, selecting 54 of them on statistical and interpretability grounds, and reducing those by factor analysis to five clusters. Each amino acid then carries five numbers:

The five Atchley factors, named as they are in the dissertation.
Factor Name What it summarises
f1 Polarity Bonded against non-bonded energy, hydrogen bond donors, polarity, hydrophilicity against hydrophobicity.
f2 Secondary structure The relative inclination to fold into a particular secondary structure, such as a coil or an alpha helix.
f3 Molecular volume Size and volume, including bulkiness, residue volume, side chain volume and molecular weight.
f4 Relative composition The number of codons coding for the amino acid, and amino acid composition.
f5 Electrostatic charge Isoelectric point and net charge.

In practice the encoding is a lookup. The twenty amino acids each map to a row of five values, so a sequence of length n becomes an array of n by 5 numbers. A residue is no longer a symbol, it is a point in a five dimensional chemical space where distances mean something.

Two features were built from this and compared head to head. The first is raw magnitude, the Atchley values themselves at each position. The second is sequential change, the difference in each factor between one residue and the next. For example, if a residue has a polarity value of -0.591 and the next has 1.831, the sequential change recorded at that position is 2.422.

Why turn a protein sequence into an image?

Because once each residue carries five numbers, the shape of the data invites it. Each protein was inscribed into a 5 by 500 matrix: five rows for the five Atchley factors, 500 columns for the maximum sequence length. Shorter sequences were zero padded to fill the matrix, which is standard practice when a convolutional network needs a fixed input size.

A 5 by 500 matrix with one channel is structurally the same thing as a small grayscale image. That matters because of what convolution does: the network slides small filters across the input, learns local patterns, then combines them into larger ones layer by layer. Applied here, it sees all five chemical properties at once across a stretch of sequence, instead of treating polarity and volume as five unrelated one dimensional signals. The encoded sequence becomes a picture of the protein's chemistry, and the network looks for the shapes that toxins tend to have.

The data were reshaped to (11792, 5, 500, 1) for training: 11,792 sequences, the 5 by 500 matrix, and a single channel, where a colour image would have three. Labels were shaped (11792, 2) for the two classes.

What does the CNN architecture look like?

It follows LeNet-5, the 1994 architecture by LeCun, Bottou, Bengio and Haffner that was built to read handwritten digits. LeNet-5 has seven layers after the input: convolution, average pooling, convolution, average pooling, two fully connected layers and an output layer. It was chosen precisely because it was designed for small single channel images, which is what an Atchley encoded sequence resembles.

Several parameters were changed to suit sequence data rather than digits, including the activation functions, the input size and the optimiser. Two dropout layers were added as regularisation against overfitting the validation set. The output uses a sigmoid activation with binary cross entropy loss, which is the standard pairing for a two class problem: sigmoid squashes the output into the range 0 to 1, and binary cross entropy measures how wrong that value is.

Training setup, kept deliberately close to TOXIFY so that results stay comparable:

  • Built with Keras on TensorFlow.
  • 50 epochs, matching the number TOXIFY trained its recurrent network for.
  • A 90% training and 10% validation split using k-fold stratified cross validation, so class proportions are preserved in both parts.
  • Optimisers compared during development, including adam and nadam, with the best performing configuration kept.

Which feature worked better, raw magnitude or sequential change?

Raw magnitude, clearly, with the identical architecture used for both.

Results on the benchmark set, same CNN architecture, two input features.
Input feature Accuracy Loss
Atchley raw magnitude 0.931 1.013
Atchley sequential change 0.820 2.153

The training curves told the same story before the benchmark did. With raw magnitudes, loss decreased and accuracy increased across the epochs, with a visible gradient and some validation peaks around epochs 32 and 34. With sequential change, both loss and accuracy flattened out, and the validation line was considerably more volatile throughout.

The cautious biological reading is that the actual chemical values along a sequence carry more signal for toxicity than the residue to residue transitions between them. Taking the difference appears to discard more than it adds.

One observation from the data analysis stage is worth recording. Averaged across sequences in both classes, Atchley values start large at the beginning of a sequence and reduce towards the end. Atoxic proteins show larger values at the start than toxic ones, and toxic proteins vary more across the whole sequence where atoxic proteins change steadily. These are averaged trends, not a diagnostic rule.

How does 93.1% compare with other tools?

On the same benchmark set, TOXIFY reported a final accuracy of 96.0% and ToxClassifier reported 99.7%. ToxinClass reached 93.1%. That is third of three, and the framing in the dissertation is deliberately plain: the result is comparable, and it comes from an architecture that had not previously been applied to this problem, which is the contribution being claimed. It is not a new best score and is not presented as one.

Two caveats run in both directions. Higher published figures from the mid 2010s are worth reading with care, since the authors of CSM-Toxin later reported that the performance of several earlier methods deteriorated significantly from their published results when re-evaluated. Equally, ToxinClass was trained with the compute available to one student, on a downsampled and length limited dataset. Varying the split ratios or training on a larger set is the obvious route to a better number.

What are the limitations of this method?

Stated flatly, because the limitations are the interesting part.

  • The answer is binary. Toxic or atoxic, with no indication of what kind of toxin a sequence is, and no grading of how strong it is. The training data carries no labels distinguishing mildly from highly toxic sequences.
  • No reasons attached. This applies equally to the other neural approaches in the field. The model takes no biological background knowledge as input, so it cannot produce a rule explaining why a sequence is toxic. It gives a verdict, not an argument.
  • Species are pooled. The dataset lumps all venomous species together. That gives a general view of what makes a sequence toxin-like, but it makes questions about how venom develops across the evolution of one species invisible to the model.
  • Length is capped. Sequences above 500 amino acids were excluded, so longer venom proteins are outside the trained range.
  • Compute bound. Downsampling from 49,764 atoxic sequences was partly a hardware decision. A larger training run is an obvious unexplored direction.

What is planned next?

Everything in this section is planned, not shipped. It is recorded here because the dissertation set it out as future work and the product roadmap follows it.

  • Reasons alongside labels. The dissertation proposes inductive logic programming, which sits at the intersection of machine learning and logic programming, as a route to predictions that come with rules rather than only scores. It names ProGolem, a system based on asymmetric relative minimal generalisation, as a candidate. The known trade-off is computational cost on high dimensional data, and the fact that the surrounding tooling is not widely used today.
  • Beyond toxic or atoxic. Moving from a binary verdict towards toxin class or mode of action, which is what researchers in this area increasingly ask for.
  • A hosted tool, batch screening and an API. The dissertation's own evaluation says it plainly: the author would have liked to produce a web tool, in the way ToxClassifier and ClanTox once did, to make this accessible to people who are not going to install anything. That is what ToxinClass is now.
  • Evolutionary tracking. Following venom protein development within specific species, acknowledged as a much higher dimensional problem than the one solved here.

Frequently asked questions

How do Atchley factors encode amino acids?

Each amino acid is replaced by five numbers summarising polarity, secondary structure propensity, molecular volume, relative composition and electrostatic charge. They come from a multivariate analysis that began with 494 amino acid attributes, narrowed to 54, and reduced by factor analysis to five clusters. The point is to give residues a meaningful numerical distance from each other, which letters cannot do.

What data was ToxinClass trained on?

UniProtKB sequences, obtained through the dataset published with TOXIFY. The raw training set held 6,001 toxic and 49,764 atoxic sequences. After excluding sequences longer than 500 amino acids and randomly downsampling the atoxic class without replacement, the balanced training set was 5,896 toxic and 5,896 atoxic sequences.

Why use a convolutional neural network on a protein sequence?

Because the Atchley encoded protein is a 5 by 500 single channel matrix, which has the same shape as a small grayscale image. Convolution lets the network read all five chemical properties together along a stretch of sequence rather than as five separate signals.

How accurate is ToxinClass?

93.1% accuracy with a loss of 1.013 on the benchmark set used by TOXIFY and ToxClassifier, using raw Atchley magnitudes. On the same benchmark, TOXIFY reported 96.0% and ToxClassifier reported 99.7%.

Which feature worked better, raw magnitude or sequential change?

Raw magnitude, at 93.1% accuracy and 1.013 loss, against 82.0% and 2.153 for sequential change under the identical architecture. The sequential change model also produced flatter training curves and a much more volatile validation line.

What are the limitations of this method?

A binary output with no toxin class or strength grading, no explanation of why a sequence was called toxic, a 500 amino acid length cap, species pooled together in the training data, and a downsampled training set chosen partly for compute reasons.

The full write-up

The dissertation covers all of the above at length, including the data distributions, the model summary, the training curves and the ethics statement.

Read the dissertation (PDF, 2.7 MB) or read the companion piece on why toxin prediction tools disappear. Questions about the method? Email hello@toxinclass.com.

References

  1. Van Sebroeck C. "Machine Learning Approaches for Classifying Animal Venom Proteins." University of Surrey, Department of Computing, 2020. Supervisor: Dr Alireza Tamaddoni-Nezhad. Read the PDF
  2. Atchley W. R., Zhao J., Fernandes A. D., Drüke T. "Solving the protein sequence metric problem." Proceedings of the National Academy of Sciences. pnas.org/content/102/18/6395
  3. Cole T. J., Brewer M. S. "TOXIFY: a deep learning approach to classify animal venom proteins." PeerJ, 2019. peerj.com/articles/7200
  4. Gacesa R., Barlow D., Long P. F. "ToxClassifier." PeerJ Computer Science, 2016. peerj.com/articles/cs-90
  5. Naamati G., Askenazi M., Linial M. "ClanTox: a classifier of short animal toxins." Nucleic Acids Research, 2009. pmc.ncbi.nlm.nih.gov/articles/PMC2703885
  6. UniProtKB. uniprot.org/help/uniprotkb
  7. Muggleton S. H., Santos J., Tamaddoni-Nezhad A. "ProGolem: a system based on relative minimal generalisation." Proceedings of the 19th International Conference on Inductive Logic Programming, Springer, 2010. doc.ic.ac.uk/~shm/Papers/progolem.pdf
  8. "CSM-Toxin: a web server for predicting protein toxicity." 2023. pmc.ncbi.nlm.nih.gov/articles/PMC9966851

Accuracy figures for TOXIFY and ToxClassifier are as reported in their own publications. ToxinClass figures come from the dissertation.