Imageomics / Imageomics/FuncaPalooza-2025

Using computer vision to better predict snake mimicry

Open
#7 2 comments 1 reaction 0 assignees View on GitHub
dataset project idea
Dominant language
No language data
Stars
5
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Hi everyone!

I am bringing a large dataset of images to the workshop for which I am hoping to build a ML image processing pipeline. I am primarily an evolutionary biologist and herpetologist with minimal training in ML/AI methods, but I'm very interested in using these methods to quantify complex traits -- which normally are discretized or highly simplified to be studied.

I study coral snake mimicry, the largest mimicry complex in vertebrates, which involves nonvenomous snakes mimicking the bright conspicuous warning color of the highly venomous coral snakes (examples in images below). I have 8,000 images of the color patterns of 2,700 specimens of snakes across 130 species, collected during a trip to Brazil last winter where I visited 9 natural history collections across the three major biomes of Brazil. My hope is to train a CNN or vision transformer to be able to extract only the color pattern phenotype of a snake from an image, and then quantify that pattern by its position in an embedding space built during training. The output would then be a quantitative and holistic measurement of a traditionally discretized phenotype, color pattern, that would allow comparison between different specimens and species. This metric would be useful to quantify things like mimetic accuracy/precision (which can then be compared to other variables), or to better train snake identification models, which tend to fail with mimics. Hoyal-Cuthill et al 2019 [https://doi.org/10.1126/sciadv.aaw4967] provide an example of a pipeline made for butterfly mimicry that is very similar to what I would like to build.

Below are two photos that show the standardized format of the entire dataset and display the color pattern phenotype. The first is a coral snake (the model) and the second is a mimic. The first one is an older specimen so the bright red we would see in life has faded to a dull peach (which is a source of variation that will need to be dealt with as well).

For some more info on coral snake mimicry --

The New World coral snake mimicry complex involves ~80 species of highly venomous coral snakes and ~150 species of nonvenomous snakes that mimic the color patterns of coral snakes. Mimicry of coral snakes has evolved independently 20 times across various clades of nonvenomous snakes, involving around 20% of all snake species in the New World. You may be familiar with butterfly mimicry in _Heliconius_ butterflies, but typically coral snake mimetic phenotypes are more variable and unpredictable, likely because predators have evolved a stronger aversion to deadly snakes than distasteful butterflies. We understand much less about snake mimicry than butterfly mimicry, as we know little about what factors structure mimetic precision, mimic spatial distributions, and phenotypic evolution in coral snake mimicry. By quantifying the mimetic phenotype in a way that retains more of its complexity and variation, my hope is that certain variables (such as the geographic proximity of two species) will be better predictors of mimetic accuracy. The implication is that some of the theoretical confusion surrounding mimetic systems arises from an inability to properly quantify the mimetic phenotype.

If you are someone who is skilled in computer vision, or if you are interested in the application of ML methods to trait evolution (there have alreday been several posts about this!), I would love to chat.

Contributor guide

No contributing guide indexed for this repository

Research direction

No repository files, tests, entry points, or implementation plan are named. Start by reviewing the 8,000-image dataset requirements and the Hoyal-Cuthill et al. 2019 butterfly-mimicry pipeline. A completed effort would define and implement a reproducible pipeline that extracts snake color-pattern phenotypes and produces a usable quantitative embedding.

Written by the indexing model from the issue text.

Assessment

Domain
computer-vision, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.