Jeremiah has Cousins and I Can Prove it
- Anushka Ring
- Apr 14
- 3 min read
Updated: Jun 22

I know what you’re thinking. Why does this look like a crazy seating chart for a big family reunion where some people have to be a certain distance from other people in order to avoid violent fights while others may want to share a table and reminisce over childhood memories? Or you saw the large fish scattered between the networks and assumed that this had to do with fish, given that a lot of my prior blog posts have been about my research barcoding Congolese fish at the American Museum of Natural History. Either way, you’d be on the right track.
The last time I updated you, we had just finished sending all our DNA samples of the amplified COI gene to the MC Lab in California to be sequenced. Thanks to the efficiency of the US postal service, we got our results back pretty quickly. What was immediately clear was that before being able to analyze and really derive meaning from these results, we had to assemble the genes and create Haplotype networks, which are the networks in the picture—surprise! The California Lab had sequenced each gene for each fish both forwards and backwards (refer to New Year New Species for how Sanger sequencing works) in order to account for any potential errors that might occur in one sequence. However, we now had to condense the two sequences into just one for each fish.
Armed with the snacks we raided from the office of the director of my internship program, we turned to the software Geneious, which really lives up to its name. This software was able to align the two sequences with one another—lining each base up with the same base on the other sequence—and also flip the reverse sequence to run in a forward direction like the other. It also showed us something called a Chromatogram (pictured below).

A chromatogram is a visual chart that displays the results of sequencing. It shows colored peaks corresponding to the four DNA nucleotide bases: A (Green), T (Red), C (Blue), and G (Black). Depending on how high the peak is, that is how confident the machine was that it was that nucleotide. Therefore, if base 62 in the forward sequence said A and base 62 in the reverse sequence said T, but the T peak was higher, the software Geneious would fill that base in as T, and so on and so forth. It’s like if my brother and I wanted different things for dinner and my mom said whoever could jump higher could choose what we ate. Obviously, we would end up eating whatever I wanted.
Once we had the finalized sequences, we assembled them on MEGA and added them to already existing sequences of our fish species found on the BOLD Systems database. We did this in order to compare our results to past research and to track changes in genetic variation. Finally, we configured Median-Joining Haplotype Networks on PopART, yet another tool. I asked my mentor if there would be a day where you could do all of this on one database, and he told me that I should do it…We’ll see.
Each network is for a different species of fish. The six species we ended up having enough data for were:
Congopanchax brichardi
Epiplatys duboisi
Aphyosemian Elegans
Congopanchax multifasciatus
Hylopanchax moke
Epiplatys chevalieri
Each color in these networks is a different location that these fish were sampled from. Every single circle is one version of the COI gene, and the distance between the circles determines how far genetically the versions are from one another. The size of the circle determines how many individuals of that species shared that version of the gene (bigger = more common).
For example, take a look at our Epiplatys Chevalieri species on the top left. I know, they might not be as colorful as the rest of his fish buddies, but they actually serve as a pretty good example. Take the two red circles on the top—there’s a number 2 on the line between them. This means that the COI gene in the first circle was only 2 nucleotides different from the COI gene in the second circle. The second circle is also around twice the size of the first circle, meaning that around two individuals shared the second COI gene and only one had the first. Poor little guy, excluded like that. Essentially, it’s a map of genetic relationships — showing which versions evolved from which, how diverse a population is, and whether groups that live far apart are genetically similar or different.
Stay tuned for an upcoming blog post, where I’ll dive into what we think these networks show. Before reading it, you should try and make some interpretations of your own based on what you’ve just learned about haplotype networks.
HINT: there's a new species somewhere!!!!!!!

Comments