Showing posts with label genetic genealogy. Show all posts
Showing posts with label genetic genealogy. Show all posts

Friday, December 5, 2014

DNA Convergence and Chicken Little

   For me, the topic of convergence in yDNA first came up early in 2014.  I had just posted a paper and one of the comments was – “What about convergence?”  I said to myself, “What convergence?”  I admit I had to look up the topic.

Convergence: A term used in genetic genealogy to describe the process whereby two different haplotypes mutate over time to become identical or near identical resulting in an accidental or coincidental match. - Turner A & Smolenyak M 2004.

My response back to the comment was - “All of the haplotypes in my paper are unique.”  My data did not exhibit convergence. 
Convergence casts a shadow on genetic genealogy
   I started to poke around on the topic of convergence within yDNA STR haplotypes and the immediate impression that I got was that folks were ready to give up on STRs in favor of SNPs and the sky was falling.  Chicken Little was running around in the genetic genealogy circles.  Here is a small sample:

Y-STRs are effectively dead” - Dienekes Pontikos, 2011

Convergence of Y chromosome STR haplotypes from different SNP haplogroups compromises accuracy of haplogroup prediction” – Wang, et al, 2013

   Okay, convergence happens, but it’s an illusion.

   Let’s take a big step backwards in this story.  Did you know that most scientific papers relating to genetic genealogy use 17 STR markers or less?  Some use as few as 9 or 10.  For any of you who ever took one of the original 12 STR marker tests, you know that the results were essentially useless for anything except deep haplogroup association and history.

   Many researchers in the last couple of years are using the AmpFLSTR® Yfiler® to get their 17 marker results.  This equipment is approved for forensic cases.  Research papers are not forensic cases and researchers don’t need to limit themselves to 17 markers.  Thirty-seven marker yDNA tests have been available since 2004.

   Why does the number of STR markers matter?  I’m going to release my inner math geek to help explain.  If we look at marker DYS19, usually listed first in science papers and third in Family Tree DNA results, it can have a value within the range of 7 to 22 across all haplogroups.  Looking at R1b specifically, DYS19 ranges from 10 to 17 and statistically at two standard deviations (2 sigma) the range of values narrows to 13, 14 and 15.  From a probability point of view, there is a 1 in 3 chance that DYS19 will be 13, 14 or 15.  Making the odds even better in our favor, 95% of the time DYS19 for R1b will already be 13, 14 or 15.  This means there is a 1 in 2 chance that DYS19 could change to another value on its way to converging with another haplotype.

   Taking standard deviation into account to determine the possible number of values for the STR markers and then multiplying each probability gives the odds that a haplotype could converge.
STR
DYS393
DYS390
DYS19
DYS391
DYS385a
DYS385b
DYS426
DYS388
DYS439
DYS389i
DYS392
DYS389ii

Total
# of possible
marker values
2
4
2
2
2
4
1
1
2
2
2
2

4096

   There is a 1 in 4096 chance that two R1b 12 marker haplotypes could converge.  This is not the probability that one marker will change.  This is the probability that all 12 markers will change enough to match another haplotype.  These are very good odds and the reason why a 12-marker test is practically useless. 

   With a high probability that 12 STR markers will converge, haplotypes start to blend together.  Two different haplogroups or family lines will appear to be the same.  Converging also means that when we calculate the time to the most recent common ancestor (TMRCA), it will look like less time has passed.  Convergence makes a 12-marker test result unusable for genealogical matching, haplogroup prediction and TMRCA calculations.  The Chicken Littles are correct, we have a problem with 12 marker STR results.

   What about 17 markers, a quasi-industry standard for science papers?  Taking the same approach with statistics and probability, a 17-marker yDNA R1b result has a 1 in 2 million chance of converging with another haplotype.  Each haplogroup has slightly different odds.  There is a 1 in 500,000 chance of an R1a 17 marker haplotype converging.  Those odds are better than any lottery.  Convergence is still a problem at 17 markers.

   When Dienekes Pontikos proclaimed the death of yDNA STRs, he was commenting on the attempt to get good TMRCA dates from 10-marker results.  I agree, you can’t get valid TMRCA dates from 10-markers.  When Wang, et al, determined that convergence compromises haplogroup prediction, they were correct, 17 marker haplotypes can converge to make one haplogroup look like another.

   In a quick analysis of 4,300 unique 37-marker R1b haplotypes, the average genetic distance is 17 steps for 37 markers.  That means there are 17 mutations required for convergence in a 37-marker haplotype.  Nearly half of the markers in the haplotypes would need to change.  When we look at the probability of 25-marker haplotype convergence, the chances are 1 in 84 million.  Considering there are about 3.6 billion men on the planet, one in 84 million is still in the realm of possibility.  By the time we get to 37-markers, the odds are 1 in 49 trillion.

   There is a 1 in 49 trillion chance that all the necessary mutations will occur in order for two 37-marker haplotypes to converge.  The odds are likely much higher.  I’ve only looked at the probable values for each marker and I haven’t taken into account the STR mutation rates, the possibility that a marker will change over time. 

   There is essentially no such thing as convergence when 37 or more markers are tested and researched.  If you eliminate the possibility of convergence by using 37 STR markers, then immediately TMRCA calculation become more accurate and haplotypes from different haplogroups no longer resemble each other.  The reports of the death of yDNA STR results have been greatly exaggerated.


   I can’t tell you why researchers are currently stuck on 17 markers.  I can tell you that any research using less than 37 markers runs the risk of convergence in their data, which in turn could lead to the wrong conclusions.  I still consider genetic genealogy to be in its infancy.  Every month new research papers are published and the new concepts introduced are latched onto immediately.  It is understandable that papers from over a decade ago used a dozen STRs and a handful of SNPs, that was the height of technology.  If the latest technology and best data are not being used in today’s research papers, is that equivalent to scientific negligence?  Or, am I missing something and this is a case of scientific ignorance on my part?

Tuesday, December 2, 2014

DNA, SNP, STR, OMG!

(Originally published May 2014 in Going In-Depth)

   Oh my gosh, there are many acronyms in genetic genealogy.  You have to agree that using the acronym DNA is better than writing deoxyribonucleic acid repeatedly.  Although, when we talk about using DNA for genealogy and we only use acronyms, they start to lose their meaning and become just another ‘thing’.  “Hey, I’ve got a SNP.  Do you have a SNP?”  “I dunno, let me check.”  Maybe I’m weird.  I like to understand what all the acronyms mean and how they play a part in the larger picture.

   Let’s start with some DNA basics.  We have DNA in every cell except the red blood cells.  Inside the nucleus of our cells, we have 46 chromosomes or 23 pairs (nuclear DNA).  One set of 23 comes from dad and one set comes from mom.  If we took the tightly coiled DNA from one cell and stretched it out it would be about six feet long.  In that six-foot double helix from one cell, there are over 3 billion base pairs.  If you picture our double helix DNA as a twisted ladder, each rung is a base pair made up from four nucleotides (DNA building blocks).  The rungs are made from either an adenine-thymine rung or a cytosine-guanine rung.



   When we talk about DNA, we often also talk about mitochondrial DNA.  Mitochondria exist outside of the nucleus as an energy source for the cell and have their own independent DNA.  Mitochondrial DNA has just over 16,000 base pairs in comparison to the 3 billion base pairs in our nuclear DNA.  We inherit our mitochondrial DNA only from our mothers.

   DNA is divided into coding regions (genes that define proteins for such things as eye color) and non-coding regions (sometimes called junk DNA).  The coding region that defines us is less than 2% of our overall DNA and within that, there are less than 25,000 genes.  A gene is a sequence of nucleotides averaging about 23,000 base pairs.  One of the largest genes, which encodes for the Caspr2 protein, has over 2.3 million base pairs.


   Within the 3 billion base pairs of our DNA there are variations (normally occurring mutations), where one base pair has been replaced with another base pair.  As an example, it was adenine (A) and now its guanine (G).  This is a single nucleotide polymorphism or SNP (pronounced snip).  There are over 15 million SNPs in our DNA.  Once a SNP occurs, it is usually permanent in the population.  The farther back in time that the SNP occurred, the more people will have that particular mutation.  To be considered a SNP, it has to exist in greater than 1% of the population.  They are found in both the coding and non-coding regions of our DNA.  In the coding regions, SNPs are often markers for genes.

   Let’s divide our DNA into four groups.  Group one, the autosomes, are the first 22 pairs of chromosomes.  The next two groups, the sex chromosomes, are one X and one Y if you are male and two Xs if you are female.  That gives us yDNA and xDNA.  The last DNA group is mitochondrial.  All types of DNA have SNPs.  Autosomal SNPs are used for health and ethnicity.  Mitochondrial and Y-DNA SNPs are used to determine world haplogroups.  While there are 1,000s of X SNPs, there doesn’t seem to be much research around them.

   SNPs have no effect on health, but their presence may predict a health risk.  If you had an autosomal test from 23andMe (prior to the FDA ruling), they would have delivered health information with your results.  They were able to report SNPs in the coding region associated with gene combinations responsible for health risks, like cancer or Alzheimer’s or basic information, like eye and hair color.  Even though you cannot get health information from 23andMe currently, you can still use your autosomal results with Promethease from SNPedia.com to research your health risks.
   Combinations of SNPs are analyzed to determine ancestry-informative markers (AIM – another new acronym for you).  AIMs are used to estimate the ethnicity or at least the geographic origins of your ancestors.  When you receive ethnicity results from an autosomal test, it will be based on the AIMs that the test company are using.  They don’t all use the same markers, so results will vary.  There are even 42 SNPs associated with having Neandertal ancestry.
   SNPs are used to organize us into larger branches of the human family tree (haplogroups).  Our maternal family tree is organized into 26 branches (A through Z) using mitochondrial DNA.  Our paternal tree is similarly organized into 20 branches (A through T) using yDNA SNPs.   As an example, take four men (I use men because the scenario works for both mitochondrial DNA and yDNA), Abe, Bob, Chaz and Dave.  Test each of them for three SNPs, X, Y and Z.  You find that they all test positive for SNP Z, Abe and Chaz test positive for X and Bob and Dave test positive for Y.  You can start to see the branches and the beginning of a tree.



   The first yDNA and mtDNA trees were built using only a few dozen SNPs.  Today, the paternal and maternal haplogroup trees are much more detailed, based on thousands of SNPs.  Complete SNP testing has been available for mitochondrial DNA for a number of years.  Starting last year, complete SNP testing is available for yDNA from companies like FamilyTreeDNA with their Big Y test.  Previously yDNA SNP tests were designed to look for specific SNPs.  With advances in technology, they can now look for all the SNPs across over 12 million yDNA base pairs.

   Just to add another acronym to the pile, there are also STRs or short tandem repeats (aka microsatellites).  STRs are short sequences of base pairs that repeat.  These repeats are found in autosomal, y and x DNA.  You may have heard the term CODIS if you watch Crime/Drama shows on television.  CODIS is the FBI’s Combined DNA Index System (more acronyms).  When DNA is collected for CODIS, they typically test for 13 STR markers across the autosomes.  When you have a yDNA STR test done, genetic genealogy companies test for up to 111 markers only on the Y chromosome.  They will also perform a basic SNP test to identify your paternal haplogroup.  SNPs and STRs are different in that SNPs appear to be permanent changes in our DNA and STRs are variable.  STRs are identified by location on the chromosome and by the number of times that the repeat occurs.  The number of repeats per STR can change over time, sometimes increasing, sometimes decreasing in number or increasing then decreasing again (known as a back mutation).  The combined set of STR markers is your haplotype and may be unique to your surname or span multiple surnames.  With the advances in yDNA SNP testing, SNPs will be found that are unique to your surname, which could make STR testing obsolete.

   We all have DNA: 23 chromosomes in our cell nuclei, half from mom and half from dad.  We also have mitochondrial DNA from our moms.  Less than 2% of our DNA is in the form of genes, which define who we are.  SNPs can be used to identify our “good” and “bad” genes.  SNPs can also help identify our ethnicity and build our paternal and maternal family trees.  STRs can organize us down to the paternal surname level.  When folks start talking DNA, don’t be afraid to question them about, “What kind of DNA?”, “What does that SNP indicate?” or “What type of STR is being tested?”.  We’ll never get away from using acronyms to simplify how we communicate genetic genealogy.  That doesn’t mean we need to let the acronyms simplify the meanings to a point where the science is lost.  Every little bit of knowledge adds to our understanding of ourselves.


© Michael Maglio

Wednesday, September 24, 2014

DNA Mysteries: Iberian R1b-V88 in Africa

   When I first heard about R1b in Africa, my immediate assumption was that the predominantly Celtic haplogroup must have been a recent transplant.  I ran some of the V88 haplotypes against the big databases (FTDNA & ySearch) expecting to see matches to European men within the African colonial timeframe.  It wasn’t that easy.  Common ancestor analysis put the R1b Africans (V88) thousands of years removed from the rest of their European R1b cousins.  Where did they come from?  How did they get there?


   I started with the given that the R1b defining mutations (SNPs) occurred in the Iberian Peninsula.  The jury is still out on this hypothesis.  There have been scientific papers for and against Iberian origins of R1b.  My own work (Iberian Origins of R1b) supports an origin prior to the Neolithic expansion.  Could V88 have made a straight-line migration from Iberia to the Lake Chad region of Africa?  Could V88 have crossed the Straits of Gibraltar, travelled across the Sahara, which 7,000 years ago was a savannah well populated with animals for hunting, and arrived at Lake Mega-Chad?  That was my early premise.  I was wrong.

   The distribution of V88 is much larger than any of the scientific papers would indicate.  While I agree with the work that’s been done correlating the spread of V88 with the spread of Chadic languages (Cruciani et al 2010), the Chadic population is only a subset.  Nobody takes into consideration the V88 populations in Europe and the Middle East.  If they do, it is a sideways glance to say were ignoring them because they don’t fit into what we are trying to prove.  If you don’t look at the entire picture, your conclusions will be skewed.

   I wanted the largest selection of V88 Y-DNA records with at least 37 markers tested.  I started with Family Tree DNA projects that had the records SNP tested.  Those haplotypes were run against the ySearch database to identify highly related records with no SNP testing.  The initial gathering of records picked up individuals with SNP M73.  These were removed.  The key differentiator between V88 and M73 was DYS464a&b.  V88 was typically 12,12 and M73 was 15,15.  Thirty-seven or more STR markers are helpful in identifying additional related haplotypes and even more necessary in determining the relationship between records.  Most studies only looks at SNPs or a small handful of STR markers.  This is shortsighted.  Imagine a reference population of 100 records all with the same SNP.  Without enough STR markers you can’t tell whether you are looking at one haplotype with minor 1 or 2 step variations or 100 unique haplotypes.  That’s the difference between a founder event starting with as few as one individual or a group with greater diversity and age.

   My final set of 119 records has at least 37 STR markers, V88 SNP testing or is highly related via STR and has the geographic location of the most distant known ancestor.  The records are processed through PHYLIP to generate a phylogenetic tree.  The phylogenetic tree give a visual depiction of the relationships in the dataset and an approximate number of years back to common ancestors, represented as the nodes between the records.


All of this is very standard genetic genealogy.  I add a twist (Biogeographical Multilateration) by converting the years back to a common ancestor to a distance using Cavalli-Sforza’s migration rate of 1 to 1.2 km per year.  This is enough for me to solve a series of cascading equations giving me the locations of the common ancestors.  Looking back at the phylogenetic tree shows us how all the nodes and locations are connected, essentially the flow of migration.


   The out of Iberia event took place about 7,700 ± 1,600 years ago.  TMRCA calculations have been shown to be very inconsistent.  Some folks use a constant mutation rate and some use rates per marker.  I include a TMRCA to give a relative chronology.  While the majority of R1b is known for its Western Atlantic migrations, V88 took a path along the Mediterranean coast and down the Adriatic.  While none of the V88 records indicated Crete as an ancestral location, it appears multiple times as a common ancestor location.  The data shows Crete as a stepping-stone in the Mediterranean as V88 migrated to the Nile River Valley.  The back to Africa event(s) occurred roughly 5,500 ± 1,000 years ago.


The majority of the Chadic records (Cameroon, Chad and Nigeria) have relatively close genetic connections to individuals in the Middle East (mainly Saudi Arabia).  The Chadic and Middle Eastern records tie back to common ancestors along the upper Nile.  There is a significant lack of information to understand what impact R1b-V88 had on the Nile Valley cultures.  Considering that there was only 1 out of 119 records with an exact Nile River location, I would venture a guess that V88 didn’t integrate well.

   While the V88 back to Africa migration has captured much attention, the data shows a more fascinating event.  There was a V88 re-migration back to Europe from Africa.   The back to Europe event took place about 3,200 ± 1,000 years ago.  Again, Crete played a role as a stepping-stone as V88 entered the Eastern Adriatic region and spread into Central and Eastern Europe.  Someone will probably notice that many of the V88 in Eastern Europe are Jewish and that the date for leaving the Nile region is close to the time of Exodus.  There is nothing in any of the data to indicate that this was the Jewish Exodus from Egypt.  The V88 group in Eastern Europe is closely related and there is phylogenetic evidence to support that this may have been a founder event with a single male or small group of closely related males.  There is no evidence to support that those founders were Jewish when they left Africa.


   By looking at the big picture, including all the data and letting the data illustrate the patterns, we can unravel what appears to be the mysterious appearance of R1b in Central Africa.  Along the way, we can uncover a previously unknown re-migration from Africa to Europe.  Too often haplogroup data is treated as discrete buckets of information living in a vacuum with no interaction to other haplogroups and no internal relationships.  Every DNA record is connected to every other record in a network.  Each haplotype is a vector with location and direction.  The sooner we treat genetic records as a network analysis, the sooner we will solve more DNA mysteries.

Out of Iberia and back to Africa.  Followed by a return to Europe.

Reference:

Maglio, MR (2014)  Y Chromosome Haplogroup R1b-V88: Biogeographical Evidence for an Iberian Origin (Link)


Thursday, February 20, 2014

Pushing the Boundaries

Tonight I made a guest appearance on Steve St. Clair's SinclairDNA BlogTalkRadio episode - Pushing the Boundaries of What Can Be Learned from Your DNA.


The episode was recorded live Friday, February 21 (2014).  You can listen to the recording through your computer by using this link - SinclairDNA

Steve and I talked about my recent experience with y-DNA research; Mapping y-DNA, William the Conqueror and a rare R1b group R-L11*.  We'll also talked a bit on the future of y-DNA testing.

Have a listen if you get a moment.  You can even download the show and take it with you.

Monday, February 10, 2014

The Third Brother: A Y-DNA Tale

   If we were to look at the Y-DNA family tree, we would see ancestors and descendants in a genetic sense. Haplogroup B is descended from A and C is descended from B. If we keep going, R is descended from P, etc. Within haplogroup R is SNP R-L11/P310 (R1b1a2a1a ISOGG 2014). There was a boy born somewhere between 3,000 and 10,000 years ago (there is much disagreement on the exact age). This boy was the first male to have this mutation on his Y-chromosome. He essentially became the ‘father’ of all R1b men in Western Europe.

   This R-L11 man had three sons, in the genetic sense, not in the literal sense. The first two sons are R-U106 and R-P312. Their stories are well known (at least in genetic genealogy circles). This is the story of the third brother, the one without a name. I’m going out on a limb in saying that this third branch exists as an independent unidentified SNP. R-DF100 has been identified as belonging to this third branch. Yet, it is too early to determine whether DF100 is the third brother or one of the many nephews (I had to keep the analogy going). Currently it is known as R-L11*/P310* (xU106,xP312), which means that folks on this branch test positive for having the L11 SNP and test negative for the U106 and P312 SNPs. Let’s call him R-x for simplicity. In case you were wondering, a SNP (single nucleotide polymorphism) is a mutation that can mark a branch point on your DNA.

Figure 1 – Three Brothers
   What do we know about R-x? They are a small group, only about 10% of the very large R1b population in Europe. They are still found in substantial numbers in Danelaw areas, the Netherlands, Pomerania, former Prussia and Denmark. U.S. President John Adams is one famous member of group R-x. A group of R-x descendants have created a site (http://www.worldfamilies.net/surnames/r1b1a2a1a) for those who are interested in tracing their family origins further back, have taken a y-DNA deep clade test and tested positive for L11 and negative for P312 / U106. 

   I was approached because of my work done on William the Conqueror’s DNA. The question was asked, what was the frequency of R-L11* (R-x) in the Conqueror study. All of the DNA records that made it into the final paper were R-L21*, which is downstream from R-P312. Unfortunately, for the R-x folks, that meant that no R-x records made it into the William the Conqueror modal haplotype.

   R-x was rare and it piqued my curiosity. I wanted to know how they fit into the bigger picture, where they came from and maybe connect them to a part of history. I’ve had some good success with geographical distribution of y-DNA data based on multiple distance measurements from reference positions (BGM). To start, I collected 26 R-x y-DNA records with close STR marker matches and known or probable SNP matches. Eight of these records were directly from the R1b1a2a1a website group data. The records were processed to determine time to most recent common ancestor (TMRCA). The neighbor-joining method was run on the results to create a phylogenetic tree.

Figure 2 – Phylogenetic Tree – R-L11*/P310* (xU106, xP312)

   Each of these records were picked because they also contained self-reported ancestral origins. The records were mapped based on these origins and a range calculated from the TMRCA was drawn as a radius representing distance to a common ancestor. See “Getting More” for additional details.

Figure 3 – Generalized Migration Flow – R-L11*/P310* (xU106,xP312)

   Migration direction is determined from phylogenetic connections. The orange arrows represent the primary migrations from the South Baltic region starting 2,000 years ago ± 200 years. The destinations for these migrations were into Scandinavia and along the Rhine River. The yellow arrows represent secondary migration events ending about 1,000 years ago. The results validate the R-x group’s origin locations (Pomerania, former Prussia and Denmark) and adds the Rhine River as a secondary origin. This is not the endgame. This just gets us 2,000 years into the past. Additional records need to be identified to push us back another 1,000 or so years. Where were the R-x ancestors before they were in the South Baltic?

   The third brother remains unnamed. Perhaps his name is R-DF100. The SNP hunters, those folks that are finding new SNPs every day, need more R-L11*/P310* (xU106,xP312) samples in order to identify a defining SNP. I’d also love to see better techniques of determining the age of a genetic branch. Someday we will know the name and the birthdate of the third brother.


Reference:
Maglio, MR (2014) Y-Chromosome Haplotype Origins via Biogeographical Multilateration (Link)

© MRMaglio 2014

Sunday, January 19, 2014

Pandora's DNA

Deep Into DNA*

   Ah, poor Pandora. So slandered. 

   As the story goes, Zeus gave her a jar and told her never to open it. Pandora’s curiosity got the better of her and she released all the ‘evils’ upon mankind. This is one of many origin stories for why there is evil in the world. This is also a metaphor for the spread or release of information. Too much knowledge can be ‘evil’.


   New knowledge discoveries are often also slandered. DNA test results are a modern example of disrupting the status quo.

...continued at The In-Depth Genealogist with a free membership.


*The Deep Into DNA article series is published each month in the new Going In-Depth
digital genealogy magazine presented by The In-Depth Genealogist.

#gDNA

Monday, November 26, 2012

Genetic Genealogy: Adding DNA to Your Toolkit


We’re all cousins! Genetic genealogy can tell us just how closely we are related.

Each of us already has a key to a library of knowledge about our ancestors, it’s in our DNA. Until recently, genealogists have relied on oral tradition and historical records. With DNA for genealogy, we now have a valuable new type of evidence.



Join us on November 28th at 6pm at the Boston Public Library as the Local & Family History Lecture Series hosts and I present Genetic Genealogy: Adding DNA to Your Toolkit.  We will look at the variety of testing options and the potential mysteries that they will unlock.  Learn about your deep ancestry, confirm your existing family history or break through brick walls in your genealogy research.

See how easy it is to add DNA to your genealogy toolkit.

Tuesday, November 20, 2012

Stephen Hopkins: Saxon DNA?


   As we approach Thanksgiving, it’s a great time to write about our Mayflower ancestors.  So far, I have found two on my wife's side, Stephen Hopkins and Stephen Hopkins.  Ok, that’s really just one, but I have two lines that trace back to him.   This isn’t unusual, estimates put the count of Stephen Hopkins’ descendants at about 2 million Americans.  

   What can Stephen Hopkins’ DNA tell us about his origins and his ancestors?  First, I should say that no one has a sample of Stephen’s DNA.  What we know about Stephen comes from tests completed by his male-line descendants with corroborating genealogical paper trails.  The Hopkins families are members of y-DNA haplogroup R1b, the largest genetic population in Europe.  R1b is often associated with the Celtic and Gallic tribes.  Hopkins’ DNA may be able to shed additional light on his birthplace, extend his genealogy further by tapping into an older family line or tell us about his deep ancestral origins.

   One of the first things I like to do is compare the haplotype, (the numeric markers from a y-DNA test) against a public database like ySearch.org.  The goal is to find other parallel lines of Hopkins with ancestry that predates Stephen.  This would allow us to work forward in time, connecting to Stephen and his father John, breaking through the current brick wall.  Unfortunately, no such records exist.

   What we do get from ySearch is list of genetic cousins and their ancestral locations.  Plotting these locations generates a distribution from Kent to Cornwall across southern England.  The highest concentration of cousins is in the historic Anglo-Saxon kingdom of Wessex.  The current research on Stephen Hopkins has him baptized in Hampshire, the heart of Wessex.
 
   What kind of R1b was Hopkins?  Was he a Celt, a Gaul, an Anglo-Saxon or something completely different?  One way to get close to the answer is to look at his genetic cousins again.  Since R1b is such a large group, it is important to focus on both the haplotype and SNP that defines his R1b subgroup.  The SNP that best defines Stephen is S493, which on the 2012 haplogroup tree is R1b1a2a1a1a2.  With the explosion of new SNPs identification and the rapidly expanding and changing subgroup nomenclature, researchers are advocating the use of the SNP rather than subgroup as a naming convention.  Let’s call Stephen Hopkins R-S493.

   When I take all these genetic cousins and run them through TribeMapper®, a pattern forms.  Ancestors start to pile up on either side of the English Channel and an approximate date of migration emerges.  Here’s where we pull out our history books.  If the date were about 2,500 years ago, I would say this was a Celtic migration.  If the date were 2,000 years ago, I might say these were Gaels fleeing the Romans.   The calculations come out to be about 1,500 years ago, putting this migration in line with the Anglo-Saxon invasion of Britain.

   Why stop there?  What flavor of Anglo-Saxon are we talking about?  Angle, Saxon, Jute?  The great thing about tribe mapping is that we can continuously turn back the clock and get a new picture.  If we find a Danish connection, then we might say Jutes or an association to the Angeln region of Germany, we could say Angles.  We have to be careful as those names and locations were just a snapshot in time when ancient historians catalogued Germanic tribes.  Those tribes, like all tribes, were just passing through.

   Stephen Hopkins’ DNA points to a genetic cluster in modern day Lithuania and Latvia.   This data most closely correlates to the Saxons and their origins on the Baltic coast.  Continuing this process gives us the following migration map.


   The R-S493 data takes us through Finland, Sweden and back to the mainland Europe to the Iberian Peninsula.  This puts the origin of R-S493 in Iberia about 4,000 years ago ± 500 years.

   We can’t be certain that Stephen Hopkins has Saxon DNA.  We can’t even say that all Saxons were haplogroup R1b.  It’s unlikely that they were a single homogenous ethnic group, but the core of the tribe would have had strong familial and genetic ties.  Were Hopkins’ ancestors at the core of this tribe or part of the fringe, picked up along the way?  A broader study of DNA associated with the same places and times would be required to answer that question.

   If we look at the surname Hopkins, its origins are from Hobbes-kin and even further back to the Germanic name Hrodberht.  Stephen Hopkins and his closest genetic cousins are found in the historic Kingdom of Wessex (West Saxons).  Time-wise, there is a correlation to the Anglo-Saxon invasion of Britain.  We can even make a connection to the proto-Saxons along the Baltic coast.  I’m going out on a limb and calling Hopkins a Saxon.

   That Saxon bloodline remained adventurous and served Stephen well as he voyaged to Bermuda, Jamestown and Plymouth colony.

   It’s never obvious where DNA will lead.  Each tribe mapping is an adventure in itself.


© Michael R. Maglio and OriginsDNA

Friday, June 8, 2012

DNA at the Genealogy Field Day

The Massachusetts Society of Genealogists presents -





When: Saturday, June 9, 2012, starting at 1:30PM

Where: Framingham Public Library, 49 Lexington St., Framingham

(free & open to the public - bring a friend)

Join Jeff Carpenter and Mike Maglio at the DNA table to talk about all your genetic genealogy questions.

Other Field Day topics include:
The 1940 Census
Scanning Demonstration - Flip Pal
Genealogy Mapping Tools
iGoogle
Scandinavian Research
Ask the Expert
Web Demonstrations
and more...



Tuesday, March 13, 2012

What’s in My gDNA Toolbox


   If I were talking about my regular genealogy toolbox, I would be listing links to all the great websites with digital records (e.g. FamilySearch).  I would also talk about great repositories like NARA, BPL or the Mass Archives.  Or, I would mention tips and techniques like Nearest Neighbor and the Hidden Treasures in old photos.



   Now that we are adding DNA as a tool for genealogy, we have to pack a new toolbox.

The Databases – record sources to compare your DNA against

· Ysearch.org – Y-DNA database
· Mitosearch.org – mtDNA database
· FTDNA.com – DNA Project database
· WorldFamilies.net – DNA Project database
· SMGF.org – DNA Project database

The Testing Companies – many different testing companies that are not all equal – do your homework

· FTDNA.com – DNA testing  (my favorite)
· 23andMe.com – DNA testing
· SMGF.org – DNA testing
· Ancestry.com – DNA testing
· GeneTree.com - DNA testing

Sources of gDNA Knowledge – There are many areas of genetic genealogy that are open for interpretation.  Read everything and come to your own conclusions.

· ISoGG.org – Advocates for the use of genetics as a tool for genealogical research
· Wikipedia  - Haplogroup details
· nationalgeographic.com/genographic

Analysis Tools – DNA results love to be compared and analyzed

· hprg.com/hapest5/index.html – Whit Athey’s Haplogroup predictor
· mymcgee.com/tools/ - Dean McGee’s Y-DNA comparison tools
· www.math.mun.ca/~dapike/FF23utils/ - David Pike’s autosomal comparison tools
· http://gedmatch.com/ - Autosomal comparison tools
· PHYLIP – phylogenetic tree creation

DNA Data Management – you need to organize and manage your DNA records

· Legacy Family Tree – supports DNA records (the one I use)
· Family Tree Maker, RootsMagic, Ancestral Quest and The Master Genealogist – supports DNA
· Excel – spreadsheet tools
Misc
· Google Maps – User defined maps – you never know when you might want to build your own custom map

   This is hardly an exhaustive list.  I use most of these tools on a weekly basis.  I’m always looking for new tools (or creating ones that don’t exist).

   What's in your toolbox?  Let me know what tools you are using.

#gDNA

Monday, December 5, 2011

Migration Mapping: Eldred the Terrible

   Genetic genealogy has been very good at identifying distant origins and for making connections along paternal and maternal lines going back a half dozen centuries.  What seems to be missing is how we got from point A to point B.

'Eldridge' clan mapping

   At some distant place in time in every genealogy the surname becomes irrelevant.  The only way to go further back is to use DNA testing.  We have to rely on Clans and Tribes, genetically related groups of individuals, to get an understanding of our history.

   Pride in your historic nationality is wonderful and can tell you much about your family, but we are all descendants of nomads.  As nomads we belong to ancient cultures just as much as we belong to any one nationality.  To know what culture you are you need to know where your tribe was and when.

   When I had my DNA tested I learned that I was part of haplogroup G with origins in the Caucasus Mountains going back about 22,000 years.  I also learned that I had no close matches in the last few centuries.  That left me with very little to work with. So, I put on my analyst hat and developed a technique for plotting the migration path of my tribe at different periods in history.  I needed to answer how my people got from the Caucasus to a little village outside of Naples, Italy.

   I knew I had hit on something after my first mapping exercise.

'Maglio' clan mapping

   The individuals that I plotted lined up along the Rhine River and down the Apennines (with a few stragglers in Wales).  Successive maps, each going back further in time, showed a pattern along the Danube and around the Black Sea back to the Caucasus Mountains.  I now have my migration answers and a plausible correlation to the Etruscan metalworking culture.

   I have been using my technique to help my clients get a deeper understanding of their history and their culture.  For all of you with the surname Eldridge, Eldredge, Aldrich and variation, I have posted a sample report on my website - "The Genetic Genealogy of Eldridge"  

   I'd love to hear about other successes mapping genetic data across time.