Ewens's sampling formula

From Wikipedia, the free encyclopedia

(Redirected from Ewens distribution)
Jump to: navigation, search

In population genetics, Ewens's sampling formula, introduced by Warren Ewens, states that under certain conditions (specified below), if a random sample of n gametes is taken from a population and classified according to the gene at a particular locus then the probability that there are a1 alleles represented once in the sample, and a2 alleles represented twice, and so on, is

\operatorname{Pr}(a_1,\dots,a_n)={n! \over \theta(\theta+1)\cdots(\theta+n-1)}\prod_{j=1}^n{\theta^{a_j} \over j^{a_j} a_j!},

for some positive number θ, whenever a1, ..., an is a sequence of nonnegative integers such that

a_1+2a_2+3a_3+\cdots+na_n=n.\,

The phrase "under certain conditions", used above, must of course be made precise. The assumptions are (1) the sample size n is small by comparison to the size of the whole population, and (2) the population is in statistical equilibrium under mutation and genetic drift and the role of selection at the locus in question is negligible, and (3) every mutant allele is novel. (See also idealised population.)

This is a probability distribution on the set of all partitions of the integer n. Among probabilists and statisticians it is often called the Ewens distribution.

When θ = 0, the probability is 1 that all n genes are the same. When θ = 1, then the distribution is precisely that of the integer partition induced by a uniformly distributed random permutation. As θ → ∞, the probability that no two of the n genes are the same approaches 1.

This family of probability distributions enjoys the property that if after the sample of n is taken, m of the n gametes are chosen without replacement, then the resulting probability distribution on the set of all partitions of the smaller integer m is just what the formula above would give if m were put in place of n.

The Ewens distribution arises naturally from the Chinese restaurant process.

  • Warren Ewens, "The sampling theory of selectively neutral alleles", Theoretical Population Biology, volume 3, pages 87—112, 1972.
  • J.F.C. Kingman, "Random partitions in population genetics", Proceedings of the Royal Society of London, Series B, Mathematical and Physical Sciences, volume 361, number 1704, 1978.
  • S. Tavare and W. J. Ewens, "The Ewens sampling formula". In Multivariate discrete distributions by N.L. Johnson, S. Kotz, and N. Balakrishnan (eds), 1997, Wiley.

Advanced Search
Included Web Search Engines


Safe Search

close

Top Matching Results

Occasionally Search.com will highlight specialized results that are based on the context of your query. Examples of specialized results include specific links to news, images, or video.

Top Matching Results may highlight information from other Search.com pages, content from the CNET Network of sites, or third party content. The listings are based purely on relevance. Search.com does not receive payment for listings in this section but our partners that provide this data may get paid for listing these products.

Sponsored Links

This section contains paid listings which have been purchased by companies that want to have their sites appear for specific search terms and related content. These listings are administered, sorted and maintained by a third party and are not endorsed by Search.com.

Search Results

Search.com sends your search query to several search engines at one time and integrates the results into one list which has been sorted by relevance using Search.com's proprietary algorithm. You can customize the list of search engines included in your metasearch from the preferences.

The search engines that are used in your metasearch may allow companies to pay to have their Web sites included within the results. To view the Paid Inclusion policy for a specific search engine, please visit their Web site. Search.com does not accept payment or share revenue with any search engine partner for listings in this section.