Order statistics | What are order statistics?

Contents

Introduction

Order statistics is a very useful concept in statistical science. They have a wide range of applications including auction modeling, auto racing and insurance policies, optimization of production processes, estimation of parameters of distributions, et al. Through this article, we will understand the idea of ​​order statistics. We will first understand its meaning and gradually proceed to its distribution., eventually covering more advanced concepts.

Suppose we have a set of random variables X1, X2, …, XNorth, that are independent and identically distributed (iid). For independence, we mean that the value taken by a variable random variable is not influenced by the values taken by other random variables. By identical distribution, we mean that the probability density function (PDF) (or equivalently, the cumulative distribution function, CDF) for random variables it is the same. The Kth The order statistic for this set of random variables is defined as kth smallest sample value.

To better understand this concept, we will take 5 random variables X1, X2, X3, X4, X5. We will observe a realization / random result of the distribution of each of these random variables. Suppose we obtain the following values:

14444capture01-min-8688645

The Kth the order statistic for this experiment is kth smallest value of the set {4, 2, 7, 11, 5}. Then, the 1S t the order statistic is 2 (smallest value), the 2North Dakota the order statistic is 4 (the next smallest), and so on. The 5th the order statistic is the fifth smallest value (the greatest value), What is it 11. We repeat this process many times, namely, we extract samples from the distribution of each of these iid random variables and find the kth smallest value for each set of observations. The probability distribution of these values ​​gives the distribution of kth order statistics.

In general, if we order random variables X1, X2, …, XNorth in ascending order, then the kth the order statistic is displayed as:

58274capture02-min-7203541

The general notation of the kth the order statistic is X(k). Note X(k) is different from Xk. Xk is the kth random variable of our set, while X(k) is the kth statistical order of our set. X(k) takes the value of Xk yes Xk is the kth Random variable when realizations are sorted in ascending order.

The 1S t X-order statistic(1) is the set of minimum values ​​of the realization of the set of 'n’ random variables. Laterth X-order statistic(North) is the set of maximum values (nth minimum values) of the realization of the set of 'n’ random variables. They can be expressed as:

11926capture03-min-4227227

Order statistics distribution

Now we will try to find out the distribution of the order statistics. We will first describe the distribution of the nth order statistics, then he 1S t statistical order and finally the kth general order statistics.

A) Distribution of nth Order statistics:

Let the probability density function (PDF) and the cumulative distribution function (CDF) our random variables let fX(x) y FX(x) respectively. By definition of CDF,

82261capture04-min-8907724

Since our random variables are identically distributed, have the same PDF fX(x) y CDF FX(X). Now we will calculate the CDF of nth order statistics (FNorth(x)) as follows:

91898capture05-min-1932533

Random variables X1, X2, …, XNorth they are also independent. Therefore, by property of independence,

71167capture06-min-7598051

The PDF of the nth statistical order (fNorth(x)) is calculated as follows:

82826capture07-min-4983818

Therefore, the expression for PDF and CDF of nth The order statistic has been obtained.

B) Distribution of 1S t Order statistics:

The CDF of a random variable can also be calculated as the one minus the probability that the random variable X takes a value greater than or equal to x. Mathematically,

47335capture08-min-6746054

We will determine the CDF of 1S t order statistics (F1(x)) as follows:

51217capture09-min-4314011

One more time, using the independence property of random variables,

56291capture10-min-6261394

The PDF of the 1S t statistical order (f1(x)) is calculated as follows:

92517capture11-min-5947183

Therefore, the expression for PDF and CDF of 1S t The order statistic has been obtained.

C) Distribution of the kth Order statistics:

Forkth order statistics, in general, the following equation describes your CDF (Fk(X)):

61621capture12-min-2601707

The PDF of kth statistical order (fk(x)) is expressed as:

63589capture13-min-3525477

To prevent confusions, we will use geometric proofs to understand the equation. As discussed above, the set of random variables has the same PDF (fX(X)). The graphic below shows a sample PDF with the kth Order statistic obtained from random sampling:

58258capture14-min-2854158

Then, the PDF of the random variables fX(x) is defined between the interval [a,b]. The k-th order statistic for a random sample is shown with the red line. The other variable realizations (for the random sample) shown by the small black lines on the x-axis.

There are exactly (k – 1) observations of random variables that fall in the yellow region of the graph (the region between & kth order statistics). The probability that a particular observation falls into this region is given by the CDF of the random variables (FX(X)). But we are aware that (k – 1) observations fell in the region, what gives us the term (for independence) (FX(X))(k – 1).

There are exactly (n – k) observations of random variables that fall in the blue region of the graph (the region between kth statistical order & b). The probability that a particular observation will fall in this region is given by the 1 – CDF of the random variables (1– FX(X)). But we are aware that (n – k) observations fell in the region, what gives us the term (for independence) (1-FX(X))(n – k).

Finally, exactly 1 observation falls exactly on the k-th order statistic with probability fX(X). Therefore, the product of 3 terms gives us an idea of ​​the geometric meaning of the equation for PDF of the k-th order statistic. But, Where does the factorial term come from? The above scenario only showed one of the many orderings. There can be many of these combinations. The total number of such combinations is shown below:

56727capture15-min-5188837

Therefore, the product of all these terms gives us the general distribution of kth order statistics.

Useful functions of order statistics

Order statistics lead to several useful functions. Among them, notable ones include the sample range and the median of the sample.

1) Sample range: It is defined as the difference between the largest and smallest value. It is expressed as follows:

67901capture16-min-3026387

2) Sample median: The sample median divides the random sample (realizations of the set of random variables) in two halves, one containing the lowest value samples and the other containing the highest value samples. It's like the middle order statistic / central. It is mathematically defined as:

54935capture17-min-7664566

Order statistics set PDF

A joint probability density function can help us better understand the relationship between two random variables (two-order statistics
in our case). The joint PDF for any statistics of 2 X orders(a) & X(B), such that 1 ≤ a ≤ b ≤ n is given by the following equation:

57960capture18-min-9865168

Example

We will use a very simple example to illustrate the distribution of the order statistics: the standard uniform distribution (U[0, 1] distribution). We will take 5 random variables X1, X2, X3, X4, X5, everyone has the U[0, 1] distribution. For this set of random variables, we will calculate and plot the 1S t, 3rd (the sample median) Y 5th (Northth) order statistics. The following figure shows the U[0, 1] distribution:

14549capture19-min-7605626

We will draw random samples as follows and find the 1S t, 3rd & 5th order statistic for each sample. Below are two of the samples:

84551capture20-min-5655291

The standard uniform distribution PDF and CDF are given as:

40491capture21-min-6236281

We will use this information and calculate X(1), X(3) & X(5) using the formulas we derived. We will take the case only when x is between 0 Y 1 (for other cases, the order statistic is zero since PDF is zero).

A) To 1S t order statistics:

88253capture22-min-5826320

Plot for f1(X):

97189capture23-min-5582206

B) To 3rd order statistics:

36859capture24-min-6027432

Plot for f5(X):

88539capture25-min-1884985

C) To 5th order statistics:

69523capture26-min-4975531

Plot for f5(X):

94366capture27-min-6960818

Conclution

Therefore, we have explored the concepts of order statistics in depth. A wide range of physical processes can be modeled through order statistics, exploiting its properties, particularly their distributions.

The media shown in this article is not the property of DataPeaker and is used at the author's discretion.

Subscribe to our Newsletter

We will not send you SPAM mail. We hate it as much as you.

Datapeaker