Outlier Detection Procedures in a Sample from a Gumbel Distribution with Unknown Scale Parameter
Published On December 5, 2023
Journal Issue LJRS Volume 23 Issue 19

Outlier Detection Procedures in a Sample from a Gumbel Distribution with Unknown Scale Parameter

Dr. Pratyasha Tripathi
Dr. Pratyasha Tripathi
Outlier Detection Procedures in a Sample from a Gumbel Distribution with Unknown Scale Parameter
Article Fingerprint
Research ID 750GC

IntelliPaper

Abstract

Lalitha and Tripathi (2018) have suggested a test statistic for the detection of a pair of outliers in a sample from a Gumbel distribution with known scale parameter σ. The statistic is not found to be suitable while dealing with the case of unknown scale parameter σ. Thus, in this paper, the test statistic suggested by Lalitha and Tripathi (2018), is suitably modified by using the modified moment estimator of the scale parameter σ for the detection of two single (upper/lower) outlying observations and one more test statistic is suggested for the detection of a pair of outlying observations. Their critical values and performance probabilities are obtained at different levels of significance by a simulation technique.

Explore Digital Article Text

I. INTRODUCTION

Lalitha and Tripathi (2018) have discussed detection of a single upper and lower outlier when the location parameter and the scale parameter of Gumbel distribution were assumed to be known. But if the scale parameter is not known then in such situation, previously discussed procedures should be suitably modified. Thus, when the scale parameter is not known, an estimator of the scale parameter is used in the test statistic. Then it's critical values and the corresponding performance study were done by simulation technique. Since the study is about outlying observation, the entire sample should not be considered for the estimation of the scale parameter. Hence, the best linear unbiased estimate suggested by Balakrishnan and Cohen (1991) for a type II censored sample was used as an estimator for the scale parameter. But with this estimator, the values of the tabulated coefficients were available only for samples of size at most 10. Hence, a modified form of the moment estimator is considered and the statistics were studied. But on using an estimator for the scale parameter, derivation of the theoretical probability distribution of the test statistic is extremely tedious. Hence, the critical values as well as the performance of the test statistic were obtained by using simulation technique.

II. TEST STATISTIC USING THE BEST LINEAR UNBIASED ESTIMATOR FOR THE SCALE PARAMETER

Three test statistic , and have been suggested to detect an upper outlier, lower outlier and a pair of outlying observations respectively to test the null hypothesis against the slippage alternative i.e. there is one or a pair of observation(s) from another Gumbel with a shifted scale parameter , where the scale parameter was unknown. Hence for , the best linear unbiased estimator suggested by Balakrishnan and Cohen (1991) given as , where r and s denotes number of trimmed observations from lower and upper side respectively and 's are known constants obtained by Balakrishnan and Chan (1992), was used. Since values of 's obtained by Balakrishnan and Chan (1992) up to sample size 30 but only up to 10 values are available, therefore the use of the test statistic, suggested in this case is restricted up to sample size 10. The test statistics for an upper, lower and a pair of observations obtained respectively for this case are as follows.

\[Z _ {1} ^ {\prime} = \frac{x _ {(n)} - x _ {(n - 1)}}{\sigma^ {*}}, Z _ {2} ^ {\prime} = \frac{x _ {(2)} - x _ {(1)}}{\sigma^ {*}}, Z _ {3} ^ {\prime} = \frac{x _ {(n)} - x _ {(1)}}{\sigma^ {*}}\]

where , , and are first, and order statistics respectively arranged in an ascending order of magnitude. However, because of the limitations in obtaining the values of 's, these statistics have only limited usage and hence a modified form of these statistics is suggested in this paper.

2.1 Modified Test Statistics using a moment estimator for the scale parameter

The moment estimator of was suggested by Johnson et.al. (1994) and is given as , where is sample standard deviation, can be used as an efficient estimator of . This is because Johnson et.al. (1994) have shown that the moment estimator is about more efficient than Cramer-Rao lower bound estimator of scale parameter. Further, as our work is concerned with the extreme events and the sample standard deviation is affected by the extreme observations, therefore to make the test statistic more efficient, the modified sample standard deviation obtained from a trimmed sample (i.e. a sample obtained after deleting the extreme observations and ) was used instead of a complete sample standard deviation. Thus, in this case, scale parameter was replaced by its modified moment estimator, given as , where obtained from a trimmed sample (i.e. a sample obtained after deleting the extreme observations and ). The test statistics so obtained are as follows

, where , and are first, (n-1)th and nth order statistics respectively arranged in an ascending order of magnitude. These three test statistics can be applied to a sample of size n=3, 4,..., 50. Performances of these three statistics were studied by simulation technique using 10,000 replications.

III. CRITICAL VALUES FOR THE TEST STATISTIC Z1 TO DETECT AN UPPER OUTLIER

The test statistic was used to detect an upper outlying observation in a sample from Gumbel distribution. In the detection of an upper outlying observation, the null hypothesis would state that there is no outlying observation in the sample. As the statistic is based on the difference of the largest and second largest observations, this test statistic should reject the null hypothesis for large values of . Thus an -level critical region will be given as where can be obtained from , where . Critical values of the test statistic for case-I and case-II of the Gumbel distribution were obtained using simulation technique with 10,000 replications which are tabulated in Table 3.1. and Table.3.2. respectively for at 1%, 5% and 10% significance levels. Just as the critical values obtained for a single upper, single lower and a pair of outliers with known scale parameter are free from the scale parameter, calculation of these critical values are also free from the scale parameters.

Table 3.1: Critical values of for case-I at different levels of significance.

$Z_{\alpha}$
n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$ n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$
34.5716433.0873162.347163114.4243063.0549172.289501
44.9978843.1712702.567949124.4103293.0228912.254245
54.8284513.1702632.382530134.4087353.0640202.237309
64.7996593.1464202.358684144.4016373.1204712.224776
74.7442533.1379732.341062154.3586553.1096702.209747
84.7309473.1203502.315699204.3744672.9878432.126154
94.6958623.1112722.313943304.7676223.2482942.304326
104.5628873.0993252.304717504.8600723.2949212.579111

Table 3.2. Critical values of for case-II at different levels of significance

$Z_{\alpha}$
n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$ n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$
32.9163632.1225241.794401111.37001301.03592460.8400556
42.9163632.1225241.794401121.38654161.04216080.8221364
52.5327881.8061031.495726131.29512160.97141100.7671696
62.1816851.5963651.277108141.19044430.95825320.7770338
71.7834071.3499301.136774151.28761760.88980450.7190456
81.6372241.1937691.004956201.13197710.80372150.6639349
91.43752881.07548330.9080605301.03537730.74090270.6022329
101.44796351.05170480.8511008500.88228680.64060250.5248657

IV. CRITICAL VALUES FOR THE TEST STATISTIC TO DETECT A LOWER OUTLIER

It was observed in Lalitha and Tripathi (2018) that, for Gumbel distribution, the density functions on the real line of case-I and case-II of lower outlier are the mirror images of the density functions of the case-II and case-I of upper outlier respectively. Thus for the test statistic , critical values obtained in section 4 for the upper outlier of case-I and case-II, can be used for detection of the lower outlier of case-II and case-I respectively.

V. THE CRITICAL VALUES FOR THE TEST STATISTIC TO DETECT A PAIR OF OUTLIERS

The critical values of the test statistic were obtained by using the simulation technique with 10,000 replications and are given in Table 5.1. The critical values of the test statistic (unknown scale parameter) for case-I and case-II of the Gumbel are close to the critical values obtained by Lalitha and Tripathi (2018) for the Gumbel distribution case-I and case-II (known scale parameter).

Table 5.1: Critical values of for case-I at different levels of significance.

$Z_{\alpha}$
n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$ n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$
36.4176284.9637214.173418117.9805356.4533465.588609
46.7283165.0209104.240807128.0721916.4819305.838556
56.9615245.2068544.584614138.1014286.7510446.060688
67.2206865.6140504.877048148.2600556.7638566.063707
77.3603215.8028355.092647158.3091066.8274936.055398
87.4877796.0617405.298490208.6770327.2647506.477429
97.5813686.0973975.398235308.9434157.5910827.051505
107.8686126.1956655.584236509.9983348.2879667.574812

VI. WHEN THE OBSERVATION OF A SAMPLE IS ON EITHER SIDE OF THE LOCATION PARAMETER

In this case it is assumed that some of the observations are below and some are above the location parameter. In this section the moment estimator of , suggested by Johnson et. al. (1994) is used in place of the scale parameter . The detection of an upper and a lower outlying observation can be done as discussed in section 4 and section 5 respectively. But while dealing with a pair of outlying observations the procedure would be slightly different and it is given as follows.

Let be a random sample from a Gumbel distribution with location parameter and scale parameter (unknown), in which some, say m, observations, are less than the location parameter and rest of the observations of the sample are greater than . Then considering the m observations lying below the location parameter, a modified test statistic can be defined for detecting the smallest observation by considering largest observation as , i.e.

\[\frac{X_{(m)} - X_{(1)}}{\frac{\sqrt{6}}{\pi} s^{*}}, m = 2, \dots, n - 1.\]

As before, the observation corresponding to cannot be declared as an outlying one, being the observation lying closest to the location parameter, when the test statistic falls in the critical region. Thus, in the event of rejection of the null hypothesis, only the smallest observation i.e. should be considered as outlying observation. In this case, the critical values given in case II of section 4 should be used for testing the null hypothesis with a sample size as m. For the rest, observations above the location parameter, a modified test statistic can be defined for detecting the largest observation by considering largest observation as , i.e.

\[Z = \frac{X_{(n)} - X_{(m+1)}}{\frac{\sqrt{6}}{\pi} s^{*}}, m = 1, \dots , n - 2.\]

Here again, the observation corresponding to would be lying closest to the location parameter and therefore this observation cannot be declared as an outlying one. In this case the critical values obtained in case I of section 4 should be used for testing the null hypothesis with a suitable modification of the sample size as . As before, in the event of rejection of the null hypothesis, the largest observation i.e. should be declared as outlying observation.

a. Case when only one observation is on either side of the location parameter and the location parameter is also known (i). When only one observation is lying below the location parameter which is assumed to be known, while all other observations are above the location parameter, then the statistic Z can be modified as

\[Z = \frac{\mu - X_{(1)}}{\frac{\sqrt{6}}{\pi} s^{*}} .\]

Here, if it is assumed that and follows a Gumbel distribution as defined in case II with location and scale parameters and respectively. The critical values can be obtained by using simulation technique with 10,000 replicates. These critical values were tabulated in Table 6.1 for different values of sample size n and different levels of significance, given as follows.

Table 6.1: The critical values of the test statistic when only one observation is lying below the location parameter (known)

$Z_{\alpha}$
n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$ n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$
51.802870.924730.65052120.652250.450470.40411
61.169720.642280.48095130.642220.446030.41108
71.092300.580270.44450140.686540.450070.41784
80.886340.551030.42177150.736760.477750.42694
90.738020.492500.39991200.633600.489890.46042
100.724270.476920.93198300.616970.543200.51112
110.683660.437310.38694500.683110.603970.55898

(ii). When only one observation is lying above the location parameter, while all other observations are below the location parameter, then the statistic Z is modified as

\[Z = \frac{X _ {(n)} - \mu}{\frac{\sqrt{6}}{\pi} s ^ {*}} .\]

Here, as it is assumed that , follows a Gumbel distribution as defined in case I with location and scale parameters and respectively. The critical values can be obtained by using simulation technique with 10,000 replications and were tabulated in Table 6.2., as given below

Table 6.2: The critical values of the test statistic when only one observation is lying above the location parameter (known)

$Z_{\alpha}$
n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$ n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$
51.818570.926330.60187120.650550.460360.40235
61.25745 $0.6888^0$ 0.49004130.677750.445930.40829
71.048260.585070.45963140.666290.458120.42150
80.944770.569330.41839150.675640.466580.42315
90.916800.480560.38863200.556920.486120.45476
100.752540.491480.40449300.605390.544560.50122
110.743740.498460.40382500.699220.608990.56216

b. Case when the location parameter is unknown

When the location parameter is unknown, it can be estimated with a trimmed sample, as suggested in section 3 and is denoted by . As the two extreme observations are suspected outliers therefore the sample should be trimmed at both the ends. The location parameter used in the test statistics given in equations (6.1) and (6.2), is replaced by this estimate . With this estimate, the number of observations on its left and right, i.e. and can be decided. Then for and/ , the above said procedures can be used, as their critical values are independent of both the location and scale parameter. Also when and/(n - m) = 1, the test statistic will be as given above in section 6(a), with the value of the location parameter replaced by its estimator . The critical values so obtained were tabulate in Table 6.3. and Table 6.4. for the two cases of the distribution i.e. when only one observation is below the location parameter and when only one observation is above the location parameter respectively, are as given below.

Table 6.3: The critical values of the test statistic when only one observation is lying below the location parameter (unknown)

$Z_{\alpha}$
n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$ n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$
51.669080.687640.47271120.616360.385360.33359
61.352440.654020.44787130.578810.408410.35085
71.005320.540350.38811140.651510.412130.36337
81.028010.473120.36072150.674390.391630.36625
90.758870.455920.33152200.570110.422480.39505
100.874990.463450.33778300.563840.479500.44343
110.708890.450180.33154500.642580.548030.50708

Table 6.4: The critical values of the test statistic when only one observation is lying above the location parameter (unknown) Here the critical values were obtained up to sample size 50, therefore the suggested test statistic, based on the trimmed estimate of the scale parameter can be used for large samples as well.

$Z_{\alpha}$
n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$ n $\alpha = 0.01$ $\alpha = 0.05$ $\alpha = 0.10$
52.376670.960490.62846120.706810.403610.33457
61.431590.633150.63315130.585280.410910.34473
70.946070.427590.33917140.673700.399650.35868
80.775050.427590.33917150.628770.395520.36069
90.790990.444930.33463200.566410.440570.40532
100.690590.421490.31121300.550850.483150.44851
110.703270.395490.32351500.628770.549540.50829

VII. PERFORMANCE STUDY FOR A SINGLE UPPER OUTLIER

Consider a set of observations which contains a single contaminant observation and which comes from a population that has a different distribution from the rest of the observations. The power of the test can be defined as the sample contains one contaminant observation], i.e. the probability of the largest observation of the sample is being identified as discordant. But it is not necessary that the test always identifies as a discordant or defining only power of a test statistic is not sufficient to describe performance of a test statistic. Barnett and Lewis (1994) have suggested probabilities, to as test performance criteria. To improve the performance criteria of a statistic, David and Nagaraja (2003) discussed the properties of five such probabilities, labeled as to as a reasonable measure of the performance of . These probabilities are defined as

, is the probability the observation tested by the test statistic is identified as the outlying one, when it is known that there is one outlier is present.

, is the probability the contaminant observation tested by the test statistic is identified as the outlying one, when it is known that there is one outlier is present.

, is the probability the contaminant observation is the extreme observation tested by the test statistic which is identified as the outlying one, when it is known that there is one outlier is present.

is the probability that the largest observation is the contaminant observation, which is being detected as discordant while the second largest is not a discordant observation; and

, is the conditional probability that when it is given that the contaminant is an outlier and it is also identified as discordant by the test, where is a general test statistic and is the corresponding value for the contaminant observation. David and Nagaraja (2003) have observed that and the information given by and are seen to be limited. Thus, only and are sufficient for defining performance of any test statistic. Further, Hayes and Kinsella (2003) discussed the performance of the discordancy test on the basis of six performance criteria and called them as non spurious power, spurious power, swamping effect, spurious Type II error, partially spurious Type II error and nonspurious Type II error. Also, it was suggested that a good discordancy test should have high value of non spurious power, low value of spurious power and low value of swamping effect. They recommended that out of six performance criteria, probability of the non-spurious power , probability of the spurious power and the probability of the non-spurious type-II error (which gives the probability that the test wrongly identifies a good observation as discordant), are required to specify the test completely. In accordance with the a...

In this section, study of the performance for detection of a lower outlier for the both the cases of Gumbel distribution with location parameter and scale parameter (unknown) are discussed. Case-I: When all observations are greater than the location parameter . The performance of the test statistic to detect an upper outlier in a sample from a Gumbel distribution, after introducing a contaminant observation from another sample from the same distribution with different scale parameter, was observed. This was replicated 10,000 times. From this, the values of the probabilities and for the test statistic were obtained and are given in Table 7.1. It can be observed that at 1% level of significance, the power , the non-spurious power and the conditional power of the test statistic are showing almost negligible changes between sample size 5 and 10, between sample sizes 10 and 15 a rapid increase is observed while beyond 15 the rate of increase is comparatively low. The non-spurious type-II error is also showing almost no change between sample size 5 and 10, between sample sizes 10 and 15 it is increasing while beyond 15 it decreases. The ratio i.e. the probability that the contaminant shows up as an outlier, is almost constant between sample sizes 5 and 10, between sample sizes 10 and 15 it increases with very high rate of increase and beyond 15 rate of increase is comparatively low. Thus, it can be interpreted that the performance of the test statistic is increasing with sample size but beyond sample size 20 there is hardly any variation.

Table 7.1: Performance of the test statistic for case-I at different levels of significance

nLevel of Significance $P_1$ $P_3$ $P_1 - P_3$ $P_5$ $P_3/P_5$
51%0.514900.424070.090830.556100.76258
5%0.653600.536360.117240.700500.76568
10%0.738900.709820.029080.869900.81598
1%0.514900.424070.090830.535100.79251
105%0.908800.846940.061860.900500.94052
10%0.948600.916990.031610.940400.97511
151%0.940090.756530.183560.835500.90548
5%0.976900.922380.054520.949800.97113
10%0.990200.956520.033680.969500.98661
201%0.950900.895120.055780.933300.95909
5%0.992000.956670.035330.969900.98636
10%0.998200.995000.003200.998400.99659

It can be seen that at 5% level of significance the power, the non spurious power, the conditional power and the probability that the contaminant is identified as an outlier all are increasing rapidly between sample sizes 5 and 10 and beyond 10 its rate of increase is comparatively low. The values of all these probabilities are very high for sample sizes up to 15 while beyond 15 it shows almost no fluctuation. The non spurious type-II error is decreasing very fast between sample sizes 5 and 10 while beyond that it decreases with comparatively low rate of decrease. Hence from inferences discussed above it can be concluded that the performance of the test statistic is good at 5% level of significance. It can be noticed from. at 10% level of significance the power, the non spurious power, the conditional power and the probability of showing up a contaminant as an outlier are very high and increasing rapidly between sample sizes 5 and 10 and beyond that its rate of increase is very low. The non spurious type-II error is increasing with a very low rate of increase with the increase of the sample size up to 15 and beyond sample size 15 it decreases slowly. Hence it can be seen that the calculated probabilities and are high and i.e. non spurious error is low, as desired. Thus, on the basis of above results it can be interpreted that the performance is increasing with sample size up to a certain level and after that the variations are almost uniform, also the performance of the statistic is found to be good for the detection of an upper outlier in a Gumbel sample at 5% and 10% levels of significance, while at 1% level it is comparatively low.

Case-II: When all observations were smaller than the location parameter . The performance of the test statistic to detect an upper outlier in a sample from a Gumbel distribution, after introducing a contaminant observation from another sample of the same distribution with different scale parameter was obtained. These performance probabilities were obtained by simulation technique with 10,000 replicates and are given in

Table 7.2: Performance of the test statistic for case-II at different levels of significance

nLevel of Significance $P_1$ $P_3$ $P_1 - P_3$ $P_5$ $P_3/P_5$
51%0.37870.34880.02990.56050.6222
5%0.52840.49180.03660.70170.7009
10%0.60740.54070.06670.76380.7079
101%0.87160.80710.06450.89610.9007
5%0.89130.83580.05540.90290.9257
10%0.94100.90900.03190.92630.9813
151%0.98400.93290.05110.94840.9837
5%0.99120.94470.04640.95830.9858
10%0.99440.97620.01820.98740.9887
201%0.99830.95000.04830.97290.9765
5%0.99710.96910.02790.98880.9801
10%0.99910.98660.01250.99120.9953

It was observed from Table 7.2. that the performance of the test statistic under consideration was found to be good (as the performance is increasing with the increase of the sample size and the calculated probabilities and are high and is low, as desired). Also, it can be seen that at 1% level of significance the power, the nonspurious power, the conditional power and the probability of identifying a contaminant as an outlier, are increasing rapidly between sample sizes 5 and 10, beyond that the changes are almost negligible. The non-spurious type-II error is increasing with a small rate of increase at 1% and 5% levels while at 10% level of significance it is decreasing with a very small rate of decrease. Thus, it can be interpreted that the performance of the statistic is increasing with the sample size up to 15 and beyond that almost negligible changes are observed.

VII. PERFORMANCE STUDY FOR A SINGLE LOWER OUTLIER

In this section the performance of was studied for both the cases of the Gumbel distribution. Case-I: When all observations were greater than the location parameter . The performance of the test statistic to detect a lower outlier in a Gumbel sample, after introducing a contaminant observation from another sample of the same distribution with different scale parameter, was studied by simulation technique with 10,000 replicates. These performance probabilities are given in Table 8.1. From this table, it can be observed that the performance is increasing with the increase of the sample size and it is found to be satisfactory beyond sample size 15.

Table 8.1: Performance of the test statistic for case-I at different levels of significanc

nLevel of Significance $P_{1}$ $P_{3}$ $P_{1}-P_{3}$ $P_{5}$ $P_{3}/P_{5}$
51%0.57040.51240.05800.70770.7240
5%0.65320.63310.02010.72430.8741
10%0.72140.68650.03490.74560.9207
101%0.73600.67910.05690.92760.7321
5%0.82220.79550.02670.92190.8629
10%0.93810.90870.02940.94780.9587
151%0.81390.7560.05790.9910.7629
5%0.96840.93590.03250.99560.9400
10%0.98420.9550.02920.99560.9592
201%0.95120.89650.05470.99910.8973
5%0.98780.95030.03750.99940.9509
10%0.99750.96990.02760.99960.9703

It can be seen that the power , the non spurious power and the conditional power are increasing rapidly between sample sizes 5 and 10 at 1% and 10% levels of significance while are increasing rapidly between sample sizes 5 and 15 at 5% level of significance. The non spurious type-II is found to be very low and it shows very small changes throughout. The probability of showing the contaminant as an outlier, is high for high values (greater than 10) of the sample size, it is increasing rapidly between sample sizes 5 and 10 at 1% level of significance. While at 5% and 10% levels of significance it is showing negligible changes between sample sizes 5 and 10, beyond 10 it increases very slowly. On the basis of above inference, it can be interpreted that the power , the non spurious power , the conditional power and the probability that the contaminant is showing up as an outlier, are found to be very high and the non spurious type-II error is low, as desired. Thus, it can be said that the test statistic is performing very well in this case for sample sizes greater than 10.

Case-II: When all observations were smaller than the location parameter . The performance of the test statistic for detection of a lower outlying observation from a Gumbel sample with a contaminant observation which was taken from a sample of the same distribution with different scale parameter was studied. All the performance probabilities were calculated by simulation technique with 10,000 replications, and are given in Table 8.2. From Table 8.2, it can be seen that the performance is increasing with the sample size and it is good enough for sample sizes greater than 10.

Table 8.2: Performance of the test statistic for case-II at different levels of significance

nLevel of Significance $P_{1}$ $P_{3}$ $P_{1}-P_{3}$ $P_{5}$ $P_{3}/P_{5}$
51%0.55810.45660.10150.56160.8130
5%0.78620.74980.03640.79340.9450
10%0.87860.85430.02430.88480.9655
101%0.79490.77860.01630.79870.9548
5%0.97940.93640.04300.97880.9767
10%0.98310.96460.01850.98370.9806
151%0.93700.92630.01070.93620.9694
5%0.99760.97680.02080.99790.9789
10%0.99870.98540.01330.99870.9867
201%0.98040.97600.00440.98310.9828
5%0.99990.98830.01160.99790.9904
10%1.00000.99810.00191.00000.9981

From this, it can be interpreted that the power, the non spurious power and the conditional power of the test statistic are high for large sample sizes; these are increasing rapidly between sample sizes 5 and 10 while beyond sample size 10 it increases with a comparatively low rate of increase. The non spurious type-II error is low and decreasing slowly with a very small rate of decrease. The probability of identifying the contaminant as an outlying observation is also very high and increasing with the increase of the sample size, as desired for a good discordancy test according to Barnett and Lewis (1994) and Hayes and Kinsella (2003). Thus it can be concluded that the test statistic under consideration is performing very well for sample sizes greater than 5

IX. PERFORMANCE STUDY FOR A PAIR OF OUTLIERS

Case-I: When all observations were greater than the location parameter . The performance of the test statistic to detect a pair of outliers in a sample from Gumbel distribution after introducing a pair of contaminants from another sample from the same distribution with a shifted scale parameter was obtained by simulation technique with 10,000 replicates. These performance probabilities are given in Table 9.1.

nLevel of Significance $P_{1}$ $P_{3}$ $P_{1}-P_{3}$ $P_{5}$ $P_{3}/P_{5}$
51%0.26780.17900.08880.67310.2659
5%0.50660.50080.00580.84990.5892
10%0.61390.53600.07790.89960.5958
1%0.55140.45500.09640.87360.5208
105%0.80190.70950.09240.96650.7341
10%0.87930.77790.10140.98490.7898
151%0.75100.63250.11850.94680.6680
5%0.91690.82160.09530.99000.8299
10%0.96680.83760.12920.99670.8404
201%0.85790.72530.13260.97740.7421
5%0.96460.85120.11340.99660.8541
10%0.98960.84060.14900.99900.8414

It can be observed from the above table that the power, the non-spurious power and the conditional power of the test statistic are increasing rapidly throughout. It can also be seen that the power, the non spurious power, the conditional power and the probability of identifying the contaminant as an outlier is increasing with the increase of the sample size and the non spurious type-II error is very low. Since the probabilities are significantly high for sample sizes greater than 10 and is low, as desired for a good discordancy test. Thus, it can be interpreted that the performance the test statistic is good for large samples.

Case-II: When all observations were smaller than the location parameter . The performance of the test statistic to detect a pair of outliers in a sample from a Gumbel distribution after introducing a pair of contaminants from another sample from the same distribution with different scale parameter was studied by simulation technique with 10,000 replications. These performance probabilities are given in Table 9.2.

Table 9.2: Performance of the test statistic for case-II at different levels of significance.

nLevel of Significance $P_{1}$ $P_{3}$ $P_{1}-P_{3}$ $P_{5}$ $P_{3}/P_{5}$
51%0.01540.00660.00880.19150.0345
5%0.08910.08330.00580.44890.1856
10%0.20080.12290.07790.63540.1934
101%0.06440.04800.01640.39160.1226
5%0.27410.18370.09040.71310.2576
10%0.45630.41610.04020.83930.4958
151%0.11490.09640.01850.48670.1981
5%0.37060.27530.09530.76510.3598
10%0.61820.52500.09320.90280.5815
201%0.15260.11000.04260.49260.2233
5%0.58480.47140.11340.86280.5463
10%0.81070.63560.17510.96480.6587

From Table 9.2. it can be observed that since the power, the non spurious power, the conditional power, are very low thus the test statistic is not found to be satisfactory for the detection of a pair of outliers in a sample from case-II of a Gumbel distribution. It can be seen that the power of the statistic is increasing throughout with almost constant rate of increase. The non spurious power, is increasing with a good rate of increase between sample sizes 5 and 15 while beyond that almost no changes are observed. The non spurious type II error of the test statistic at 1% level of significance is found to be very low. The conditional power of the statistic at 1% level of significance shows that the conditional power increases between sample sizes 5 and 10, between sample sizes 10 and 15 almost no changes are observed while beyond sample size 15 again it increases uniformly. The probability of detecting an outlying pair at 1% level of significance is found to be low. It increases rapidly between sample sizes 5 and 15, while beyond 15 the rate of increase is low.

It can also be seen that the power increases with the increase of sample size throughout almost uniformly but small changes are occurring between sample sizes 10 and 15. The non-spurious power of the test statistic at 5% level of significance increases rapidly from sample size 5 to 15 and beyond sample size15, its rate of increase is comparatively low. The spurious power of the test statistic at 5% level of significance increases rapidly between sample sizes 5 and 10, the rate of increase is comparatively low between sample sizes 10 and 25 while beyond sample size 25, almost no changes are seen. The conditional power increases with good rate of increase between sample sizes 5 and 10 but gives relatively low rate of increase between sample sizes 10 and 20, beyond sample size 20 it increases uniformly. The probability of detecting contaminant(s) as an outlying pair at 5% level of significance is found to be good and is increasing with the increase of the sample size.

It can be noticed at 10% level of significance that the power of the test statistic under consideration increases with the increase of sample size up to 15 and beyond sample size 15 it gives relatively small changes. The non-spurious power of the test statistic at 10% level of significance increases rapidly with the increase of sample size from sample size 5 to 10 and beyond sample size 10 it increases uniformly with comparatively low rate of increase. The spurious power of the test statistic at 10% level of significance is found to be low. It decreases between sample sizes 5 and 10 and then starts to increase beyond 10 with a good rate of increase. The conditional power of the statistic at 10% level of significance shows that the conditional power increases rapidly between sample sizes 5 and 20. The probability of detecting contaminant(s) as an outlying pair at 10% level of significance is found to be high and increases with the increase of the sample size up to 10, beyond 10 the rate of increase is comparatively low. It can be concluded from the above inference that the test statistic is not found to be suitable for the detection of a pair of outliers in a sample from a Gumbel distribution. Therefore, it cannot be recommended for detection of a pair of outliers in a sample from a Gumbel distribution for case-II.

Conclusion: It can be concluded from the above study that the two suggested test statistics and (for detection of the single upper and lower outlier) are performing very well, while the performance of the test statistic suggested for the detection of a pair of outliers is very poor, especially for case-II. Thus, use of the test statistics and for detection of an upper and lower outlying observation respectively, can be used for the Gumbel distribution with unknown scale parameter but the test statistic cannot be suggested for efficient results in a Gumbel distribution.

Conflict of Interest

The authors declare no conflict of interest.

Ethical Approval

Not applicable

Data Availability

The datasets used in this study are openly available at [repository link] and the source code is available on GitHub at [GitHub link].

Funding

This work did not receive any external funding.

References

7 Cites in Article

Cite this article

Generating citation...

Related Research

  • LCC Code: QA273.6
  • Version of record

    v1.0

  • Issue date

    05 December 2023

  • Language

    en

Outlier Detection Procedures in a Sample from a Gumbel Distribution with Unknown Scale Parameter
Open Access
Research Article
CC-BY-NC 4.0
Views 646
Downloads 33
Special Issue

Launch a focused special issue to highlight research, emerging trends, and expert insights in your academic field.

Support