Scaled inverse chi-squared distribution

From Wikipedia, the free encyclopedia
Jump to navigation Jump to search
Scaled inverse chi-squared
Probability density function
Cumulative distribution function
Parameters ν>0
τ2>0
Support x(0,)
PDF (τ2ν/2)ν/2Γ(ν/2)exp[ντ22x]x1+ν/2
CDF Γ(ν2,τ2ν2x)/Γ(ν2)
Mean ντ2ν2 for ν>2
Mode ντ2ν+2
Variance 2ν2τ4(ν2)2(ν4)for ν>4
Skewness 4ν62(ν4)for ν>6
Excess kurtosis 12(5ν22)(ν6)(ν8)for ν>8
Entropy

ν2+ln(τ2ν2Γ(ν2))

(1+ν2)ψ(ν2)
MGF 2Γ(ν2)(τ2νt2)ν4Kν2(2τ2νt)
CF 2Γ(ν2)(iτ2νt2)ν4Kν2(2iτ2νt)

The scaled inverse chi-squared distribution ψinv-χ2(ν), where ψ is the scale parameter, equals the univariate inverse Wishart distribution 𝒲1(ψ,ν) with degrees of freedom ν.

This family of scaled inverse chi-squared distributions is linked to the inverse-chi-squared distribution and to the chi-squared distribution:

If Xψinv-χ2(ν) then X/ψinv-χ2(ν) as well as ψ/Xχ2(ν) and 1/Xψ1χ2(ν).

Instead of ψ, the scaled inverse chi-squared distribution is however most frequently parametrized by the scale parameter τ2=ψ/ν and the distribution ντ2inv-χ2(ν) is denoted by Scale-inv-χ2(ν,τ2).


In terms of τ2 the above relations can be written as follows:

If XScale-inv-χ2(ν,τ2) then Xντ2inv-χ2(ν) as well as ντ2Xχ2(ν) and 1/X1ντ2χ2(ν).


This family of scaled inverse chi-squared distributions is a reparametrization of the inverse-gamma distribution.

Specifically, if

Xψinv-χ2(ν)=Scale-inv-χ2(ν,τ2)   then   XInv-Gamma(ν2,ψ2)=Inv-Gamma(ν2,ντ22)


Either form may be used to represent the maximum entropy distribution for a fixed first inverse moment (E(1/X)) and first logarithmic moment (E(ln(X)).

The scaled inverse chi-squared distribution also has a particular use in Bayesian statistics. Specifically, the scaled inverse chi-squared distribution can be used as a conjugate prior for the variance parameter of a normal distribution. The same prior in alternative parametrization is given by the inverse-gamma distribution.

Characterization

[edit | edit source]

The probability density function of the scaled inverse chi-squared distribution extends over the domain x>0 and is

f(x;ν,τ2)=(τ2ν/2)ν/2Γ(ν/2)exp[ντ22x]x1+ν/2

where ν is the degrees of freedom parameter and τ2 is the scale parameter. The cumulative distribution function is

F(x;ν,τ2)=Γ(ν2,τ2ν2x)/Γ(ν2)
=Q(ν2,τ2ν2x)

where Γ(a,x) is the incomplete gamma function, Γ(x) is the gamma function and Q(a,x) is a regularized gamma function. The characteristic function is

φ(t;ν,τ2)=
2Γ(ν2)(iτ2νt2)ν4Kν2(2iτ2νt),

where Kν2(z) is the modified Modified Bessel function of the second kind.

Parameter estimation

[edit | edit source]

The maximum likelihood estimate of τ2 is

τ2=n/i=1n1xi.

The maximum likelihood estimate of ν2 can be found using Newton's method on:

ln(ν2)ψ(ν2)=1ni=1nln(xi)ln(τ2),

where ψ(x) is the digamma function. An initial estimate can be found by taking the formula for mean and solving it for ν. Let x¯=1ni=1nxi be the sample mean. Then an initial estimate for ν is given by:

ν2=x¯x¯τ2.

Bayesian estimation of the variance of a normal distribution

[edit | edit source]

The scaled inverse chi-squared distribution has a second important application, in the Bayesian estimation of the variance of a Normal distribution.

According to Bayes' theorem, the posterior probability distribution for quantities of interest is proportional to the product of a prior distribution for the quantities and a likelihood function:

p(σ2|D,I)p(σ2|I)p(D|σ2)

where D represents the data and I represents any initial information about σ2 that we may already have.

The simplest scenario arises if the mean μ is already known; or, alternatively, if it is the conditional distribution of σ2 that is sought, for a particular assumed value of μ.

Then the likelihood term L2|D) = p(D2) has the familiar form

(σ2|D,μ)=1(2πσ)nexp[in(xiμ)22σ2]

Combining this with the rescaling-invariant prior p(σ2|I) = 1/σ2, which can be argued (e.g. following Jeffreys) to be the least informative possible prior for σ2 in this problem, gives a combined posterior probability

p(σ2|D,I,μ)1σn+2exp[in(xiμ)22σ2]

This form can be recognised as that of a scaled inverse chi-squared distribution, with parameters ν = n and τ2 = s2 = (1/n) Σ (xi-μ)2

Gelman and co-authors remark that the re-appearance of this distribution, previously seen in a sampling context, may seem remarkable; but given the choice of prior "this result is not surprising."[1]

In particular, the choice of a rescaling-invariant prior for σ2 has the result that the probability for the ratio of σ2 / s2 has the same form (independent of the conditioning variable) when conditioned on s2 as when conditioned on σ2:

p(σ2s2|s2)=p(σ2s2|σ2)

In the sampling-theory case, conditioned on σ2, the probability distribution for (1/s2) is a scaled inverse chi-squared distribution; and so the probability distribution for σ2 conditioned on s2, given a scale-agnostic prior, is also a scaled inverse chi-squared distribution.

Use as an informative prior

[edit | edit source]

If more is known about the possible values of σ2, a distribution from the scaled inverse chi-squared family, such as Scale-inv-χ2(n0, s02) can be a convenient form to represent a more informative prior for σ2, as if from the result of n0 previous observations (though n0 need not necessarily be a whole number):

p(σ2|I,μ)1σn0+2exp[n0s022σ2]

Such a prior would lead to the posterior distribution

p(σ2|D,I,μ)1σn+n0+2exp[ns2+n0s022σ2]

which is itself a scaled inverse chi-squared distribution. The scaled inverse chi-squared distributions are thus a convenient conjugate prior family for σ2 estimation.

Estimation of variance when mean is unknown

[edit | edit source]

If the mean is not known, the most uninformative prior that can be taken for it is arguably the translation-invariant prior p(μ|I) ∝ const., which gives the following joint posterior distribution for μ and σ2,

p(μ,σ2D,I)1σn+2exp[in(xiμ)22σ2]=1σn+2exp[in(xix¯)22σ2]exp[n(μx¯)22σ2]

The marginal posterior distribution for σ2 is obtained from the joint posterior distribution by integrating out over μ,

p(σ2|D,I)1σn+2exp[in(xix¯)22σ2]exp[n(μx¯)22σ2]dμ=1σn+2exp[in(xix¯)22σ2]2πσ2/n(σ2)(n+1)/2exp[(n1)s22σ2]

This is again a scaled inverse chi-squared distribution, with parameters n1 and s2=(xix¯)2/(n1).

[edit | edit source]

References

[edit | edit source]
  • Lua error in Module:Citation/CS1/Configuration at line 2172: attempt to index field '?' (a nil value).
  1. ^ Lua error in Module:Citation/CS1/Configuration at line 2172: attempt to index field '?' (a nil value).