Quoted message said:A. Jakulin said:P-values can be understood as a decision-theoretical
approach. Basically, if your loss function has a
particular distribution, the p-value of 0.01 means that
in 0.01 the null model will have equal lower loss than
the alternative model.
I agree that p-values are not generally considered to be a
decision-theoretical approach. However, I am trying to point
out the parallels.
Herman Rubin said:In any decision-theoretical approach, the loss function
must come only from the consequences as seen by the user,
and cannot use any purely statistical input. One has to
start out with the consequences for acceptance or
rejection under the various states of nature, assessed
directly.
In many situations, the subject of the decision is a
probability distribution, not a particular action. It
remains unclear how to define loss functions on the space of
probability density functions, but as far as I am concerned,
KL-divergence is a reliable probabilistic loss function,
measuring the loss incurred by using a different PDF to
approximate the "true" PDF. Having decided on the
probability density function, the probabilistic model in the
first phase, one can employ decision-theoretic analysis of
actions with with this model in the second phase.
Quoted message said:Now if there is loss when the null is improperly rejected,
the probability of rejection is ONE component of the risk.
Precisely. There are two components, the reward and the
risk. The reward is the (expected) utility of the *decision*
under a particular model. The risk is the probability that
the *model* under which you're computing the utility is
better (as measured by a different kind of utility) than
some trivial null model. In a practical decision situation,
I have a model that the medication works, and a model that
the medication does not work. I want to decide upon a single
model. The expected reward is greater for the model that the
medication works, with p=0.9. However, with a p-value of 0.4
in favor of the null, I can only say that the evidence for
comparing the two models is inconclusive.
Some formulate this in a minmax style. One should pick the
least risky model, and pick the most rewarding decision. The
maximum entropy principle is an example of this approach.
Although this is not the core of this discussion, those who
are interested might take a look at the paper of
Grunwald&Dawid imstat.orgissue 32 4.htmlOpen ↗
Quoted message said:As to the meaning of the p-value, it changes drastically
from problem to problem. For a two-sided one-dimensional
normal, the change in posterior odds for symmetric prior
at a p-value of .05 is not 1/.05 = 20, but is at most
around 3.7. It does not mean what those who do not know
decision analysis think it does.
Agreed. However, I have given a precise definition of the
decision-theoretic meaning of the p-value (the probability
that the null model's loss on a new sample will be equal or
lower than the loss of the alternative model on the original
sample). Clearly, if you try to redefine the meaning of
risk, this definition will not fit.
Quoted message said:Even worse is that the null hypothesis is always false as
stated. This problem is much harder, although if the
region in the parameter space in which one wants to accept
is small relative to the precision of estimation, an
approximation by one with a prior probability of the point
null can be good. But the action to be taken will
correspond to SOME p-value, but this has to be computed by
taking into account everything.
I agree about this, but we are are not looking for the
"true" model (no model is true), we are trying to refute the
alternative informative hypothesis with a null uninformative
hypothesis. Both of them are somehow chosen. Usually, the
alternative hypothesis is the one that allows one to
demonstrate the utility of a particular intervention, and
the null one claims that the intervention is independent of
the outcome.
One problem is that the loss distribution incurred by the
alternative hypothesis is not fairly estimated, further
biasing results in favor of the null hypothesis. In the
resampling context, we could say that the distribution of
the null loss is estimated via resampling, but the
alternative loss is a point estimate without resampling. It
is very easy to reformulate and resample both at the same
time, that's more in the spirit of Neyman-Pearson
hypothesis testing. The so-defined "v-value" is defined as
the probability that the null model's loss on a new sample
will be equal or lower than the loss of the alternative
model on the same sample. The trouble is that the analytic
solution (without resampling) to this reformulated problem
is unwieldy.
Quoted message said:Suppose that we must make a decision based on a bivariate
normal random variable X, which has mean \theta and
variance vI; the null hypothesis is \theta = 0. If the
total risk is the probability of improper rejection plus
1/(2\pi) times the integral of the probability of improper
acceptance over the rest of the parameter space, the
best p-value can be calculated to be v. This might not
be a good example, but it can be done in closed form;
the more realistic ones are similar.
Well, agreed, but you're using a different definition of
risk.
In conclusion, I disagree with the categorical denial of p-
values, since the p-values are just about refuting the
alternative hypothesis, not about hypothesis selection. They
fully complement the analysis of expected utility, the usual
"testing" framework of decision theory. Their role is the
simplicity bias, answering the question of "Do we have
enough data to employ the alternative model?" If so, we can
employ decision theory to see what is the optimal decision
based on the alternative model. If not so, we don't have
enough data to even discuss the alternative model. In other
words, p-values are the Occam's razor with respect to what
model we are allowed to use to perform the analysis of
expected utility.
--
mag. Aleks Jakulin ai.fri.uni-lj.sialeksOpen ↗ Artificial
Intelligence Laboratory, Faculty of Computer and Information
Science, University of Ljubljana.