General fitness, health and nutrition · Public discussion

Clinical trials and P values

Started by Kathy · · Last activity · 28 posts · 1,366 views

Thread navigation

Jump through the discussion

Go to the original post, the replies on this page, or the latest preserved contribution.

Thread details

What we know about this thread

Original section
General fitness, health and nutrition
Published
4 March 2004
Last activity
15 March 2004
Original author
Kathy
Posts
28
Discussion status
Public discussion
Total views
1,366
Views / 30 days
0

The navigation and discussion metadata provide context. Posts remain in their original chronological order.

Showing posts 1–20 of 28
Posts remain in their original chronological order.

Text size
  1. Suppose that in clinical trials of medical treatments, only 10% of all trials result in a positive
    result. Positive meaning that the determination was made that the treatment had a statisically
    significant effect.

    Suppose that the outcome of a specific trial is a P value of .01.

    What is the overall probability that the results of this trial were do to chance?

    Thanks, Kathy

  2. In article <[email hidden]>,

    Kathy said:

    Suppose that in clinical trials of medical treatments, only 10% of all trials result in a positive
    result. Positive meaning that the determination was made that the treatment had a statisically
    significant effect.

    In this context, "statistically significant" has no meaning whatever. Given a sufficiently large
    sample, even a minuscule effect is going to have a high probability of showing up as statistically
    significant.

    Quoted message said:

    Suppose that the outcome of a specific trial is a P value of .01.

    Quoted message said:

    What is the overall probability that the results of this trial were do to chance?

    With more refinements, one can come up with an answer if the information was of the form is P <=
    .01. For the information that it comes out at the P value of .01, the posterior probability depends
    on the form of the distributions involved.

    P values need to be abandoned in favor of a good decision theoretical approach. They do mean
    something, but not what matters in medical decision making.
    --
    This address is for information only. I do not claim that these views are those of the Statistics
    Department or of Purdue University. Herman Rubin, Department of Statistics, Purdue University
    [email hidden] Phone: (765)494-6054 FAX: (765)494-0558

  3. Kathy said:

    Suppose that in clinical trials of medical treatments, only 10% of all trials result in a positive
    result. Positive meaning that the

    ... treatment is effective.

    Quoted message said:

    determination was made that the treatment had a statisically significant effect.

    No.

    Quoted message said:

    Suppose that the outcome of a specific trial is a P value of .01.

    What is the overall probability that the results of this trial were do to chance?

    The cumulative probability at 0.01 is

    .01*(1-0.1) / (0.01*(1-0.1) + 0.1*t(p))

    where 't' is the cumulative distribution function associated with p-values for the trials with the
    "positive effect", assuming the same common effect: t(p) = 1 - H(F^{-1} (1-p)), where F, H are the
    test statistic distributions under the null, and under the presence of effect situations, and {-1}
    denotes the inverse.

    You can differentiate t(p) to get the density, plug it in the formula above (and also replace 0.01
    by 1) to get the probability *at* 0.01.

    DZ

    Quoted message said:

    Thanks, Kathy

  4. P.S. t(p) is evaluated at p=0.01

    DZ said:
    Kathy said:

    Suppose that in clinical trials of medical treatments, only 10% of all trials result in a
    positive result. Positive meaning that the

    ... treatment is effective.

    Quoted message said:

    determination was made that the treatment had a statisically significant effect.

    No.

    Quoted message said:

    Suppose that the outcome of a specific trial is a P value of .01.

    What is the overall probability that the results of this trial were do to chance?

    The cumulative probability at 0.01 is

    .01*(1-0.1) / (0.01*(1-0.1) + 0.1*t(p))

    where 't' is the cumulative distribution function associated with p-values for the trials with the
    "positive effect", assuming the same common effect: t(p) = 1 - H(F^{-1} (1-p)), where F, H are the
    test statistic distributions under the null, and under the presence of effect situations, and {-1}
    denotes the inverse.

    You can differentiate t(p) to get the density, plug it in the formula above (and also replace 0.01
    by 1) to get the probability *at* 0.01.

    DZ

    Quoted message said:

    Thanks, Kathy

  5. [email hidden] (Herman Rubin) wrote in message news:<[email hidden]>...

    Quoted message said:
    In article said:

    Suppose that in clinical trials of medical treatments, only 10% of all trials result in a
    positive result. Positive meaning that the determination was made that the treatment had a
    statisically significant effect.

    In this context, "statistically significant" has no meaning whatever. Given a sufficiently large
    sample, even a minuscule effect is going to have a high probability of showing up as statistically
    significant.

    I disagree. A p value of .01 means the same thing whether the effect is 10% or .01%. It means that
    the observed effect is unlikely to occur due chance, and effects of this magnitude occur by chance
    only 1% of the time.

    Whether a treatment with an effect of .01% is worth taking is not answered by the p value. It
    probably is a "value judgement" and cannot be answered objectively.

    Quoted message said:


    Quoted message said:

    Suppose that the outcome of a specific trial is a P value of .01. What is the overall probability
    that the results of this trial were do to chance?

    With more refinements, one can come up with an answer if the information was of the form is P <=
    .01. For the information that it comes out at the P value of .01, the posterior probability
    depends on the form of the distributions involved.

    Do you mean the distribution of P values seen across all trails? Is that the info needed?

    Quoted message said:


    P values need to be abandoned in favor of a good decision theoretical approach. They do mean
    something, but not what matters in medical decision making.

    Maybe so, but it is heavly used right now. We should know how the interpretation should be effected
    by the overall success/failure rate of trials.

    Kathy

  6. [email hidden] (Kathy) wrote in message news:<[email hidden]>...

    Quoted message said:

    Suppose that in clinical trials of medical treatments, only 10% of all trials result in a positive
    result. Positive meaning that the determination was made that the treatment had a statisically
    significant effect.

    Suppose that the outcome of a specific trial is a P value of .01.

    What is the overall probability that the results of this trial were do to chance?

    Thanks, Kathy

    Unless you argue some sort of dependency, the answer is obvious.

    js

  7. In article <[email hidden]>,

    Kathy said:

    [email hidden] (Herman Rubin) wrote in message
    news:<[email hidden]>...

    Quoted message said:
    In article said:

    Suppose that in clinical trials of medical treatments, only 10% of all trials result in a
    positive result. Positive meaning that the determination was made that the treatment had a
    statisically significant effect.

    Quoted message said:
    Quoted message said:

    In this context, "statistically significant" has no meaning whatever. Given a sufficiently large
    sample, even a minuscule effect is going to have a high probability of showing up as
    statistically significant.

    Quoted message said:

    I disagree. A p value of .01 means the same thing whether the effect is 10% or .01%. It means that
    the observed effect is unlikely to occur due chance, and effects of this magnitude occur by chance
    only 1% of the time.

    Only if the null hypothesis is exactly correct.

    Quoted message said:

    Whether a treatment with an effect of .01% is worth taking is not answered by the p value. It
    probably is a "value judgement" and cannot be answered objectively.

    The p value, and not the magnitude of an effect, is a better guide to whether a treatment is
    worth taking?

    Quoted message said:
    Quoted message said:
    Quoted message said:

    Suppose that the outcome of a specific trial is a P value of .01. What is the overall
    probability that the results of this trial were do to chance?

    Quoted message said:
    Quoted message said:

    With more refinements, one can come up with an answer if the information was of the form is P <=
    .01. For the information that it comes out at the P value of .01, the posterior probability
    depends on the form of the distributions involved.

    Quoted message said:

    Do you mean the distribution of P values seen across all trails? Is that the info needed?

    Not at all. The information needed is the likelihood function for the p value exactly equal to .01.
    What is the situation for the possible observations with p values less than .01 is irrelevant; it is
    this sample which matters, and not how it ranks among others.

    Quoted message said:
    Quoted message said:

    P values need to be abandoned in favor of a good decision theoretical approach. They do mean
    something, but not what matters in medical decision making.

    Quoted message said:

    Maybe so, but it is heavly used right now. We should know how the interpretation should be effected
    by the overall success/failure rate of trials.

    This is the same argument offered to justify using alchemy instead of chemistry, or possibly more
    like using religious arguments instead of science. Statistics has been called the religion of
    medicine, and the way that p values are used is only justified on religious grounds.

    --
    This address is for information only. I do not claim that these views are those of the Statistics
    Department or of Purdue University. Herman Rubin, Department of Statistics, Purdue University
    [email hidden] Phone: (765)494-6054 FAX: (765)494-0558

  8. My problems with most of this is the focus on statistical
    significance. My question then becomes, shouldn't we be
    using stat significance as the first cut? After we find
    something is more effective than placebo, is there a way to
    look at clinical significance in a well, clinical manner? Is
    Clin sig too much an eye-of-the-beholder thing?

    --
    I'd certainly be far more supportive of Bush's
    "Axis of Evil" concept if he'd been honest about it
    and added Microsoft to the list.
    _Chris J in abt-c

  9. In article <[email hidden]>,

    Kurt Ullman said:

    My problems with most of this is the focus on
    statistical significance. My question then becomes,
    shouldn't we be using stat significance as the first
    cut? After we find something is more effective than
    placebo, is there a way to look at clinical
    significance in a well, clinical manner? Is Clin sig
    too much an eye-of-the-beholder thing?

    Something can be quite effective, but lost in statistical
    significance. I recall an article on the carcinogenicity of
    a substance, with (I may have the numbers not quite correct)
    with one out of 25 rats in the control group getting cancer,
    and 12 of the experimental group. This was called a moderate
    effect because the p value was on the order of .035.

    In the British study of Type 2 diabetics, something was
    called unimportant because the p value was .052.

    I have seen studies in which the balanced set of subjects
    became imbalanced because of dropouts presumably unrelated
    to the experiment. This can cause a larger effect than what
    is being measured, but the lack of statistical significance
    in the proportion of the various groups causes this to be
    disregarded.

    Many problems are very high dimensional. In these, if one
    considers the variables and attempts to find what is
    relevant, classical statistics cannot be used at all. But
    the medical investigators, steeped in p values, cannot even
    attempts what should be done.

    --
    This address is for information only. I do not claim that
    these views are those of the Statistics Department or of
    Purdue University. Herman Rubin, Department of Statistics,
    Purdue University [email hidden] Phone: (765)494-
    6054 FAX: (765)494-0558

  10. In article <[email hidden]>,

    Kurt Ullman said:

    My problems with most of this is the focus on
    statistical significance. My question then becomes,
    shouldn't we be using stat significance as the first
    cut? After we find something is more effective than
    placebo, is there a way to look at clinical
    significance in a well, clinical manner? Is Clin sig
    too much an eye-of-the-beholder thing?

    Something can be quite effective, but lost in statistical
    significance. I recall an article on the carcinogenicity of
    a substance, with (I may have the numbers not quite correct)
    with one out of 25 rats in the control group getting cancer,
    and 12 of the experimental group. This was called a moderate
    effect because the p value was on the order of .035.

    In the British study of Type 2 diabetics, something was
    called unimportant because the p value was .052.

    I have seen studies in which the balanced set of subjects
    became imbalanced because of dropouts presumably unrelated
    to the experiment. This can cause a larger effect than what
    is being measured, but the lack of statistical significance
    in the proportion of the various groups causes this to be
    disregarded.

    Many problems are very high dimensional. In these, if one
    considers the variables and attempts to find what is
    relevant, classical statistics cannot be used at all. But
    the medical investigators, steeped in p values, cannot even
    attempts what should be done.

    --
    This address is for information only. I do not claim that
    these views are those of the Statistics Department or of
    Purdue University. Herman Rubin, Department of Statistics,
    Purdue University [email hidden] Phone: (765)494-
    6054 FAX: (765)494-0558

  11. I'm responding to Herman, concerning proper description of
    studies. I'm also responding to the original question, which
    Kurt brings up again, about statistical significance.

    On 6 Mar 2004 08:58:57 -0500, [email hidden]

    (Herman Rubin) said:

    In article
    <[email hidden]>,

    Kurt Ullman said:

    My problems with most of this is the focus on
    statistical significance. My question then becomes,
    shouldn't we be using stat significance as the first
    cut? After we find something is more effective than
    placebo, is there a way to look at clinical
    significance in a well, clinical manner? Is Clin sig
    too much an eye-of-the-beholder thing?

    HR >

    Quoted message said:

    Something can be quite effective, but lost in statistical
    significance. I recall an article on the carcinogenicity
    of a substance, with (I may have the numbers not quite
    correct) with one out of 25 rats in the control group
    getting cancer, and 12 of the experimental group. This was
    called a moderate effect because the p value was on the
    order of .035.

    I would call that "bad reporting". It should have been
    called, say, "moderate statistical support for a clinical
    effect which was very large in these rats."

    Quoted message said:


    In the British study of Type 2 diabetics, something was
    called unimportant because the p value was .052. [ snip,
    rest of Herman's ]

    Similarly - though Herman leaves us without 'effect size' in
    meaningful terms - that *sounds* like bad reporting. (Did
    the original study also omit Odds ratios, or other units of
    'effect size'? Were there dozens of 'effects' in the study
    that had much smaller p-values, which accounted for most of
    the same substance?)

    Back to the question of 'statistical significance' -- It was
    *stated* that 10% of trials met a 1% test; it was asked (I
    think), 'Is this significant?'

    - In the narrowest statistical terms, with assumptions that
    are probably not met, the answer depends on the number of
    trials. Here is the 'test of hypothesis': One in 10: no;
    two in 20: at 2%; three in 30: pretty good. Assumed:
    Equal weights, implying equal sizes for the studies, and
    equal quality; trials performed essentially the same,
    providing interchangeable elements that allow the
    binomial to pertain.

    Since there are better techniques that tell us more, the
    *most* that this 3/30 tells us is that "There is *something*
    more to say about these data; if the result was 2 in 30 or
    only 1 in 100, we might feel like we should move on to other
    topics." As Kurt suggests, though, this "10%" does provide a
    'first cut' that encourages us (perhaps) to dig deeper.

    However.
    - Since the burgeoning of meta-analysis in the last 10
    years, nobody combines 'drug trials' in this way. For one
    style of analysis, you insist on exact p-levels. For a
    better style of analysis, you insist on exact effect
    sizes: even though "alive/ dead within x days" is
    criterion that is hard to match for being unambiguous. If
    the studies are equivalent, they you determine (in some
    fashion) to "average the results."

    - Another standard emerging from meta analysis -- though
    many of the ones performed have failed to live up to it
    -- is that homogeneity of outcome is essential, if you
    are going to average their results. So, if some of the
    trials definitely *failed*, by showing the opposite
    effect, say, then there is something serious to explain,
    about how and why these studies differed, when they were
    supposed to test the same thing. - if they did not
    differ, more than by chance, then that is well and
    good... if they differ, that suggests an analysis of what
    *mattered* between them.

    The lousiness of the practice of meta-analysis so far also
    emphasizes that it takes a practiced eye and good reading,
    in order to combine studies. If it starts out as an
    inappropriate 'review of the literature', it is not going to
    be fixed by stretching it out on a statistical apparatus.
    Should the study that was carefully controlled, randomized
    and blinded, with N=300, and experienced
    clinician/scientists be weighted the same as every N=20 hack
    job? or N=300 hack job?

    So: How numerous and comparable were those studies where 10%
    beat a 1% test?

    --
    Rich Ulrich, [email hidden]
    pitt.eduindex.html
    - I need a new job, after March 31. Openings? -

  12. In article <[email hidden]>,

    Rich Ulrich said:

    I'm responding to Herman, concerning proper description of
    studies. I'm also responding to the original question,
    which Kurt brings up again, about statistical significance.

    Quoted message said:

    On 6 Mar 2004 08:58:57 -0500, [email hidden]
    (Herman Rubin) wrote:

    Quoted message said:
    Quoted message said:

    In article <[email hidden]-
    thlink.net>, Kurt Ullman <[email hidden]> wrote:

    Quoted message said:
    Quoted message said:
    Quoted message said:

    My problems with most of this is the focus on
    statistical significance. My question then becomes,
    shouldn't we be using stat significance as the first
    cut? After we find something is more effective than
    placebo, is there a way to look at clinical
    significance in a well, clinical manner? Is Clin sig
    too much an eye-of-the-beholder thing?

    Quoted message said:

    HR >


    <> Something can be quite effective, but lost in statistical
    <> significance. I recall an article on the carcinogenicity
    <> of a substance, with (I may have the numbers not quite <>
    correct) with one out of 25 rats in the control group <>
    getting cancer, and 12 of the experimental group. This <>
    was called a moderate effect because the p value was on <>
    the order of .035.

    Quoted message said:

    I would call that "bad reporting". It should have been
    called, say, "moderate statistical support for a clinical
    effect which was very large in these rats."

    This "bad reporting" is the rule, not the exception.

    <> In the British study of Type 2 diabetics, something was
    <> called unimportant because the p value was .052. <> [
    snip, rest of Herman's ]

    Quoted message said:

    Similarly - though Herman leaves us without 'effect size'
    in meaningful terms - that *sounds* like bad reporting.
    (Did the original study also omit Odds ratios, or other
    units of 'effect size'? Were there dozens of 'effects' in
    the study that had much smaller p-values, which accounted
    for most of the same substance?)

    This very large study gave little information on effect size
    for anything, and I doubt that any of those involved have
    any idea of what "odds ratios" means.

    Quoted message said:

    Back to the question of 'statistical significance' -- It
    was *stated* that 10% of trials met a 1% test; it was asked
    (I think), 'Is this significant?'

    Quoted message said:

    - In the narrowest statistical terms, with assumptions
    that are probably not met, the answer depends on the
    number of trials. Here is the 'test of hypothesis': One
    in 10: no; two in 20: at 2%; three in 30: pretty good.
    Assumed: Equal weights, implying equal sizes for the
    studies, and equal quality; trials performed essentially
    the same, providing interchangeable elements that allow
    the binomial to pertain.

    Quoted message said:

    Since there are better techniques that tell us more, the
    *most* that this 3/30 tells us is that "There is
    *something* more to say about these data; if the result was
    2 in 30 or only 1 in 100, we might feel like we should move
    on to other topics." As Kurt suggests, though, this "10%"
    does provide a 'first cut' that encourages us (perhaps) to
    dig deeper.

    Quoted message said:

    However.
    - Since the burgeoning of meta-analysis in the last 10
    years, nobody combines 'drug trials' in this way. For
    one style of analysis, you insist on exact p-levels. For
    a better style of analysis, you insist on exact effect
    sizes: even though "alive/ dead within x days" is
    criterion that is hard to match for being unambiguous.
    If the studies are equivalent, they you determine (in
    some fashion) to "average the results."

    Quoted message said:

    - Another standard emerging from meta analysis -- though
    many of the ones performed have failed to live up to it
    -- is that homogeneity of outcome is essential, if you
    are going to average their results. So, if some of the
    trials definitely *failed*, by showing the opposite
    effect, say, then there is something serious to explain,
    about how and why these studies differed, when they were
    supposed to test the same thing. - if they did not
    differ, more than by chance, then that is well and
    good... if they differ, that suggests an analysis of
    what *mattered* between them.

    This still suffers from the underreporting of failed trials.
    There are ways around this, but they are not easy to carry
    out, and the assumptions are hard to assess. A failed trial
    for one effect might have been important for another, but it
    is very difficult to assess the failed trials.

    If only significant results are reported, we can look at the
    distribution of the p values. If they are uniform between 0
    and .05, this would be an indication of pure chance. If
    there is a discreteness effect, it is even more complicated.

    Quoted message said:

    The lousiness of the practice of meta-analysis so far also
    emphasizes that it takes a practiced eye and good reading,
    in order to combine studies. If it starts out as an
    inappropriate 'review of the literature', it is not going
    to be fixed by stretching it out on a statistical
    apparatus. Should the study that was carefully controlled,
    randomized and blinded, with N=300, and experienced
    clinician/scientists be weighted the same as every N=20
    hack job? or N=300 hack job?

    Quoted message said:

    So: How numerous and comparable were those studies where
    10% beat a 1% test?

    Quoted message said:

    --
    Rich Ulrich, [email hidden]
    pitt.eduindex.html
    - I need a new job, after March 31. Openings? -

    --
    This address is for information only. I do not claim that
    these views are those of the Statistics Department or of
    Purdue University. Herman Rubin, Department of Statistics,
    Purdue University [email hidden] Phone: (765)494-
    6054 FAX: (765)494-0558

  13. In article <[email hidden]>,

    Rich Ulrich said:

    I'm responding to Herman, concerning proper description of
    studies. I'm also responding to the original question,
    which Kurt brings up again, about statistical significance.

    Quoted message said:

    On 6 Mar 2004 08:58:57 -0500, [email hidden]
    (Herman Rubin) wrote:

    Quoted message said:
    Quoted message said:

    In article <[email hidden]-
    thlink.net>, Kurt Ullman <[email hidden]> wrote:

    Quoted message said:
    Quoted message said:
    Quoted message said:

    My problems with most of this is the focus on
    statistical significance. My question then becomes,
    shouldn't we be using stat significance as the first
    cut? After we find something is more effective than
    placebo, is there a way to look at clinical
    significance in a well, clinical manner? Is Clin sig
    too much an eye-of-the-beholder thing?

    Quoted message said:

    HR >


    <> Something can be quite effective, but lost in statistical
    <> significance. I recall an article on the carcinogenicity
    <> of a substance, with (I may have the numbers not quite <>
    correct) with one out of 25 rats in the control group <>
    getting cancer, and 12 of the experimental group. This <>
    was called a moderate effect because the p value was on <>
    the order of .035.

    Quoted message said:

    I would call that "bad reporting". It should have been
    called, say, "moderate statistical support for a clinical
    effect which was very large in these rats."

    This "bad reporting" is the rule, not the exception.

    <> In the British study of Type 2 diabetics, something was
    <> called unimportant because the p value was .052. <> [
    snip, rest of Herman's ]

    Quoted message said:

    Similarly - though Herman leaves us without 'effect size'
    in meaningful terms - that *sounds* like bad reporting.
    (Did the original study also omit Odds ratios, or other
    units of 'effect size'? Were there dozens of 'effects' in
    the study that had much smaller p-values, which accounted
    for most of the same substance?)

    This very large study gave little information on effect size
    for anything, and I doubt that any of those involved have
    any idea of what "odds ratios" means.

    Quoted message said:

    Back to the question of 'statistical significance' -- It
    was *stated* that 10% of trials met a 1% test; it was asked
    (I think), 'Is this significant?'

    Quoted message said:

    - In the narrowest statistical terms, with assumptions
    that are probably not met, the answer depends on the
    number of trials. Here is the 'test of hypothesis': One
    in 10: no; two in 20: at 2%; three in 30: pretty good.
    Assumed: Equal weights, implying equal sizes for the
    studies, and equal quality; trials performed essentially
    the same, providing interchangeable elements that allow
    the binomial to pertain.

    Quoted message said:

    Since there are better techniques that tell us more, the
    *most* that this 3/30 tells us is that "There is
    *something* more to say about these data; if the result was
    2 in 30 or only 1 in 100, we might feel like we should move
    on to other topics." As Kurt suggests, though, this "10%"
    does provide a 'first cut' that encourages us (perhaps) to
    dig deeper.

    Quoted message said:

    However.
    - Since the burgeoning of meta-analysis in the last 10
    years, nobody combines 'drug trials' in this way. For
    one style of analysis, you insist on exact p-levels. For
    a better style of analysis, you insist on exact effect
    sizes: even though "alive/ dead within x days" is
    criterion that is hard to match for being unambiguous.
    If the studies are equivalent, they you determine (in
    some fashion) to "average the results."

    Quoted message said:

    - Another standard emerging from meta analysis -- though
    many of the ones performed have failed to live up to it
    -- is that homogeneity of outcome is essential, if you
    are going to average their results. So, if some of the
    trials definitely *failed*, by showing the opposite
    effect, say, then there is something serious to explain,
    about how and why these studies differed, when they were
    supposed to test the same thing. - if they did not
    differ, more than by chance, then that is well and
    good... if they differ, that suggests an analysis of
    what *mattered* between them.

    This still suffers from the underreporting of failed trials.
    There are ways around this, but they are not easy to carry
    out, and the assumptions are hard to assess. A failed trial
    for one effect might have been important for another, but it
    is very difficult to assess the failed trials.

    If only significant results are reported, we can look at the
    distribution of the p values. If they are uniform between 0
    and .05, this would be an indication of pure chance. If
    there is a discreteness effect, it is even more complicated.

    Quoted message said:

    The lousiness of the practice of meta-analysis so far also
    emphasizes that it takes a practiced eye and good reading,
    in order to combine studies. If it starts out as an
    inappropriate 'review of the literature', it is not going
    to be fixed by stretching it out on a statistical
    apparatus. Should the study that was carefully controlled,
    randomized and blinded, with N=300, and experienced
    clinician/scientists be weighted the same as every N=20
    hack job? or N=300 hack job?

    Quoted message said:

    So: How numerous and comparable were those studies where
    10% beat a 1% test?

    Quoted message said:

    --
    Rich Ulrich, [email hidden]
    pitt.eduindex.html
    - I need a new job, after March 31. Openings? -

    --
    This address is for information only. I do not claim that
    these views are those of the Statistics Department or of
    Purdue University. Herman Rubin, Department of Statistics,
    Purdue University [email hidden] Phone: (765)494-
    6054 FAX: (765)494-0558

  14. Herman Rubin wrote:

    HR:

    Quoted message said:

    <> Something can be quite effective, but lost in
    statistical <> significance. I recall an article on the
    carcinogenicity <> of a substance, with (I may have the
    numbers not quite <> correct) with one out of 25 rats in
    the control group <> getting cancer, and 12 of the
    experimental group. This <> was called a moderate effect
    because the p value was on <> the order of .035.

    RU:

    Quoted message said:
    Quoted message said:

    I would call that "bad reporting". It should have been
    called, say, "moderate statistical support for a clinical
    effect which was very large in these rats."


    HR:

    Quoted message said:

    This "bad reporting" is the rule, not the exception.

    Is there any reason to think reporting would improve if
    medical researchers adopted the decision approach Professor
    Rubin and others promote?

    Cheers, Bruce
    --
    Bruce Weaver [email hidden]
    www.angelfire.com/wv/bwhomedir/

  15. Herman Rubin wrote:

    HR:

    Quoted message said:

    <> Something can be quite effective, but lost in
    statistical <> significance. I recall an article on the
    carcinogenicity <> of a substance, with (I may have the
    numbers not quite <> correct) with one out of 25 rats in
    the control group <> getting cancer, and 12 of the
    experimental group. This <> was called a moderate effect
    because the p value was on <> the order of .035.

    RU:

    Quoted message said:
    Quoted message said:

    I would call that "bad reporting". It should have been
    called, say, "moderate statistical support for a clinical
    effect which was very large in these rats."


    HR:

    Quoted message said:

    This "bad reporting" is the rule, not the exception.

    Is there any reason to think reporting would improve if
    medical researchers adopted the decision approach Professor
    Rubin and others promote?

    Cheers, Bruce
    --
    Bruce Weaver [email hidden]
    www.angelfire.com/wv/bwhomedir/

  16. Herman Rubin said:

    With more refinements, one can come up with an answer if
    the information was of the form is P <= .01. For the
    information that it comes out at the P value of .01, the
    posterior probability depends on the form of the
    distributions involved.

    P values need to be abandoned in favor of a good decision
    theoretical approach. They do mean something, but not what
    matters in medical decision making.

    P-values can be understood as a decision-theoretical
    approach. Basically, if your loss function has a particular
    distribution, the p-value of 0.01 means that in 0.01 the
    null model will have equal or lower loss than the
    alternative model. For example, it is known that log-loss
    (Kullback-Leibler divergence) multiplied by 2N (N being the
    number of samples) has a chi-squared distribution in
    categorical models. The null model has well-defined
    semantics: the null model: the drug has no influence; the
    alternative model: the drug affects the outcome.

    What one can do, however, is to place the question in a
    different way. Assume that the alternative model is
    *practically* significant if it demonstrates a benefit on
    100 samples (the number of patients in an hypothetical
    medical department).

    So I agree that a decision-theoretic view is better, but we
    should not ignore the decision-theoretic interpretation of
    p-values when the statistic under study corresponds to a
    loss function.

    --
    mag. Aleks Jakulin ai.fri.uni-lj.sialeks Artificial
    Intelligence Laboratory, Faculty of Computer and Information
    Science, University of Ljubljana.

  17. In article <[email hidden]>,

    Aleks Jakulin jakulin@@ieee.org said:
    Herman Rubin said:

    With more refinements, one can come up with an answer if
    the information was of the form is P <= .01. For the
    information that it comes out at the P value of .01, the
    posterior probability depends on the form of the
    distributions involved.

    Quoted message said:
    Quoted message said:

    P values need to be abandoned in favor of a good decision
    theoretical approach. They do mean something, but not
    what matters in medical decision making.

    Quoted message said:

    P-values can be understood as a decision-theoretical
    approach. Basically, if your loss function has a particular
    distribution, the p-value of 0.01 means that in 0.01 the
    null model will have equal or lower loss than the
    alternative model.

    For one thing, this is not so. For another, it is not a decision-
    theoretical approach.

    In any decision-theoretical approach, the loss function must
    come only from the consequences as seen by the user, and
    cannot use any purely statistical input. One has to start
    out with the consequences for acceptance or rejection under
    the various states of nature, assessed directly.

    Now if there is loss when the null is improperly rejected,
    the probability of rejection is ONE component of the risk.
    There is also the probability of improper acceptance as a
    function of each non-null state of nature, and it can be
    argued that one should use a linear functional of this, and
    a linear combination with the significance level.

    As to the meaning of the p-value, it changes drastically
    from problem to problem. For a two-sided one-dimensional
    normal, the change in posterior odds for symmetric priors at
    a p-value of .05 is not 1/.05 = 20, but is at most around
    3.7. It does not mean what those who do not know decision
    analysis think it does.

    Even worse is that the null hypothesis is always false as
    stated. This problem is much harder, although if the region
    in the parameter space in which one wants to accept is small
    relative to the precision of estimation, an approximation by
    one with a prior probability of the point null can be good.
    But the action to be taken will correspond to SOME p-value,
    but this has to be computed by taking into account
    everything.

    Suppose that we must make a decision based on a bivariate
    normal random variable X, which has mean \theta and variance
    vI; the null hypothesis is \theta = 0. If the total risk is
    the probability of improper rejection plus
    1/(2\pi) times the integral of the probability of improper
    acceptance over the rest of the parameter space, the best
    p-value can be calculated to be v. This might not be a
    good example, but it can be done in closed form; the more
    realistic ones are similar.
    --
    This address is for information only. I do not claim that
    these views are those of the Statistics Department or of
    Purdue University. Herman Rubin, Department of Statistics,
    Purdue University [email hidden] Phone: (765)494-
    6054 FAX: (765)494-0558

  18. In article <[email hidden]>,

    Bruce Weaver said:

    Herman Rubin wrote:

    Quoted message said:

    HR:

    Quoted message said:

    <> Something can be quite effective, but lost in
    statistical <> significance. I recall an article on the
    carcinogenicity <> of a substance, with (I may have the
    numbers not quite <> correct) with one out of 25 rats in
    the control group <> getting cancer, and 12 of the
    experimental group. This <> was called a moderate effect
    because the p value was on <> the order of .035.

    Quoted message said:

    RU:

    Quoted message said:
    Quoted message said:

    I would call that "bad reporting". It should have been
    called, say, "moderate statistical support for a clinical
    effect which was very large in these rats."

    Quoted message said:

    HR:

    Quoted message said:
    Quoted message said:

    This "bad reporting" is the rule, not the exception.

    Quoted message said:

    Is there any reason to think reporting would improve if
    medical researchers adopted the decision approach Professor
    Rubin and others promote?

    The reporting would not use p-values in this manner. They
    could use them as one component of risk.
    --
    This address is for information only. I do not claim that
    these views are those of the Statistics Department or of
    Purdue University. Herman Rubin, Department of Statistics,
    Purdue University [email hidden] Phone: (765)494-
    6054 FAX: (765)494-0558

  19. Quoted message said:
    A. Jakulin said:

    P-values can be understood as a decision-theoretical
    approach. Basically, if your loss function has a
    particular distribution, the p-value of 0.01 means that
    in 0.01 the null model will have equal lower loss than
    the alternative model.

    I agree that p-values are not generally considered to be a
    decision-theoretical approach. However, I am trying to point
    out the parallels.

    Herman Rubin said:

    In any decision-theoretical approach, the loss function
    must come only from the consequences as seen by the user,
    and cannot use any purely statistical input. One has to
    start out with the consequences for acceptance or
    rejection under the various states of nature, assessed
    directly.

    In many situations, the subject of the decision is a
    probability distribution, not a particular action. It
    remains unclear how to define loss functions on the space of
    probability density functions, but as far as I am concerned,
    KL-divergence is a reliable probabilistic loss function,
    measuring the loss incurred by using a different PDF to
    approximate the "true" PDF. Having decided on the
    probability density function, the probabilistic model in the
    first phase, one can employ decision-theoretic analysis of
    actions with with this model in the second phase.

    Quoted message said:

    Now if there is loss when the null is improperly rejected,
    the probability of rejection is ONE component of the risk.

    Precisely. There are two components, the reward and the
    risk. The reward is the (expected) utility of the *decision*
    under a particular model. The risk is the probability that
    the *model* under which you're computing the utility is
    better (as measured by a different kind of utility) than
    some trivial null model. In a practical decision situation,
    I have a model that the medication works, and a model that
    the medication does not work. I want to decide upon a single
    model. The expected reward is greater for the model that the
    medication works, with p=0.9. However, with a p-value of 0.4
    in favor of the null, I can only say that the evidence for
    comparing the two models is inconclusive.

    Some formulate this in a minmax style. One should pick the
    least risky model, and pick the most rewarding decision. The
    maximum entropy principle is an example of this approach.
    Although this is not the core of this discussion, those who
    are interested might take a look at the paper of
    Grunwald&Dawid imstat.orgissue 32 4.html

    Quoted message said:

    As to the meaning of the p-value, it changes drastically
    from problem to problem. For a two-sided one-dimensional
    normal, the change in posterior odds for symmetric prior
    at a p-value of .05 is not 1/.05 = 20, but is at most
    around 3.7. It does not mean what those who do not know
    decision analysis think it does.

    Agreed. However, I have given a precise definition of the
    decision-theoretic meaning of the p-value (the probability
    that the null model's loss on a new sample will be equal or
    lower than the loss of the alternative model on the original
    sample). Clearly, if you try to redefine the meaning of
    risk, this definition will not fit.

    Quoted message said:

    Even worse is that the null hypothesis is always false as
    stated. This problem is much harder, although if the
    region in the parameter space in which one wants to accept
    is small relative to the precision of estimation, an
    approximation by one with a prior probability of the point
    null can be good. But the action to be taken will
    correspond to SOME p-value, but this has to be computed by
    taking into account everything.

    I agree about this, but we are are not looking for the
    "true" model (no model is true), we are trying to refute the
    alternative informative hypothesis with a null uninformative
    hypothesis. Both of them are somehow chosen. Usually, the
    alternative hypothesis is the one that allows one to
    demonstrate the utility of a particular intervention, and
    the null one claims that the intervention is independent of
    the outcome.

    One problem is that the loss distribution incurred by the
    alternative hypothesis is not fairly estimated, further
    biasing results in favor of the null hypothesis. In the
    resampling context, we could say that the distribution of
    the null loss is estimated via resampling, but the
    alternative loss is a point estimate without resampling. It
    is very easy to reformulate and resample both at the same
    time, that's more in the spirit of Neyman-Pearson
    hypothesis testing. The so-defined "v-value" is defined as
    the probability that the null model's loss on a new sample
    will be equal or lower than the loss of the alternative
    model on the same sample. The trouble is that the analytic
    solution (without resampling) to this reformulated problem
    is unwieldy.

    Quoted message said:

    Suppose that we must make a decision based on a bivariate
    normal random variable X, which has mean \theta and
    variance vI; the null hypothesis is \theta = 0. If the
    total risk is the probability of improper rejection plus
    1/(2\pi) times the integral of the probability of improper
    acceptance over the rest of the parameter space, the
    best p-value can be calculated to be v. This might not
    be a good example, but it can be done in closed form;
    the more realistic ones are similar.

    Well, agreed, but you're using a different definition of
    risk.

    In conclusion, I disagree with the categorical denial of p-
    values, since the p-values are just about refuting the
    alternative hypothesis, not about hypothesis selection. They
    fully complement the analysis of expected utility, the usual
    "testing" framework of decision theory. Their role is the
    simplicity bias, answering the question of "Do we have
    enough data to employ the alternative model?" If so, we can
    employ decision theory to see what is the optimal decision
    based on the alternative model. If not so, we don't have
    enough data to even discuss the alternative model. In other
    words, p-values are the Occam's razor with respect to what
    model we are allowed to use to perform the analysis of
    expected utility.

    --
    mag. Aleks Jakulin ai.fri.uni-lj.sialeks Artificial
    Intelligence Laboratory, Faculty of Computer and Information
    Science, University of Ljubljana.

  20. [email hidden] (Herman Rubin) wrote in message news:<[email hidden]>...

    Quoted message said:

    In article
    <[email hidden]>,

    Kurt Ullman said:

    My problems with most of this is the focus on
    statistical significance. My question then becomes,
    shouldn't we be using stat significance as the first
    cut? After we find something is more effective than
    placebo, is there a way to look at clinical
    significance in a well, clinical manner? Is Clin sig
    too much an eye-of-the-beholder thing?

    Something can be quite effective, but lost in statistical
    significance.

    If you under-powered your study you don't have anything more
    than an exploratory result that is hypothesis generating.

    Quoted message said:

    I recall an article on the carcinogenicity of a substance,
    with (I may have the numbers not quite correct) with one
    out of 25 rats in the control group getting cancer, and 12
    of the experimental group. This was called a moderate
    effect because the p value was on the order of .035.

    And statistically, it is.

    Quoted message said:

    In the British study of Type 2 diabetics, something was
    called unimportant because the p value was .052.

    It likely wasn't called unimportant, Herman. However, if you
    are testing a drug effect and your primary endpoint fails to
    reject the null - you lose. Period. Sure would have been
    nice to have had a more specific reference.

    Here is the largest UK Diabetes trial (the UKPDS) and the
    most recent large cut analysis (UKPDS 61).

    Are Lower fasting Plasma Glucose Levels at Diagnosis of Type
    2 Diabetes Associated With Improved Outcomes? Stephen
    Coalagiuri, Carol A Cull, Rury R Holman for the UKPDS Group
    Diabetes Care (2002); 25: 1410-1417

    It's available full text at the Diabetes Care website.

    Quoted message said:

    I have seen studies in which the balanced set of subjects
    became imbalanced because of dropouts presumably unrelated
    to the experiment. This can cause a larger effect than
    what is being measured, but the lack of statistical
    significance in the proportion of the various groups
    causes this to be disregarded.

    Huh? Sorry, but experimental mortality at random is
    accounted for - if you are suggesting a mortality effect not
    completely at random, there are accomodations for this. You
    know better than to through up this red herring, Herman.

    Quoted message said:

    Many problems are very high dimensional. In these, if one
    considers the variables and attempts to find what is
    relevant, classical statistics cannot be used at all. But
    the medical investigators, steeped in p values, cannot
    even attempts what should be done.

    Herman - read what you wrote and try again, please.

    Confirmatory studies is the point.

    js

Active in the last 60 minutes

Active in this thread

0 users · 0 guests ·0 bots ·0 total

No signed-in users are active right now.

No known search crawlers active right now.