# Predicting survival times

**URL:** <https://discourse.datamethods.org/t/predicting-survival-times/3752>\
**Category:** general\
**Tags:** survival-analysis, prediction\
**Created:** [September 18, 2020, 9:43am UTC](https://discourse.datamethods.org/t/predicting-survival-times/3752 "2020-09-18T09:43:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jlevy13](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/jlevy13/32/1892_2.png) [@jlevy13](https://discourse.datamethods.org/u/jlevy13)\
**Post date:** [September 18, 2020, 9:43am UTC](https://discourse.datamethods.org/t/predicting-survival-times/3752/1 "2020-09-18T09:43:51Z")

</div>

Do I need to care for the PH assumption if I want to use the Cox model just for prediction?  
One can think of this model as just another black box, and it performs well on the test set, that just fine. What do you think?

---

<div class="post-metadata">

**Author:** ![albertoca](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/albertoca/32/297_2.png) [@albertoca](https://discourse.datamethods.org/u/albertoca)\
**Post date:** [September 18, 2020, 9:53am UTC](https://discourse.datamethods.org/t/predicting-survival-times/3752/2 "2020-09-18T09:53:08Z")

</div>

I believe that the PH assumption also matters for predictions, which can be biased if the deviation is important.

---

<div class="post-metadata">

**Author:** ![albertoca](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/albertoca/32/297_2.png) [@albertoca](https://discourse.datamethods.org/u/albertoca)\
**Post date:** [September 18, 2020, 6:54pm UTC](https://discourse.datamethods.org/t/predicting-survival-times/3752/3 "2020-09-18T18:54:12Z")

</div>

Here is an example of a plot showing the bias in the prediction of survival, from a Cox model with severe departure from the PH assumption. You can see how the prediction (in red) deviates substantially from the estimate by Kaplan-Meier (in black).

> [@Is the interaction between Clinical Researchers and Statisticians working? My insight after reading four failed clinical trials on immunotherapy](https://discourse.datamethods.org/t/is-the-interaction-between-clinical-researchers-and-statisticians-working-my-insight-after-reading-four-failed-clinical-trials-on-immunotherapy/1949/41):
>
> For example, in the case of the Keynote-061 trial, I show a plot with the prediction from the Cox model (in red) against Kaplan-Meier estimates (black). The inability to capture the real effect of immunotherapy in this scenario is very striking. The reason for bias and reduction in statistical power can be clearly seen in this plot. Not being taught this, the next clinical trial in an almost similar population, with the same strategy, repeated the same statistical plan. Both trials were declared…

---

<div class="post-metadata">

**Author:** ![iabbass](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/iabbass/32/856_2.png) [@iabbass](https://discourse.datamethods.org/u/iabbass)\
**Post date:** [September 19, 2020, 4:31am UTC](https://discourse.datamethods.org/t/predicting-survival-times/3752/4 "2020-09-19T04:31:11Z")

</div>

> [@jlevy13](#):
>
> One can think of this model as just another black box, and it performs well on the test set, that just fine. What do you think?

Have you considered using parametric or flexible parametric survival models if you are only interested in prediction? My understanding is Cox model does not model the underlying risk/hazard which is needed for prediction.

---

<div class="post-metadata">

**Author:** ![jlevy13](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/jlevy13/32/1892_2.png) [@jlevy13](https://discourse.datamethods.org/u/jlevy13)\
**Post date:** [September 20, 2020, 10:46am UTC](https://discourse.datamethods.org/t/predicting-survival-times/3752/5 "2020-09-20T10:46:21Z")

</div>

Yes. This is what I actually do. I am just interested in other arguments against using the Cox model under these circumstances.
