{
  "id": 277809,
  "title": "Clinically useful?",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/277809",
  "author_name": "",
  "post_date": "2021-10-11T11:37:21.123492300Z",
  "votes": 8,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm starting to think that there isn't sufficient data to create something that eventually would be clinically useful here.<br>\nIt's me or someone else is as skeptical as me?</p>",
  "messages": [
    {
      "id": "1541296",
      "postDate": "10/11/2021 11:37:21",
      "content": "<p>I'm starting to think that there isn't sufficient data to create something that eventually would be clinically useful here.<br>\nIt's me or someone else is as skeptical as me?</p>",
      "rawMarkdown": "I'm starting to think that there isn't sufficient data to create something that eventually would be clinically useful here.\nIt's me or someone else is as skeptical as me?",
      "votes": null
    },
    {
      "id": "1541642",
      "postDate": "10/11/2021 17:58:28",
      "content": "<p>A simple proof that the models are not useful is the output probabilities. I would expect some predictions to be close to 0% or 100%, but most of my probabilities are around 50% (10-fold RepeatedStratifiedKFold with AUC 0.65, 100-fold StratifiedKFold with AUC 0.82).</p>\n<p>If someone has really high or low probabilities (0 or 1), then there is a chance that the model is actually useful. However, it is also quite easy to artificially increase the probabilities by simply using non-calibrated or even non-probabilistic models (linear regression, decision tree, …). Another way to get higher probabilities is train_test_split (with random seed tuning).</p>\n<p>For this reason, I am also very skeptical. Maybe the best strategy to win this competition is to set <code>MGMT_value=0.5</code> and then randomly generate some indices where you set <code>MGMT_value=1</code> or <code>MGMT_value=0</code>. However, I still hope that someone has solved this classification problem somehow. </p>",
      "rawMarkdown": "A simple proof that the models are not useful is the output probabilities. I would expect some predictions to be close to 0% or 100%, but most of my probabilities are around 50% (10-fold RepeatedStratifiedKFold with AUC 0.65, 100-fold StratifiedKFold with AUC 0.82).\n\nIf someone has really high or low probabilities (0 or 1), then there is a chance that the model is actually useful. However, it is also quite easy to artificially increase the probabilities by simply using non-calibrated or even non-probabilistic models (linear regression, decision tree, ...). Another way to get higher probabilities is train_test_split (with random seed tuning).\n\nFor this reason, I am also very skeptical. Maybe the best strategy to win this competition is to set `MGMT_value=0.5` and then randomly generate some indices where you set `MGMT_value=1` or `MGMT_value=0`. However, I still hope that someone has solved this classification problem somehow.",
      "votes": null
    },
    {
      "id": "1542640",
      "postDate": "10/12/2021 17:47:26",
      "content": "<p>Same…          </p>",
      "rawMarkdown": "Same...",
      "votes": null
    },
    {
      "id": "1542882",
      "postDate": "10/12/2021 22:21:10",
      "content": "<p>Correct me if I'm wrong, but I don't think that is the intention in this or any Kaggle comps.</p>\n<p>More likely, the information gathered by the organisers about which model / approach works best for this problem, will be used on a deeper dataset to make something for clinical application.</p>",
      "rawMarkdown": "Correct me if I'm wrong, but I don't think that is the intention in this or any Kaggle comps.\n\nMore likely, the information gathered by the organisers about which model / approach works best for this problem, will be used on a deeper dataset to make something for clinical application.",
      "votes": null
    },
    {
      "id": "1543272",
      "postDate": "10/13/2021 09:58:02",
      "content": "<p>Good point. This is what I think too. I think they should have required 0 or 1 as an output. In this case the ranking list would have been different.</p>",
      "rawMarkdown": "Good point. This is what I think too. I think they should have required 0 or 1 as an output. In this case the ranking list would have been different.",
      "votes": null
    },
    {
      "id": "1543506",
      "postDate": "10/13/2021 14:43:40",
      "content": "<p>Understand <a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">@reubenschmidt</a> but do you think the approach they might have from here is going te be useful for usage on a deeper dataset? </p>",
      "rawMarkdown": "Understand @reubenschmidt but do you think the approach they might have from here is going te be useful for usage on a deeper dataset?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1541642,
      "author_name": "lars123",
      "author_url": "",
      "post_date": "10/11/2021 17:58:28",
      "content": "<p>A simple proof that the models are not useful is the output probabilities. I would expect some predictions to be close to 0% or 100%, but most of my probabilities are around 50% (10-fold RepeatedStratifiedKFold with AUC 0.65, 100-fold StratifiedKFold with AUC 0.82).</p>\n<p>If someone has really high or low probabilities (0 or 1), then there is a chance that the model is actually useful. However, it is also quite easy to artificially increase the probabilities by simply using non-calibrated or even non-probabilistic models (linear regression, decision tree, …). Another way to get higher probabilities is train_test_split (with random seed tuning).</p>\n<p>For this reason, I am also very skeptical. Maybe the best strategy to win this competition is to set <code>MGMT_value=0.5</code> and then randomly generate some indices where you set <code>MGMT_value=1</code> or <code>MGMT_value=0</code>. However, I still hope that someone has solved this classification problem somehow. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1543272,
          "author_name": "jhasanov",
          "author_url": "",
          "post_date": "10/13/2021 09:58:02",
          "content": "<p>Good point. This is what I think too. I think they should have required 0 or 1 as an output. In this case the ranking list would have been different.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1542640,
      "author_name": "kfk42kfk",
      "author_url": "",
      "post_date": "10/12/2021 17:47:26",
      "content": "<p>Same…          </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1542882,
      "author_name": "reubenschmidt",
      "author_url": "",
      "post_date": "10/12/2021 22:21:10",
      "content": "<p>Correct me if I'm wrong, but I don't think that is the intention in this or any Kaggle comps.</p>\n<p>More likely, the information gathered by the organisers about which model / approach works best for this problem, will be used on a deeper dataset to make something for clinical application.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1543506,
          "author_name": "jolasa",
          "author_url": "",
          "post_date": "10/13/2021 14:43:40",
          "content": "<p>Understand <a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">@reubenschmidt</a> but do you think the approach they might have from here is going te be useful for usage on a deeper dataset? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1541296": "I'm starting to think that there isn't sufficient data to create something that eventually would be clinically useful here.\nIt's me or someone else is as skeptical as me?",
    "1541642": "A simple proof that the models are not useful is the output probabilities. I would expect some predictions to be close to 0% or 100%, but most of my probabilities are around 50% (10-fold RepeatedStratifiedKFold with AUC 0.65, 100-fold StratifiedKFold with AUC 0.82).\n\nIf someone has really high or low probabilities (0 or 1), then there is a chance that the model is actually useful. However, it is also quite easy to artificially increase the probabilities by simply using non-calibrated or even non-probabilistic models (linear regression, decision tree, ...). Another way to get higher probabilities is train_test_split (with random seed tuning).\n\nFor this reason, I am also very skeptical. Maybe the best strategy to win this competition is to set `MGMT_value=0.5` and then randomly generate some indices where you set `MGMT_value=1` or `MGMT_value=0`. However, I still hope that someone has solved this classification problem somehow.",
    "1542640": "Same...",
    "1542882": "Correct me if I'm wrong, but I don't think that is the intention in this or any Kaggle comps.\n\nMore likely, the information gathered by the organisers about which model / approach works best for this problem, will be used on a deeper dataset to make something for clinical application.",
    "1543272": "Good point. This is what I think too. I think they should have required 0 or 1 as an output. In this case the ranking list would have been different.",
    "1543506": "Understand @reubenschmidt but do you think the approach they might have from here is going te be useful for usage on a deeper dataset?"
  },
  "source": "meta"
}