{
  "id": 189301,
  "title": "Two constant values  won 251st. This might imply reasons for the huge shakeup",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/189301",
  "author_name": "",
  "post_date": "2020-10-07T08:12:23.620763100Z",
  "votes": 33,
  "comment_count": 6,
  "views": 0,
  "content": "<p>One of my solutions is only made of two magic constant values, 144.998 and 284.107.  I assume that the FVC of all patients will decreased by 144.998, and the confidence is 284.107 for all cases. <strong>No any feature</strong> is used in this solution. This solution got a private score of -6.8736, and the rank is 251st. Considering that the best private score is -6.8305, I think all the solutions we found in this competition are far from satisfactory.  We can still learn a lot  from this competition. But we should not over-interpret the winning solutions because most of them are due to luck.   </p>",
  "messages": [
    {
      "id": "1040557",
      "postDate": "10/07/2020 08:12:23",
      "content": "<p>One of my solutions is only made of two magic constant values, 144.998 and 284.107.  I assume that the FVC of all patients will decreased by 144.998, and the confidence is 284.107 for all cases. <strong>No any feature</strong> is used in this solution. This solution got a private score of -6.8736, and the rank is 251st. Considering that the best private score is -6.8305, I think all the solutions we found in this competition are far from satisfactory.  We can still learn a lot  from this competition. But we should not over-interpret the winning solutions because most of them are due to luck.   </p>",
      "rawMarkdown": "One of my solutions is only made of two magic constant values, 144.998 and 284.107.  I assume that the FVC of all patients will decreased by 144.998, and the confidence is 284.107 for all cases. **No any feature** is used in this solution. This solution got a private score of -6.8736, and the rank is 251st. Considering that the best private score is -6.8305, I think all the solutions we found in this competition are far from satisfactory.  We can still learn a lot  from this competition. But we should not over-interpret the winning solutions because most of them are due to luck.",
      "votes": null
    },
    {
      "id": "1040561",
      "postDate": "10/07/2020 08:18:38",
      "content": "<p>Oh my goodness! So my trained models did far worse than a constant guess! That is so depressing </p>",
      "rawMarkdown": "Oh my goodness! So my trained models did far worse than a constant guess! That is so depressing",
      "votes": null
    },
    {
      "id": "1040577",
      "postDate": "10/07/2020 08:25:58",
      "content": "<p>I also felt depressed when I saw that all my trained models are worse than the constant-guess model. This means that there is super weak correlation between features and temporal change of FVC. </p>",
      "rawMarkdown": "I also felt depressed when I saw that all my trained models are worse than the constant-guess model. This means that there is super weak correlation between features and temporal change of FVC.",
      "votes": null
    },
    {
      "id": "1040619",
      "postDate": "10/07/2020 09:01:48",
      "content": "<p>My best score is based on somewhat similar approach, but instead of constant values make predictions proportional to initial FVC observations. Namely: pred_fvc = 0.9418 * (initial FVC), confidence=0.1108 * (initial FVC). In private LB this gives 258th place and score of -6.8745</p>\n<p>Given that the winners score is only about 0.04 higher, I kinda wonder if some small linear decay + randomness would already be enough to get us there.  Kinda disappointing, but matches my experience with this competion and dataset. I was only focusing on CT scans/images, but wasn't able to get anything to work properly.</p>",
      "rawMarkdown": "My best score is based on somewhat similar approach, but instead of constant values make predictions proportional to initial FVC observations. Namely: pred_fvc = 0.9418 * (initial FVC), confidence=0.1108 * (initial FVC). In private LB this gives 258th place and score of -6.8745\n\nGiven that the winners score is only about 0.04 higher, I kinda wonder if some small linear decay + randomness would already be enough to get us there.  Kinda disappointing, but matches my experience with this competion and dataset. I was only focusing on CT scans/images, but wasn't able to get anything to work properly.",
      "votes": null
    },
    {
      "id": "1040649",
      "postDate": "10/07/2020 09:15:16",
      "content": "<p>I agree. If you add the 'correct' randomness, you may get a better score and much better rank. The randomness fully determines the final rank, which is really unfortunate for a competition. </p>",
      "rawMarkdown": "I agree. If you add the 'correct' randomness, you may get a better score and much better rank. The randomness fully determines the final rank, which is really unfortunate for a competition.",
      "votes": null
    },
    {
      "id": "1041238",
      "postDate": "10/07/2020 16:28:41",
      "content": "<p>To be fair, there was relatively little training data considering the number of features, especially if one tried to use the imaging data in any kind of sophisticated way (personally my final predictions were a confidence-weighted average of the predictions from my tabular-data-only model and the one also using imaging). The observations we had for the target variable were also noisy due to imprecise measurements (I dealt with that issue using a probabilistic smoothing module on the training data, to help my model focus on the underlying trends rather than the noise).</p>",
      "rawMarkdown": "To be fair, there was relatively little training data considering the number of features, especially if one tried to use the imaging data in any kind of sophisticated way (personally my final predictions were a confidence-weighted average of the predictions from my tabular-data-only model and the one also using imaging). The observations we had for the target variable were also noisy due to imprecise measurements (I dealt with that issue using a probabilistic smoothing module on the training data, to help my model focus on the underlying trends rather than the noise).",
      "votes": null
    },
    {
      "id": "1041284",
      "postDate": "10/07/2020 17:07:58",
      "content": "<p>This raises the question whether there was something to learn in the first place. Machine Learning can only work if there is something in the data that can help make predictions and by the looks of it that seems to be very little.<br>\nGiven that the best submission is only 0.04 better than such a simple baseline makes me wonder if the solutions we have produced here on Kaggle have any worth at all (there was still a gap of about 2 to the perfect score).</p>\n<p>I'm not a domain expert so I can only wonder if the predictions made by these Machine Learning methods are in any useful range.</p>",
      "rawMarkdown": "This raises the question whether there was something to learn in the first place. Machine Learning can only work if there is something in the data that can help make predictions and by the looks of it that seems to be very little.\nGiven that the best submission is only 0.04 better than such a simple baseline makes me wonder if the solutions we have produced here on Kaggle have any worth at all (there was still a gap of about 2 to the perfect score).\n\nI'm not a domain expert so I can only wonder if the predictions made by these Machine Learning methods are in any useful range.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1040561,
      "author_name": "cascadenite",
      "author_url": "",
      "post_date": "10/07/2020 08:18:38",
      "content": "<p>Oh my goodness! So my trained models did far worse than a constant guess! That is so depressing </p>",
      "votes": null,
      "replies": [
        {
          "id": 1040577,
          "author_name": "ningwang1990",
          "author_url": "",
          "post_date": "10/07/2020 08:25:58",
          "content": "<p>I also felt depressed when I saw that all my trained models are worse than the constant-guess model. This means that there is super weak correlation between features and temporal change of FVC. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040619,
      "author_name": "herrahuu",
      "author_url": "",
      "post_date": "10/07/2020 09:01:48",
      "content": "<p>My best score is based on somewhat similar approach, but instead of constant values make predictions proportional to initial FVC observations. Namely: pred_fvc = 0.9418 * (initial FVC), confidence=0.1108 * (initial FVC). In private LB this gives 258th place and score of -6.8745</p>\n<p>Given that the winners score is only about 0.04 higher, I kinda wonder if some small linear decay + randomness would already be enough to get us there.  Kinda disappointing, but matches my experience with this competion and dataset. I was only focusing on CT scans/images, but wasn't able to get anything to work properly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040649,
          "author_name": "ningwang1990",
          "author_url": "",
          "post_date": "10/07/2020 09:15:16",
          "content": "<p>I agree. If you add the 'correct' randomness, you may get a better score and much better rank. The randomness fully determines the final rank, which is really unfortunate for a competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041238,
      "author_name": "archimedus",
      "author_url": "",
      "post_date": "10/07/2020 16:28:41",
      "content": "<p>To be fair, there was relatively little training data considering the number of features, especially if one tried to use the imaging data in any kind of sophisticated way (personally my final predictions were a confidence-weighted average of the predictions from my tabular-data-only model and the one also using imaging). The observations we had for the target variable were also noisy due to imprecise measurements (I dealt with that issue using a probabilistic smoothing module on the training data, to help my model focus on the underlying trends rather than the noise).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1041284,
      "author_name": "bjaeger",
      "author_url": "",
      "post_date": "10/07/2020 17:07:58",
      "content": "<p>This raises the question whether there was something to learn in the first place. Machine Learning can only work if there is something in the data that can help make predictions and by the looks of it that seems to be very little.<br>\nGiven that the best submission is only 0.04 better than such a simple baseline makes me wonder if the solutions we have produced here on Kaggle have any worth at all (there was still a gap of about 2 to the perfect score).</p>\n<p>I'm not a domain expert so I can only wonder if the predictions made by these Machine Learning methods are in any useful range.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040557": "One of my solutions is only made of two magic constant values, 144.998 and 284.107.  I assume that the FVC of all patients will decreased by 144.998, and the confidence is 284.107 for all cases. **No any feature** is used in this solution. This solution got a private score of -6.8736, and the rank is 251st. Considering that the best private score is -6.8305, I think all the solutions we found in this competition are far from satisfactory.  We can still learn a lot  from this competition. But we should not over-interpret the winning solutions because most of them are due to luck.",
    "1040561": "Oh my goodness! So my trained models did far worse than a constant guess! That is so depressing",
    "1040577": "I also felt depressed when I saw that all my trained models are worse than the constant-guess model. This means that there is super weak correlation between features and temporal change of FVC.",
    "1040619": "My best score is based on somewhat similar approach, but instead of constant values make predictions proportional to initial FVC observations. Namely: pred_fvc = 0.9418 * (initial FVC), confidence=0.1108 * (initial FVC). In private LB this gives 258th place and score of -6.8745\n\nGiven that the winners score is only about 0.04 higher, I kinda wonder if some small linear decay + randomness would already be enough to get us there.  Kinda disappointing, but matches my experience with this competion and dataset. I was only focusing on CT scans/images, but wasn't able to get anything to work properly.",
    "1040649": "I agree. If you add the 'correct' randomness, you may get a better score and much better rank. The randomness fully determines the final rank, which is really unfortunate for a competition.",
    "1041238": "To be fair, there was relatively little training data considering the number of features, especially if one tried to use the imaging data in any kind of sophisticated way (personally my final predictions were a confidence-weighted average of the predictions from my tabular-data-only model and the one also using imaging). The observations we had for the target variable were also noisy due to imprecise measurements (I dealt with that issue using a probabilistic smoothing module on the training data, to help my model focus on the underlying trends rather than the noise).",
    "1041284": "This raises the question whether there was something to learn in the first place. Machine Learning can only work if there is something in the data that can help make predictions and by the looks of it that seems to be very little.\nGiven that the best submission is only 0.04 better than such a simple baseline makes me wonder if the solutions we have produced here on Kaggle have any worth at all (there was still a gap of about 2 to the perfect score).\n\nI'm not a domain expert so I can only wonder if the predictions made by these Machine Learning methods are in any useful range."
  },
  "source": "meta"
}