{
  "id": 176826,
  "title": "Are the test FVC values free of outliers?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/176826",
  "author_name": "",
  "post_date": "2020-08-23T16:05:24.842110800Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Dear <a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a>,<br>\nThe FVC values in the training data seem to have outliers. I was wondering how reliable the test values are. Because, if there are no outliers in the test data, then the training data needs to be cleaned.</p>",
  "messages": [
    {
      "id": "982729",
      "postDate": "08/23/2020 16:05:24",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a>,<br>\nThe FVC values in the training data seem to have outliers. I was wondering how reliable the test values are. Because, if there are no outliers in the test data, then the training data needs to be cleaned.</p>",
      "rawMarkdown": "Dear @ahmedhshahin,\n\nThe FVC values in the training data seem to have outliers. I was wondering how reliable the test values are. Because, if there are no outliers in the test data, then the training data needs to be cleaned.",
      "votes": null
    },
    {
      "id": "983646",
      "postDate": "08/24/2020 13:23:48",
      "content": "<p>Hi Tolga, <br>\nThe test and training sets come from the same distribution, thus any patterns in the training set should be expected to happen in the test one. The data has been collected from the clinical sites and split for the sake of competition evaluation. Worth mentioning, it is known that the spirometer - which is used to calculate the FVC - has an inherent noise in its measurements. Also, it is possible for humans to have relatively higher FVC values for several reasons, for example, if these humans are swimmers.</p>",
      "rawMarkdown": "Hi Tolga, \nThe test and training sets come from the same distribution, thus any patterns in the training set should be expected to happen in the test one. The data has been collected from the clinical sites and split for the sake of competition evaluation. Worth mentioning, it is known that the spirometer - which is used to calculate the FVC - has an inherent noise in its measurements. Also, it is possible for humans to have relatively higher FVC values for several reasons, for example, if these humans are swimmers.",
      "votes": null
    },
    {
      "id": "983666",
      "postDate": "08/24/2020 13:40:16",
      "content": "<p>Thank you for clarifying some of the points regarding the distribution of the data.</p>\n<p>I was thinking of outliers that cause abrupt changes in the trend of a given patient. For instance, if a patient couldn't blow the spirometer properly in the 3rd measurement of 10 measurements. I see some patients with such outlier behaviors. I was wondering if the same exists in the private test data.</p>",
      "rawMarkdown": "Thank you for clarifying some of the points regarding the distribution of the data.\n\nI was thinking of outliers that cause abrupt changes in the trend of a given patient. For instance, if a patient couldn't blow the spirometer properly in the 3rd measurement of 10 measurements. I see some patients with such outlier behaviors. I was wondering if the same exists in the private test data.",
      "votes": null
    },
    {
      "id": "983679",
      "postDate": "08/24/2020 13:55:23",
      "content": "<p>Yes, I see your point. Actually this is one of the drawbacks of the FVC tests. However, even with this noise and outliers that can happen due to several reasons (some of them are related to the patient as you mentioned), FVC remains widely considered for IPF diagnosis and is routinely used as a primary endpoint from a clinical perspective. </p>",
      "rawMarkdown": "Yes, I see your point. Actually this is one of the drawbacks of the FVC tests. However, even with this noise and outliers that can happen due to several reasons (some of them are related to the patient as you mentioned), FVC remains widely considered for IPF diagnosis and is routinely used as a primary endpoint from a clinical perspective.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 983646,
      "author_name": "ahmedhshahin",
      "author_url": "",
      "post_date": "08/24/2020 13:23:48",
      "content": "<p>Hi Tolga, <br>\nThe test and training sets come from the same distribution, thus any patterns in the training set should be expected to happen in the test one. The data has been collected from the clinical sites and split for the sake of competition evaluation. Worth mentioning, it is known that the spirometer - which is used to calculate the FVC - has an inherent noise in its measurements. Also, it is possible for humans to have relatively higher FVC values for several reasons, for example, if these humans are swimmers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 983666,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "08/24/2020 13:40:16",
          "content": "<p>Thank you for clarifying some of the points regarding the distribution of the data.</p>\n<p>I was thinking of outliers that cause abrupt changes in the trend of a given patient. For instance, if a patient couldn't blow the spirometer properly in the 3rd measurement of 10 measurements. I see some patients with such outlier behaviors. I was wondering if the same exists in the private test data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 983679,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "08/24/2020 13:55:23",
          "content": "<p>Yes, I see your point. Actually this is one of the drawbacks of the FVC tests. However, even with this noise and outliers that can happen due to several reasons (some of them are related to the patient as you mentioned), FVC remains widely considered for IPF diagnosis and is routinely used as a primary endpoint from a clinical perspective. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "982729": "Dear @ahmedhshahin,\n\nThe FVC values in the training data seem to have outliers. I was wondering how reliable the test values are. Because, if there are no outliers in the test data, then the training data needs to be cleaned.",
    "983646": "Hi Tolga, \nThe test and training sets come from the same distribution, thus any patterns in the training set should be expected to happen in the test one. The data has been collected from the clinical sites and split for the sake of competition evaluation. Worth mentioning, it is known that the spirometer - which is used to calculate the FVC - has an inherent noise in its measurements. Also, it is possible for humans to have relatively higher FVC values for several reasons, for example, if these humans are swimmers.",
    "983666": "Thank you for clarifying some of the points regarding the distribution of the data.\n\nI was thinking of outliers that cause abrupt changes in the trend of a given patient. For instance, if a patient couldn't blow the spirometer properly in the 3rd measurement of 10 measurements. I see some patients with such outlier behaviors. I was wondering if the same exists in the private test data.",
    "983679": "Yes, I see your point. Actually this is one of the drawbacks of the FVC tests. However, even with this noise and outliers that can happen due to several reasons (some of them are related to the patient as you mentioned), FVC remains widely considered for IPF diagnosis and is routinely used as a primary endpoint from a clinical perspective."
  },
  "source": "meta"
}