{
  "id": 187327,
  "title": "CV Evolution",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/187327",
  "author_name": "",
  "post_date": "2020-09-28T13:28:04.305390200Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I want to resume my thoughts about cross validation. My local experiments prove them in some degree, but any discussions your points of view are welcome. So, that's what I call an evolution: </p>\n<ol>\n<li><p><em>Patients independent splitting</em>. It means, that one patient may be included in both train and validation set. That's absolutely unfair: after training on second to last measurement, we can predict FVC inside last measurement exactly enough just with constant prediction from previous FVC.</p></li>\n<li><p><em>Patients splitting</em>: training on one part of patients, validation on the other part. It may seem fair, but it's not :) Another example:</p></li>\n</ol>\n<blockquote>\n  <p>FVC_true = [2000, 2100, 2200, 2300, 2400, 2500]<br>\n  FVC_pred = [2000, 2000, 2000, 2000, 2000, 2000]</p>\n</blockquote>\n<p>SO, we can see, that <code>mae(FVC_true, FVC_pred) = 250</code>. But in this competition only last three measurements are included in scoring, so the most appropriate score is <code>mae(FVC_true[-3:], FVC_pred[-3:]) = 400</code>. That's the valuable difference that may push your optimal solution to constant transformation by this type of score. (mae score was taken as the simplest metric, discussion interpolates on any score)</p>\n<ol>\n<li><em>Patients splitting with last three scoring</em>. That's what I mention above and it seems the best score for cross validation.</li>\n</ol>",
  "messages": [
    {
      "id": "1030204",
      "postDate": "09/28/2020 13:28:04",
      "content": "<p>I want to resume my thoughts about cross validation. My local experiments prove them in some degree, but any discussions your points of view are welcome. So, that's what I call an evolution: </p>\n<ol>\n<li><p><em>Patients independent splitting</em>. It means, that one patient may be included in both train and validation set. That's absolutely unfair: after training on second to last measurement, we can predict FVC inside last measurement exactly enough just with constant prediction from previous FVC.</p></li>\n<li><p><em>Patients splitting</em>: training on one part of patients, validation on the other part. It may seem fair, but it's not :) Another example:</p></li>\n</ol>\n<blockquote>\n  <p>FVC_true = [2000, 2100, 2200, 2300, 2400, 2500]<br>\n  FVC_pred = [2000, 2000, 2000, 2000, 2000, 2000]</p>\n</blockquote>\n<p>SO, we can see, that <code>mae(FVC_true, FVC_pred) = 250</code>. But in this competition only last three measurements are included in scoring, so the most appropriate score is <code>mae(FVC_true[-3:], FVC_pred[-3:]) = 400</code>. That's the valuable difference that may push your optimal solution to constant transformation by this type of score. (mae score was taken as the simplest metric, discussion interpolates on any score)</p>\n<ol>\n<li><em>Patients splitting with last three scoring</em>. That's what I mention above and it seems the best score for cross validation.</li>\n</ol>",
      "rawMarkdown": "I want to resume my thoughts about cross validation. My local experiments prove them in some degree, but any discussions your points of view are welcome. So, that's what I call an evolution: \n\n1.  *Patients independent splitting*. It means, that one patient may be included in both train and validation set. That's absolutely unfair: after training on second to last measurement, we can predict FVC inside last measurement exactly enough just with constant prediction from previous FVC.\n\n2. *Patients splitting*: training on one part of patients, validation on the other part. It may seem fair, but it's not :) Another example:\n> FVC_true = [2000, 2100, 2200, 2300, 2400, 2500]\n> FVC_pred = [2000, 2000, 2000, 2000, 2000, 2000]\n\nSO, we can see, that `mae(FVC_true, FVC_pred) = 250`. But in this competition only last three measurements are included in scoring, so the most appropriate score is `mae(FVC_true[-3:], FVC_pred[-3:]) = 400`. That's the valuable difference that may push your optimal solution to constant transformation by this type of score. (mae score was taken as the simplest metric, discussion interpolates on any score)\n\n3. *Patients splitting with last three scoring*. That's what I mention above and it seems the best score for cross validation.",
      "votes": null
    },
    {
      "id": "1031165",
      "postDate": "09/29/2020 09:08:22",
      "content": "<p>good job ! good</p>",
      "rawMarkdown": "good job ! good",
      "votes": null
    },
    {
      "id": "1031168",
      "postDate": "09/29/2020 09:14:20",
      "content": "<p>I'm doing CV with this method.</p>",
      "rawMarkdown": "I'm doing CV with this method.",
      "votes": null
    },
    {
      "id": "1031213",
      "postDate": "09/29/2020 10:11:36",
      "content": "<p>Nice to hear</p>",
      "rawMarkdown": "Nice to hear",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1031168,
      "author_name": "resistance0108",
      "author_url": "",
      "post_date": "09/29/2020 09:14:20",
      "content": "<p>I'm doing CV with this method.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1031213,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/29/2020 10:11:36",
          "content": "<p>Nice to hear</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1031165,
      "author_name": "naim99",
      "author_url": "",
      "post_date": "09/29/2020 09:08:22",
      "content": "<p>good job ! good</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1030204": "I want to resume my thoughts about cross validation. My local experiments prove them in some degree, but any discussions your points of view are welcome. So, that's what I call an evolution: \n\n1.  *Patients independent splitting*. It means, that one patient may be included in both train and validation set. That's absolutely unfair: after training on second to last measurement, we can predict FVC inside last measurement exactly enough just with constant prediction from previous FVC.\n\n2. *Patients splitting*: training on one part of patients, validation on the other part. It may seem fair, but it's not :) Another example:\n> FVC_true = [2000, 2100, 2200, 2300, 2400, 2500]\n> FVC_pred = [2000, 2000, 2000, 2000, 2000, 2000]\n\nSO, we can see, that `mae(FVC_true, FVC_pred) = 250`. But in this competition only last three measurements are included in scoring, so the most appropriate score is `mae(FVC_true[-3:], FVC_pred[-3:]) = 400`. That's the valuable difference that may push your optimal solution to constant transformation by this type of score. (mae score was taken as the simplest metric, discussion interpolates on any score)\n\n3. *Patients splitting with last three scoring*. That's what I mention above and it seems the best score for cross validation.",
    "1031165": "good job ! good",
    "1031168": "I'm doing CV with this method.",
    "1031213": "Nice to hear"
  },
  "source": "meta"
}