{
  "id": 87405,
  "title": "CV for Histopathologic Cancer",
  "url": "/competitions/histopathologic-cancer-detection/discussion/87405",
  "author_name": "",
  "post_date": "2019-03-31T11:11:48.729905800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi </p>\n\n<p>I have tried a lot of ways to define the CV for this data-set( including taking into account the WSI). But there is always a significant difference between local CV score and Leadership score. Can someone please share the thought about how to define the Cross Validation for this data-set so that there is minimal difference between Local CV Score and Leadership Score.</p>",
  "messages": [
    {
      "id": "504299",
      "postDate": "03/31/2019 11:11:48",
      "content": "<p>Hi </p>\n\n<p>I have tried a lot of ways to define the CV for this data-set( including taking into account the WSI). But there is always a significant difference between local CV score and Leadership score. Can someone please share the thought about how to define the Cross Validation for this data-set so that there is minimal difference between Local CV Score and Leadership Score.</p>",
      "rawMarkdown": "Hi \n\nI have tried a lot of ways to define the CV for this data-set( including taking into account the WSI). But there is always a significant difference between local CV score and Leadership score. Can someone please share the thought about how to define the Cross Validation for this data-set so that there is minimal difference between Local CV Score and Leadership Score.",
      "votes": null
    },
    {
      "id": "504949",
      "postDate": "04/01/2019 11:11:09",
      "content": "<p>I tried a combination of things. Splitting by WSI + data augmentation (this will help if test data differs from training data, which might be the reason for the difference between CV score and LB score) + k-fold validation (but it's obviously more computationally expensive, so I personally didn't use it). </p>",
      "rawMarkdown": "I tried a combination of things. Splitting by WSI + data augmentation (this will help if test data differs from training data, which might be the reason for the difference between CV score and LB score) + k-fold validation (but it's obviously more computationally expensive, so I personally didn't use it).",
      "votes": null
    },
    {
      "id": "505131",
      "postDate": "04/01/2019 15:45:41",
      "content": "<p>WSI + 10 fold CV for each model</p>",
      "rawMarkdown": "WSI + 10 fold CV for each model",
      "votes": null
    },
    {
      "id": "505316",
      "postDate": "04/01/2019 21:16:01",
      "content": "<p>My suggestion is to do LB shakeup prediction. Because the std is fixed in LB, you can do nothing about it other than predict the expected std.</p>",
      "rawMarkdown": "My suggestion is to do LB shakeup prediction. Because the std is fixed in LB, you can do nothing about it other than predict the expected std.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 504949,
      "author_name": "ivanpan",
      "author_url": "",
      "post_date": "04/01/2019 11:11:09",
      "content": "<p>I tried a combination of things. Splitting by WSI + data augmentation (this will help if test data differs from training data, which might be the reason for the difference between CV score and LB score) + k-fold validation (but it's obviously more computationally expensive, so I personally didn't use it). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 505131,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "04/01/2019 15:45:41",
      "content": "<p>WSI + 10 fold CV for each model</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 505316,
      "author_name": "kokecacao",
      "author_url": "",
      "post_date": "04/01/2019 21:16:01",
      "content": "<p>My suggestion is to do LB shakeup prediction. Because the std is fixed in LB, you can do nothing about it other than predict the expected std.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "504299": "Hi \n\nI have tried a lot of ways to define the CV for this data-set( including taking into account the WSI). But there is always a significant difference between local CV score and Leadership score. Can someone please share the thought about how to define the Cross Validation for this data-set so that there is minimal difference between Local CV Score and Leadership Score.",
    "504949": "I tried a combination of things. Splitting by WSI + data augmentation (this will help if test data differs from training data, which might be the reason for the difference between CV score and LB score) + k-fold validation (but it's obviously more computationally expensive, so I personally didn't use it).",
    "505131": "WSI + 10 fold CV for each model",
    "505316": "My suggestion is to do LB shakeup prediction. Because the std is fixed in LB, you can do nothing about it other than predict the expected std."
  },
  "source": "meta"
}