{
  "id": 189212,
  "title": "Lesson learned sharing",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/189212",
  "author_name": "",
  "post_date": "2020-10-07T02:00:32.134160100Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First thing first, congratulations for the winner, what an amazing Competition</p>\n<p>It would be nice if we can share lesson learned from this competition here</p>\n<p>Especially about CV, shakeups, What kind of CV winners use, how to choose the best results? etc</p>",
  "messages": [
    {
      "id": "1040142",
      "postDate": "10/07/2020 02:00:32",
      "content": "<p>First thing first, congratulations for the winner, what an amazing Competition</p>\n<p>It would be nice if we can share lesson learned from this competition here</p>\n<p>Especially about CV, shakeups, What kind of CV winners use, how to choose the best results? etc</p>",
      "rawMarkdown": "First thing first, congratulations for the winner, what an amazing Competition\n\nIt would be nice if we can share lesson learned from this competition here\n\nEspecially about CV, shakeups, What kind of CV winners use, how to choose the best results? etc",
      "votes": null
    },
    {
      "id": "1044220",
      "postDate": "10/09/2020 15:53:07",
      "content": "<p>I'm interested in learning others' learned lessons too =) I'll share mines here:</p>\n<p>For my last submissions I was using an 8-fold cross-validation (since I can divide 176 patients in 8 groups of the same size) and I computed mean and standard deviation of the test error. This way I could know the 95% confidence interval and know what to expect. I also balanced each fold for having the same distribution of male and female patients. That's more than enough for model selection, but for doing the last selection I'll be doing randomized and repeated k-fold validations, from now on.</p>\n<p>It's also important to understand how data-leakage may occur. I always check the test data thoroughly, so that I can build my internal tests the same way. In this case, I went for dividing the data patient-wise (avoiding leaking of patient data between sets) and I only used data that I could extract from the first week's row.</p>",
      "rawMarkdown": "I'm interested in learning others' learned lessons too =) I'll share mines here:\n\nFor my last submissions I was using an 8-fold cross-validation (since I can divide 176 patients in 8 groups of the same size) and I computed mean and standard deviation of the test error. This way I could know the 95% confidence interval and know what to expect. I also balanced each fold for having the same distribution of male and female patients. That's more than enough for model selection, but for doing the last selection I'll be doing randomized and repeated k-fold validations, from now on.\n\nIt's also important to understand how data-leakage may occur. I always check the test data thoroughly, so that I can build my internal tests the same way. In this case, I went for dividing the data patient-wise (avoiding leaking of patient data between sets) and I only used data that I could extract from the first week's row.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1044220,
      "author_name": "dcasbol",
      "author_url": "",
      "post_date": "10/09/2020 15:53:07",
      "content": "<p>I'm interested in learning others' learned lessons too =) I'll share mines here:</p>\n<p>For my last submissions I was using an 8-fold cross-validation (since I can divide 176 patients in 8 groups of the same size) and I computed mean and standard deviation of the test error. This way I could know the 95% confidence interval and know what to expect. I also balanced each fold for having the same distribution of male and female patients. That's more than enough for model selection, but for doing the last selection I'll be doing randomized and repeated k-fold validations, from now on.</p>\n<p>It's also important to understand how data-leakage may occur. I always check the test data thoroughly, so that I can build my internal tests the same way. In this case, I went for dividing the data patient-wise (avoiding leaking of patient data between sets) and I only used data that I could extract from the first week's row.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040142": "First thing first, congratulations for the winner, what an amazing Competition\n\nIt would be nice if we can share lesson learned from this competition here\n\nEspecially about CV, shakeups, What kind of CV winners use, how to choose the best results? etc",
    "1044220": "I'm interested in learning others' learned lessons too =) I'll share mines here:\n\nFor my last submissions I was using an 8-fold cross-validation (since I can divide 176 patients in 8 groups of the same size) and I computed mean and standard deviation of the test error. This way I could know the 95% confidence interval and know what to expect. I also balanced each fold for having the same distribution of male and female patients. That's more than enough for model selection, but for doing the last selection I'll be doing randomized and repeated k-fold validations, from now on.\n\nIt's also important to understand how data-leakage may occur. I always check the test data thoroughly, so that I can build my internal tests the same way. In this case, I went for dividing the data patient-wise (avoiding leaking of patient data between sets) and I only used data that I could extract from the first week's row."
  },
  "source": "meta"
}