{
  "id": 71603,
  "title": "Single split vs CV",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71603",
  "author_name": "Thomas Yokota",
  "post_date": "2018-11-15T04:01:41.897000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I've been battling a flu so perhaps my thinking is somewhat cloudy...\nI've spent some time today running single split iterations and found some interesting things:\n- the validation score per model using single split is lower than CV\n- threshold remains around 0.33-0.34 for \"optimal\" f1 cut-off\n- a subjective blend based on ranking order of single split validation improves overall score - closer to my single model approach.</p>\n\n<p>Does the blend of many single split models stabilize the variance enough? Or are single split approaches at greater risk of sliding on private? Just curious on everyone elses' thoughts as I've never competed in a kernel-only competition, but this time constraint is quite interesting in terms of deciding modeling, validation, etc.</p>\n\n<p>Also, I know the competition only began, but if anyone would like to team up to discuss more about ideas, findings, meaning of life, then please lmk - please use the internal Kaggle message feature so as not to steer the thread. Mahalo!</p>",
  "messages": [
    {
      "id": 421484,
      "postDate": "2018-11-15T04:01:41.897Z",
      "content": "<p>I've been battling a flu so perhaps my thinking is somewhat cloudy...\nI've spent some time today running single split iterations and found some interesting things:\n- the validation score per model using single split is lower than CV\n- threshold remains around 0.33-0.34 for \"optimal\" f1 cut-off\n- a subjective blend based on ranking order of single split validation improves overall score - closer to my single model approach.</p>\n\n<p>Does the blend of many single split models stabilize the variance enough? Or are single split approaches at greater risk of sliding on private? Just curious on everyone elses' thoughts as I've never competed in a kernel-only competition, but this time constraint is quite interesting in terms of deciding modeling, validation, etc.</p>\n\n<p>Also, I know the competition only began, but if anyone would like to team up to discuss more about ideas, findings, meaning of life, then please lmk - please use the internal Kaggle message feature so as not to steer the thread. Mahalo!</p>",
      "rawMarkdown": "I've been battling a flu so perhaps my thinking is somewhat cloudy...\nI've spent some time today running single split iterations and found some interesting things:\n- the validation score per model using single split is lower than CV\n- threshold remains around 0.33-0.34 for \"optimal\" f1 cut-off\n- a subjective blend based on ranking order of single split validation improves overall score - closer to my single model approach.\n\nDoes the blend of many single split models stabilize the variance enough? Or are single split approaches at greater risk of sliding on private? Just curious on everyone elses' thoughts as I've never competed in a kernel-only competition, but this time constraint is quite interesting in terms of deciding modeling, validation, etc.\n\nAlso, I know the competition only began, but if anyone would like to team up to discuss more about ideas, findings, meaning of life, then please lmk - please use the internal Kaggle message feature so as not to steer the thread. Mahalo!",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "421484": "I've been battling a flu so perhaps my thinking is somewhat cloudy...\nI've spent some time today running single split iterations and found some interesting things:\n- the validation score per model using single split is lower than CV\n- threshold remains around 0.33-0.34 for \"optimal\" f1 cut-off\n- a subjective blend based on ranking order of single split validation improves overall score - closer to my single model approach.\n\nDoes the blend of many single split models stabilize the variance enough? Or are single split approaches at greater risk of sliding on private? Just curious on everyone elses' thoughts as I've never competed in a kernel-only competition, but this time constraint is quite interesting in terms of deciding modeling, validation, etc.\n\nAlso, I know the competition only began, but if anyone would like to team up to discuss more about ideas, findings, meaning of life, then please lmk - please use the internal Kaggle message feature so as not to steer the thread. Mahalo!"
  }
}