{
  "id": 44366,
  "title": "Train/Val Split",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/44366",
  "author_name": "",
  "post_date": "2017-11-27T19:14:54.379002300Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello everybody,</p>\n\n<p>I was wondering what kind of splits have been used in this competition. \nI see a lot of kernels with random 0.8/0.2 split but that just seems way wrong. May anyone share his/her approach?</p>",
  "messages": [
    {
      "id": "249128",
      "postDate": "11/27/2017 19:14:54",
      "content": "<p>Hello everybody,</p>\n\n<p>I was wondering what kind of splits have been used in this competition. \nI see a lot of kernels with random 0.8/0.2 split but that just seems way wrong. May anyone share his/her approach?</p>",
      "rawMarkdown": "Hello everybody,\n\nI was wondering what kind of splits have been used in this competition. \nI see a lot of kernels with random 0.8/0.2 split but that just seems way wrong. May anyone share his/her approach?",
      "votes": null
    },
    {
      "id": "249232",
      "postDate": "11/28/2017 01:33:03",
      "content": "<p>I'm using random split 0.95:0.05</p>",
      "rawMarkdown": "I'm using random split 0.95:0.05",
      "votes": null
    },
    {
      "id": "249287",
      "postDate": "11/28/2017 03:56:33",
      "content": "<p>split products into 0.95:0.05.</p>",
      "rawMarkdown": "split products into 0.95:0.05.",
      "votes": null
    },
    {
      "id": "249362",
      "postDate": "11/28/2017 09:13:58",
      "content": "<p>but stratified, right?</p>",
      "rawMarkdown": "but stratified, right?",
      "votes": null
    },
    {
      "id": "249364",
      "postDate": "11/28/2017 09:15:52",
      "content": "<p>Aren't you careful to make sure that the train/val data contains all classes?</p>",
      "rawMarkdown": "Aren't you careful to make sure that the train/val data contains all classes?",
      "votes": null
    },
    {
      "id": "249724",
      "postDate": "11/29/2017 04:21:09",
      "content": "<p>This dataset is huge, so you don't need such a high percentage of the data for validation. I use 0.1% of the products stratified by class as my validation set. </p>\n\n<p>For computer vision, it is more important to have as much data go into training your model as you can, than it is to know to five significant digits how much you're overfitting by.</p>\n\n<p>Better would be to train with k-fold cross validation, both for validation accuracy and for ensembling, but there isn't enough time for that for most people's hardware.</p>",
      "rawMarkdown": "This dataset is huge, so you don't need such a high percentage of the data for validation. I use 0.1% of the products stratified by class as my validation set. \n\nFor computer vision, it is more important to have as much data go into training your model as you can, than it is to know to five significant digits how much you're overfitting by.\n\nBetter would be to train with k-fold cross validation, both for validation accuracy and for ensembling, but there isn't enough time for that for most people's hardware.",
      "votes": null
    },
    {
      "id": "250646",
      "postDate": "11/30/2017 07:01:12",
      "content": "<p>no. i used only leaf category.</p>",
      "rawMarkdown": "no. i used only leaf category.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 249232,
      "author_name": "blackarrow3542",
      "author_url": "",
      "post_date": "11/28/2017 01:33:03",
      "content": "<p>I'm using random split 0.95:0.05</p>",
      "votes": null,
      "replies": [
        {
          "id": 249364,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "11/28/2017 09:15:52",
          "content": "<p>Aren't you careful to make sure that the train/val data contains all classes?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 249287,
      "author_name": "mwbyeon",
      "author_url": "",
      "post_date": "11/28/2017 03:56:33",
      "content": "<p>split products into 0.95:0.05.</p>",
      "votes": null,
      "replies": [
        {
          "id": 249362,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "11/28/2017 09:13:58",
          "content": "<p>but stratified, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250646,
          "author_name": "mwbyeon",
          "author_url": "",
          "post_date": "11/30/2017 07:01:12",
          "content": "<p>no. i used only leaf category.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 249724,
      "author_name": "eachshadow",
      "author_url": "",
      "post_date": "11/29/2017 04:21:09",
      "content": "<p>This dataset is huge, so you don't need such a high percentage of the data for validation. I use 0.1% of the products stratified by class as my validation set. </p>\n\n<p>For computer vision, it is more important to have as much data go into training your model as you can, than it is to know to five significant digits how much you're overfitting by.</p>\n\n<p>Better would be to train with k-fold cross validation, both for validation accuracy and for ensembling, but there isn't enough time for that for most people's hardware.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "249128": "Hello everybody,\n\nI was wondering what kind of splits have been used in this competition. \nI see a lot of kernels with random 0.8/0.2 split but that just seems way wrong. May anyone share his/her approach?",
    "249232": "I'm using random split 0.95:0.05",
    "249287": "split products into 0.95:0.05.",
    "249362": "but stratified, right?",
    "249364": "Aren't you careful to make sure that the train/val data contains all classes?",
    "249724": "This dataset is huge, so you don't need such a high percentage of the data for validation. I use 0.1% of the products stratified by class as my validation set. \n\nFor computer vision, it is more important to have as much data go into training your model as you can, than it is to know to five significant digits how much you're overfitting by.\n\nBetter would be to train with k-fold cross validation, both for validation accuracy and for ensembling, but there isn't enough time for that for most people's hardware.",
    "250646": "no. i used only leaf category."
  },
  "source": "meta"
}