{
  "id": 39925,
  "title": "Validation vs test scores?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/39925",
  "author_name": "anokas",
  "post_date": "2017-09-24T07:15:29.961000",
  "votes": 8,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I'm just wondering what scores people are getting locally (on stuff they've submitted to the LB). Is the leaderboard score consistent with a random validation split for you?</p>\n\n<p>EDIT: Personally, my validation scores are basically identical to LB scores, with a random split validation set.</p>",
  "messages": [
    {
      "id": 223902,
      "postDate": "2017-09-24T07:15:29.960Z",
      "content": "<p>I'm just wondering what scores people are getting locally (on stuff they've submitted to the LB). Is the leaderboard score consistent with a random validation split for you?</p>\n\n<p>EDIT: Personally, my validation scores are basically identical to LB scores, with a random split validation set.</p>",
      "rawMarkdown": "I'm just wondering what scores people are getting locally (on stuff they've submitted to the LB). Is the leaderboard score consistent with a random validation split for you?\n\nEDIT: Personally, my validation scores are basically identical to LB scores, with a random split validation set.",
      "votes": 8
    },
    {
      "id": 228387,
      "postDate": "2017-10-06T15:23:32.687Z",
      "content": "<p>So, it seems that using 80/20 or 90/10 you can have the same score in validation set than in public leaderboard.\nEither using a random split (Anokas) or a stratified split (Human Analog)</p>\n\n<p>Since there are articles with more than one photo.  I was wondering if it was important to keep all the item images in the same fold.</p>\n\n<p>On one hand, distributing the images of the same item among the folds, validation score could mislead us into thinking that the model is better than it actually is.  I'm assuming that there are images too similar that could activate the high level features in the same way.</p>\n\n<p>On the other hand, if it doesn't matter, the more variety in train set, the better</p>\n\n<p>Did you distribute randomly (or stratifiedly) the images of the same item as different observations?\nOr did you take care distributing together all the images for a item?</p>",
      "rawMarkdown": "So, it seems that using 80/20 or 90/10 you can have the same score in validation set than in public leaderboard.\nEither using a random split (Anokas) or a stratified split (Human Analog)\n\n\nSince there are articles with more than one photo.  I was wondering if it was important to keep all the item images in the same fold.\n\nOn one hand, distributing the images of the same item among the folds, validation score could mislead us into thinking that the model is better than it actually is.  I'm assuming that there are images too similar that could activate the high level features in the same way.\n\nOn the other hand, if it doesn't matter, the more variety in train set, the better\n\nDid you distribute randomly (or stratifiedly) the images of the same item as different observations?\nOr did you take care distributing together all the images for a item?",
      "votes": 2
    },
    {
      "id": 224709,
      "postDate": "2017-09-27T11:56:09.473Z",
      "content": "<p>I just did a 80/20 train/validation split on the entire dataset (making sure to do this split on each individual category, since some categories contain only a few products and I wanted to guarantee these got split properly too). Using this split, I fine-tuned a pre-trained conv net for 3 epochs. (I didn't do anything special yet, no data preprocessing or augmentation etc.) My local validation score was 35% correct, and my LB score is the same. So consistent results so far. ;-)</p>",
      "rawMarkdown": "I just did a 80/20 train/validation split on the entire dataset (making sure to do this split on each individual category, since some categories contain only a few products and I wanted to guarantee these got split properly too). Using this split, I fine-tuned a pre-trained conv net for 3 epochs. (I didn't do anything special yet, no data preprocessing or augmentation etc.) My local validation score was 35% correct, and my LB score is the same. So consistent results so far. ;-)",
      "votes": 2
    },
    {
      "id": 225423,
      "postDate": "2017-09-29T02:19:58.297Z",
      "content": "<p>Juat wondering what is your loss function?</p>",
      "rawMarkdown": "Juat wondering what is your loss function?"
    },
    {
      "id": 225201,
      "postDate": "2017-09-28T13:38:27.170Z",
      "content": "<p>I consistently get 0.5-1% gain compared to my valid set. Keeps me wondering if there is some kind of bias</p>",
      "rawMarkdown": "I consistently get 0.5-1% gain compared to my valid set. Keeps me wondering if there is some kind of bias"
    },
    {
      "id": 225147,
      "postDate": "2017-09-28T11:23:39.327Z",
      "content": "<p>Hi Anokas,\ndid you do a stratified split? I think this may help with the class imbalance!</p>\n\n<p>ps: I am looking for someone to team up from the beginning to win this competition :) </p>",
      "rawMarkdown": "Hi Anokas,\ndid you do a stratified split? I think this may help with the class imbalance!\n\nps: I am looking for someone to team up from the beginning to win this competition :) ",
      "replies": [
        {
          "id": 226022,
          "postDate": "2017-09-30T18:21:27.767Z",
          "content": "<p>Hi, I would like to join the team. I am familiar with Keras but have the limit experience with computer vision.</p>",
          "rawMarkdown": "Hi, I would like to join the team. I am familiar with Keras but have the limit experience with computer vision."
        }
      ]
    },
    {
      "id": 224799,
      "postDate": "2017-09-27T17:30:24.613Z",
      "content": "<p>@Anokas: Can you tell us which hardware you are using for this competition? And how much time do you need to train your models on your hardware?</p>",
      "rawMarkdown": "@Anokas: Can you tell us which hardware you are using for this competition? And how much time do you need to train your models on your hardware?",
      "replies": [
        {
          "id": 224838,
          "postDate": "2017-09-27T19:00:13.687Z",
          "content": "<p>I'm using a 1080 Ti. My models so far have taken around 24 hours to train.</p>",
          "rawMarkdown": "I'm using a 1080 Ti. My models so far have taken around 24 hours to train.",
          "votes": 8
        },
        {
          "id": 224907,
          "postDate": "2017-09-27T22:35:40.343Z",
          "content": "<p>Interesting! Are you training on downsampled images?</p>",
          "rawMarkdown": "Interesting! Are you training on downsampled images?"
        }
      ]
    },
    {
      "id": 224708,
      "postDate": "2017-09-27T11:36:34.113Z",
      "content": "<p>Hi @anokas,\nI have just started working on this competition and haven't had any results to submit yet but in early experimentation and not having trained on all data yet I am getting accuracies of around 0.50 for both train and validation (obviously I don't know what these would translate to in LB).</p>\n\n<p>I am currently using a simple CNN I put together with images scaled to 32x32 (I only have a GTX 960 with 2GB memory).\nCheers,\n@bk0000.</p>",
      "rawMarkdown": "Hi @anokas,\nI have just started working on this competition and haven't had any results to submit yet but in early experimentation and not having trained on all data yet I am getting accuracies of around 0.50 for both train and validation (obviously I don't know what these would translate to in LB).\n\nI am currently using a simple CNN I put together with images scaled to 32x32 (I only have a GTX 960 with 2GB memory).\nCheers,\n@bk0000.\n"
    },
    {
      "id": 224757,
      "postDate": "2017-09-27T15:34:56.267Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 224760,
          "postDate": "2017-09-27T15:40:15.607Z",
          "content": "<p>Hi, can you share your model parameters number and training time? I trained my model with 90% images on a Nvidia Tesla K20Xm and the time is longer than 200 hours.</p>",
          "rawMarkdown": "Hi, can you share your model parameters number and training time? I trained my model with 90% images on a Nvidia Tesla K20Xm and the time is longer than 200 hours."
        },
        {
          "id": 225106,
          "postDate": "2017-09-28T09:13:45.033Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 228387,
      "author_name": "Virilo Tejedor Aguilera",
      "author_url": "",
      "post_date": "2017-10-06T15:23:32.687000",
      "content": "<p>So, it seems that using 80/20 or 90/10 you can have the same score in validation set than in public leaderboard.\nEither using a random split (Anokas) or a stratified split (Human Analog)</p>\n\n<p>Since there are articles with more than one photo.  I was wondering if it was important to keep all the item images in the same fold.</p>\n\n<p>On one hand, distributing the images of the same item among the folds, validation score could mislead us into thinking that the model is better than it actually is.  I'm assuming that there are images too similar that could activate the high level features in the same way.</p>\n\n<p>On the other hand, if it doesn't matter, the more variety in train set, the better</p>\n\n<p>Did you distribute randomly (or stratifiedly) the images of the same item as different observations?\nOr did you take care distributing together all the images for a item?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 224709,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2017-09-27T11:56:09.473000",
      "content": "<p>I just did a 80/20 train/validation split on the entire dataset (making sure to do this split on each individual category, since some categories contain only a few products and I wanted to guarantee these got split properly too). Using this split, I fine-tuned a pre-trained conv net for 3 epochs. (I didn't do anything special yet, no data preprocessing or augmentation etc.) My local validation score was 35% correct, and my LB score is the same. So consistent results so far. ;-)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 225423,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2017-09-29T02:19:58.297000",
      "content": "<p>Juat wondering what is your loss function?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 225201,
      "author_name": "Adam Blazek",
      "author_url": "",
      "post_date": "2017-09-28T13:38:27.170000",
      "content": "<p>I consistently get 0.5-1% gain compared to my valid set. Keeps me wondering if there is some kind of bias</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 225147,
      "author_name": "Tim Joseph",
      "author_url": "",
      "post_date": "2017-09-28T11:23:39.327000",
      "content": "<p>Hi Anokas,\ndid you do a stratified split? I think this may help with the class imbalance!</p>\n\n<p>ps: I am looking for someone to team up from the beginning to win this competition :) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 226022,
          "author_name": "Zhijian Li",
          "author_url": "",
          "post_date": "2017-09-30T18:21:27.767000",
          "content": "<p>Hi, I would like to join the team. I am familiar with Keras but have the limit experience with computer vision.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 224799,
      "author_name": "Thomas SELECK",
      "author_url": "",
      "post_date": "2017-09-27T17:30:24.613000",
      "content": "<p>@Anokas: Can you tell us which hardware you are using for this competition? And how much time do you need to train your models on your hardware?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 224838,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2017-09-27T19:00:13.687000",
          "content": "<p>I'm using a 1080 Ti. My models so far have taken around 24 hours to train.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 224907,
          "author_name": "Tim Joseph",
          "author_url": "",
          "post_date": "2017-09-27T22:35:40.343000",
          "content": "<p>Interesting! Are you training on downsampled images?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 224708,
      "author_name": "Baris Kanber",
      "author_url": "",
      "post_date": "2017-09-27T11:36:34.113000",
      "content": "<p>Hi @anokas,\nI have just started working on this competition and haven't had any results to submit yet but in early experimentation and not having trained on all data yet I am getting accuracies of around 0.50 for both train and validation (obviously I don't know what these would translate to in LB).</p>\n\n<p>I am currently using a simple CNN I put together with images scaled to 32x32 (I only have a GTX 960 with 2GB memory).\nCheers,\n@bk0000.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 224757,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-09-27T15:34:56.267000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 224760,
          "author_name": "Zhijian Li",
          "author_url": "",
          "post_date": "2017-09-27T15:40:15.607000",
          "content": "<p>Hi, can you share your model parameters number and training time? I trained my model with 90% images on a Nvidia Tesla K20Xm and the time is longer than 200 hours.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 225106,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-09-28T09:13:45.033000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "223902": "I'm just wondering what scores people are getting locally (on stuff they've submitted to the LB). Is the leaderboard score consistent with a random validation split for you?\n\nEDIT: Personally, my validation scores are basically identical to LB scores, with a random split validation set.",
    "228387": "So, it seems that using 80/20 or 90/10 you can have the same score in validation set than in public leaderboard.\nEither using a random split (Anokas) or a stratified split (Human Analog)\n\n\nSince there are articles with more than one photo.  I was wondering if it was important to keep all the item images in the same fold.\n\nOn one hand, distributing the images of the same item among the folds, validation score could mislead us into thinking that the model is better than it actually is.  I'm assuming that there are images too similar that could activate the high level features in the same way.\n\nOn the other hand, if it doesn't matter, the more variety in train set, the better\n\nDid you distribute randomly (or stratifiedly) the images of the same item as different observations?\nOr did you take care distributing together all the images for a item?",
    "224709": "I just did a 80/20 train/validation split on the entire dataset (making sure to do this split on each individual category, since some categories contain only a few products and I wanted to guarantee these got split properly too). Using this split, I fine-tuned a pre-trained conv net for 3 epochs. (I didn't do anything special yet, no data preprocessing or augmentation etc.) My local validation score was 35% correct, and my LB score is the same. So consistent results so far. ;-)",
    "225423": "Juat wondering what is your loss function?",
    "225201": "I consistently get 0.5-1% gain compared to my valid set. Keeps me wondering if there is some kind of bias",
    "225147": "Hi Anokas,\ndid you do a stratified split? I think this may help with the class imbalance!\n\nps: I am looking for someone to team up from the beginning to win this competition :) ",
    "224799": "@Anokas: Can you tell us which hardware you are using for this competition? And how much time do you need to train your models on your hardware?",
    "224708": "Hi @anokas,\nI have just started working on this competition and haven't had any results to submit yet but in early experimentation and not having trained on all data yet I am getting accuracies of around 0.50 for both train and validation (obviously I don't know what these would translate to in LB).\n\nI am currently using a simple CNN I put together with images scaled to 32x32 (I only have a GTX 960 with 2GB memory).\nCheers,\n@bk0000.\n",
    "224757": ""
  }
}