{
  "id": 317630,
  "title": "How do you approach the gap in scoring performance?",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/317630",
  "author_name": "",
  "post_date": "2022-04-08T05:21:09.665116500Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm pretty new to all this, and assuming this is entirely normal, but I'm noticing a huge gap between my model's performance on a validation slice, and performance on the test data once submitted. This isn't just a little jump, suggesting something is quite different.</p>\n<p>My development path has been to start from Michal's baseline notebooks, work out how to migrate that to ArcFace and tune its parameters, and got some small improvements in performance out of that. However, in the same way their baseline notebook goes from 0.609 on validation set to 0.175 in scoring, my model goes from 0.628 on validation to 0.215 in scoring.</p>\n<p>So far I've tried weighting classes to deal with certain hotels being over-represented, and tweaking the augmentation chain to include more severe occlusions (in case the test set had worse/larger masks than I was training on). Both of those changes damaged performance, the latter quite badly.</p>\n<p>Anyway, I'm nearly out of GPU credits, so time for a break from experiments, but keen to hear what others are trying!</p>",
  "messages": [
    {
      "id": "1748924",
      "postDate": "04/08/2022 05:21:09",
      "content": "<p>I'm pretty new to all this, and assuming this is entirely normal, but I'm noticing a huge gap between my model's performance on a validation slice, and performance on the test data once submitted. This isn't just a little jump, suggesting something is quite different.</p>\n<p>My development path has been to start from Michal's baseline notebooks, work out how to migrate that to ArcFace and tune its parameters, and got some small improvements in performance out of that. However, in the same way their baseline notebook goes from 0.609 on validation set to 0.175 in scoring, my model goes from 0.628 on validation to 0.215 in scoring.</p>\n<p>So far I've tried weighting classes to deal with certain hotels being over-represented, and tweaking the augmentation chain to include more severe occlusions (in case the test set had worse/larger masks than I was training on). Both of those changes damaged performance, the latter quite badly.</p>\n<p>Anyway, I'm nearly out of GPU credits, so time for a break from experiments, but keen to hear what others are trying!</p>",
      "rawMarkdown": "I'm pretty new to all this, and assuming this is entirely normal, but I'm noticing a huge gap between my model's performance on a validation slice, and performance on the test data once submitted. This isn't just a little jump, suggesting something is quite different.\n\nMy development path has been to start from Michal's baseline notebooks, work out how to migrate that to ArcFace and tune its parameters, and got some small improvements in performance out of that. However, in the same way their baseline notebook goes from 0.609 on validation set to 0.175 in scoring, my model goes from 0.628 on validation to 0.215 in scoring.\n\nSo far I've tried weighting classes to deal with certain hotels being over-represented, and tweaking the augmentation chain to include more severe occlusions (in case the test set had worse/larger masks than I was training on). Both of those changes damaged performance, the latter quite badly.\n\nAnyway, I'm nearly out of GPU credits, so time for a break from experiments, but keen to hear what others are trying!",
      "votes": null
    },
    {
      "id": "1749269",
      "postDate": "04/08/2022 12:12:26",
      "content": "<p>Well the huge gap is not normal and you should aim to avoid it. Having good validation strategy is important so you can test your ideas but as long as the CV and LB score is correlated (CV goes up then LB goes up too) it's usable.</p>\n<p>We know what the occlusions in the test set look like as they are in the train_masks folder (it's called train_masks but they are masks for the test set, explained <a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/313547#1734687\" target=\"_blank\">here</a> by the competition host).<br>\nThere are 5000 images and 4950 mask files (so basically every image has occlusions), median occlusion area is around 25% of the image and they are mainly located in the middle (you can check my <a href=\"https://www.kaggle.com/code/michaln/masks-and-occlusions\" target=\"_blank\">Masks and occlusions notebook</a>)</p>\n<p>In the starter notebook I use augmentation to generate the occlusions for validation set. They have random location and probability of 0.75. So you should definitely generate it for every validation image (p=1) and maybe you can try masking the center of the image instead of random location (center might have more important information).</p>\n<p>Another thing is the way we split the data. I just sample 1 image per hotel and use it for validation. There are 11 hotels that have only 1 image though, so they are only in validation set and not used for training which can hurt the score a little.</p>\n<p>The images in the test set can also be very different than the ones in the training data. In <a href=\"https://youtu.be/OrGsQ_pHg2k?t=196\" target=\"_blank\">this video</a> describing dataset of the last year competition they mention that test images were selected ensuring that it was taken by different user than the ones in the training set. I guess it will be similar this year. We don't have information about users but maybe you can use image size or color histogram to split the dataset in better way or try some image normalization to mitigate the problem.</p>\n<p>Anyway my current best solution has accuracy 0.6470 and top 5 accuracy 0.7779 while LB score only 0.410. So I don't have solution either yet.</p>",
      "rawMarkdown": "Well the huge gap is not normal and you should aim to avoid it. Having good validation strategy is important so you can test your ideas but as long as the CV and LB score is correlated (CV goes up then LB goes up too) it's usable.\n\nWe know what the occlusions in the test set look like as they are in the train_masks folder (it's called train_masks but they are masks for the test set, explained [here](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/313547#1734687) by the competition host).\nThere are 5000 images and 4950 mask files (so basically every image has occlusions), median occlusion area is around 25% of the image and they are mainly located in the middle (you can check my [Masks and occlusions notebook](https://www.kaggle.com/code/michaln/masks-and-occlusions))\n\nIn the starter notebook I use augmentation to generate the occlusions for validation set. They have random location and probability of 0.75. So you should definitely generate it for every validation image (p=1) and maybe you can try masking the center of the image instead of random location (center might have more important information).\n\nAnother thing is the way we split the data. I just sample 1 image per hotel and use it for validation. There are 11 hotels that have only 1 image though, so they are only in validation set and not used for training which can hurt the score a little.\n\nThe images in the test set can also be very different than the ones in the training data. In [this video](https://youtu.be/OrGsQ_pHg2k?t=196) describing dataset of the last year competition they mention that test images were selected ensuring that it was taken by different user than the ones in the training set. I guess it will be similar this year. We don't have information about users but maybe you can use image size or color histogram to split the dataset in better way or try some image normalization to mitigate the problem.\n\nAnyway my current best solution has accuracy 0.6470 and top 5 accuracy 0.7779 while LB score only 0.410. So I don't have solution either yet.",
      "votes": null
    },
    {
      "id": "1752099",
      "postDate": "04/11/2022 11:58:30",
      "content": "<p>Thank you, some very useful lines to consider there. It's odd having so little to go on, but probably reflects reality better!</p>",
      "rawMarkdown": "Thank you, some very useful lines to consider there. It's odd having so little to go on, but probably reflects reality better!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1749269,
      "author_name": "michaln",
      "author_url": "",
      "post_date": "04/08/2022 12:12:26",
      "content": "<p>Well the huge gap is not normal and you should aim to avoid it. Having good validation strategy is important so you can test your ideas but as long as the CV and LB score is correlated (CV goes up then LB goes up too) it's usable.</p>\n<p>We know what the occlusions in the test set look like as they are in the train_masks folder (it's called train_masks but they are masks for the test set, explained <a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/313547#1734687\" target=\"_blank\">here</a> by the competition host).<br>\nThere are 5000 images and 4950 mask files (so basically every image has occlusions), median occlusion area is around 25% of the image and they are mainly located in the middle (you can check my <a href=\"https://www.kaggle.com/code/michaln/masks-and-occlusions\" target=\"_blank\">Masks and occlusions notebook</a>)</p>\n<p>In the starter notebook I use augmentation to generate the occlusions for validation set. They have random location and probability of 0.75. So you should definitely generate it for every validation image (p=1) and maybe you can try masking the center of the image instead of random location (center might have more important information).</p>\n<p>Another thing is the way we split the data. I just sample 1 image per hotel and use it for validation. There are 11 hotels that have only 1 image though, so they are only in validation set and not used for training which can hurt the score a little.</p>\n<p>The images in the test set can also be very different than the ones in the training data. In <a href=\"https://youtu.be/OrGsQ_pHg2k?t=196\" target=\"_blank\">this video</a> describing dataset of the last year competition they mention that test images were selected ensuring that it was taken by different user than the ones in the training set. I guess it will be similar this year. We don't have information about users but maybe you can use image size or color histogram to split the dataset in better way or try some image normalization to mitigate the problem.</p>\n<p>Anyway my current best solution has accuracy 0.6470 and top 5 accuracy 0.7779 while LB score only 0.410. So I don't have solution either yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1752099,
      "author_name": "prubyg",
      "author_url": "",
      "post_date": "04/11/2022 11:58:30",
      "content": "<p>Thank you, some very useful lines to consider there. It's odd having so little to go on, but probably reflects reality better!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1748924": "I'm pretty new to all this, and assuming this is entirely normal, but I'm noticing a huge gap between my model's performance on a validation slice, and performance on the test data once submitted. This isn't just a little jump, suggesting something is quite different.\n\nMy development path has been to start from Michal's baseline notebooks, work out how to migrate that to ArcFace and tune its parameters, and got some small improvements in performance out of that. However, in the same way their baseline notebook goes from 0.609 on validation set to 0.175 in scoring, my model goes from 0.628 on validation to 0.215 in scoring.\n\nSo far I've tried weighting classes to deal with certain hotels being over-represented, and tweaking the augmentation chain to include more severe occlusions (in case the test set had worse/larger masks than I was training on). Both of those changes damaged performance, the latter quite badly.\n\nAnyway, I'm nearly out of GPU credits, so time for a break from experiments, but keen to hear what others are trying!",
    "1749269": "Well the huge gap is not normal and you should aim to avoid it. Having good validation strategy is important so you can test your ideas but as long as the CV and LB score is correlated (CV goes up then LB goes up too) it's usable.\n\nWe know what the occlusions in the test set look like as they are in the train_masks folder (it's called train_masks but they are masks for the test set, explained [here](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/313547#1734687) by the competition host).\nThere are 5000 images and 4950 mask files (so basically every image has occlusions), median occlusion area is around 25% of the image and they are mainly located in the middle (you can check my [Masks and occlusions notebook](https://www.kaggle.com/code/michaln/masks-and-occlusions))\n\nIn the starter notebook I use augmentation to generate the occlusions for validation set. They have random location and probability of 0.75. So you should definitely generate it for every validation image (p=1) and maybe you can try masking the center of the image instead of random location (center might have more important information).\n\nAnother thing is the way we split the data. I just sample 1 image per hotel and use it for validation. There are 11 hotels that have only 1 image though, so they are only in validation set and not used for training which can hurt the score a little.\n\nThe images in the test set can also be very different than the ones in the training data. In [this video](https://youtu.be/OrGsQ_pHg2k?t=196) describing dataset of the last year competition they mention that test images were selected ensuring that it was taken by different user than the ones in the training set. I guess it will be similar this year. We don't have information about users but maybe you can use image size or color histogram to split the dataset in better way or try some image normalization to mitigate the problem.\n\nAnyway my current best solution has accuracy 0.6470 and top 5 accuracy 0.7779 while LB score only 0.410. So I don't have solution either yet.",
    "1752099": "Thank you, some very useful lines to consider there. It's odd having so little to go on, but probably reflects reality better!"
  },
  "source": "meta"
}