{
  "id": 135973,
  "title": "What I learnt as a beginner",
  "url": "/competitions/bengaliai-cv19/discussion/135973",
  "author_name": "",
  "post_date": "2020-03-17T01:01:21.695540900Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This was my first serious competition on Kaggle and these are what I learnt as a beginner:</p>\n\n<ul>\n<li><p>Trust thy CV: This has cost me my first bronze medal. It was well known (declared in competition description) that there is going to be unseen graphemes and this competition as a whole is a test of how well your model performs with unknown graphemes. Due to tight resource constraints, I was not able to test my models for a long time with a proper CV setup(something like what Qishen Ha provided). Turns out my best single model gave 0.9324 and best ensemble was giving 0.9326 either of them could have given a bronze medal but due to relying on Public LB, I got mislead and chose two non-optimal models.</p></li>\n<li><p>Discussions: The discussions forum is a gem. Almost all the tips and tricks needed to get a medal was in the discussions. The likes of @haqishen, @hengck23 and many others gave invaluable information to everyone. No wonder they are Discussions Master and Grandmaster.</p></li>\n<li><p>Augmentations: Augmentations is the King in this competition. This competition was in a way, Do Not Overfit type of competition. The CNNs with their humongus capacity to learn are prone to overfitting. The solution ? Heavy Augmentations. After trying ShiftScaleRotate, ElasticTransform, Cutout, Gridmask, Cutmix, Mixup, Gridmask, Augmix, what finally worked for me is a combination of ElasticTransfrom, ShiftScaleRotate and Cutmix with OHEM Loss.</p></li>\n<li><p>Compute resources are important: In competitions like this, people with access to Turing and Volta GPUs had a lot of advantage. A V100 can run one epoch of SeResNext50 with Augmentations under 6 mins (with validations) and the same took 12mins in a P100 provided by Colab. This is mostly due to the fact that the newer GPUs support FP16 mixed training while the older ones do not. I was too eager in training the models right away without a proper cv setup as a result I spent 2/3rds of my free gcp credits before I was training my best models. Do not do that. Start on colab, train and test there and finally move to GCP for the final training. I believe people with their own Turing GPUs had some advantage over others.</p></li>\n<li><p>Do not copy paste others inference kernels: While it looks like an easy way to top, it is not. It is merely a good starting point but you will have to train your own model and test in your own way to get the best results. I had lots of problems while making my own inference kernel and had to copy someone's kernel to able to submit. At first I used their model as a result, I got far higher score in public LB than I deserved. That score gave me motivation of what to beat which I finally did, after 3 days and my score finally became my own. But seriously, copy pasting blindly is not a good idea as you learn nothing. 300+ people above me in the LB had the same score with &lt;5 submissions was kind of disheartening and I slowly learnt to ignore that.</p></li>\n</ul>\n\n<p>Overall this competition gave me invaluable learning experience. I started at the 2nd Last position in the public LB, climbed to late 600, fell back to 800+ as I couldn't improve my score but was confident that a huge shakeup is going to happen and it finally did.</p>",
  "messages": [
    {
      "id": "775789",
      "postDate": "03/17/2020 01:01:21",
      "content": "<p>This was my first serious competition on Kaggle and these are what I learnt as a beginner:</p>\n\n<ul>\n<li><p>Trust thy CV: This has cost me my first bronze medal. It was well known (declared in competition description) that there is going to be unseen graphemes and this competition as a whole is a test of how well your model performs with unknown graphemes. Due to tight resource constraints, I was not able to test my models for a long time with a proper CV setup(something like what Qishen Ha provided). Turns out my best single model gave 0.9324 and best ensemble was giving 0.9326 either of them could have given a bronze medal but due to relying on Public LB, I got mislead and chose two non-optimal models.</p></li>\n<li><p>Discussions: The discussions forum is a gem. Almost all the tips and tricks needed to get a medal was in the discussions. The likes of @haqishen, @hengck23 and many others gave invaluable information to everyone. No wonder they are Discussions Master and Grandmaster.</p></li>\n<li><p>Augmentations: Augmentations is the King in this competition. This competition was in a way, Do Not Overfit type of competition. The CNNs with their humongus capacity to learn are prone to overfitting. The solution ? Heavy Augmentations. After trying ShiftScaleRotate, ElasticTransform, Cutout, Gridmask, Cutmix, Mixup, Gridmask, Augmix, what finally worked for me is a combination of ElasticTransfrom, ShiftScaleRotate and Cutmix with OHEM Loss.</p></li>\n<li><p>Compute resources are important: In competitions like this, people with access to Turing and Volta GPUs had a lot of advantage. A V100 can run one epoch of SeResNext50 with Augmentations under 6 mins (with validations) and the same took 12mins in a P100 provided by Colab. This is mostly due to the fact that the newer GPUs support FP16 mixed training while the older ones do not. I was too eager in training the models right away without a proper cv setup as a result I spent 2/3rds of my free gcp credits before I was training my best models. Do not do that. Start on colab, train and test there and finally move to GCP for the final training. I believe people with their own Turing GPUs had some advantage over others.</p></li>\n<li><p>Do not copy paste others inference kernels: While it looks like an easy way to top, it is not. It is merely a good starting point but you will have to train your own model and test in your own way to get the best results. I had lots of problems while making my own inference kernel and had to copy someone's kernel to able to submit. At first I used their model as a result, I got far higher score in public LB than I deserved. That score gave me motivation of what to beat which I finally did, after 3 days and my score finally became my own. But seriously, copy pasting blindly is not a good idea as you learn nothing. 300+ people above me in the LB had the same score with &lt;5 submissions was kind of disheartening and I slowly learnt to ignore that.</p></li>\n</ul>\n\n<p>Overall this competition gave me invaluable learning experience. I started at the 2nd Last position in the public LB, climbed to late 600, fell back to 800+ as I couldn't improve my score but was confident that a huge shakeup is going to happen and it finally did.</p>",
      "rawMarkdown": "This was my first serious competition on Kaggle and these are what I learnt as a beginner:\n\n-  Trust thy CV: This has cost me my first bronze medal. It was well known (declared in competition description) that there is going to be unseen graphemes and this competition as a whole is a test of how well your model performs with unknown graphemes. Due to tight resource constraints, I was not able to test my models for a long time with a proper CV setup(something like what Qishen Ha provided). Turns out my best single model gave 0.9324 and best ensemble was giving 0.9326 either of them could have given a bronze medal but due to relying on Public LB, I got mislead and chose two non-optimal models.\n\n-  Discussions: The discussions forum is a gem. Almost all the tips and tricks needed to get a medal was in the discussions. The likes of @haqishen, @hengck23 and many others gave invaluable information to everyone. No wonder they are Discussions Master and Grandmaster.\n\n-  Augmentations: Augmentations is the King in this competition. This competition was in a way, Do Not Overfit type of competition. The CNNs with their humongus capacity to learn are prone to overfitting. The solution ? Heavy Augmentations. After trying ShiftScaleRotate, ElasticTransform, Cutout, Gridmask, Cutmix, Mixup, Gridmask, Augmix, what finally worked for me is a combination of ElasticTransfrom, ShiftScaleRotate and Cutmix with OHEM Loss.\n\n-  Compute resources are important: In competitions like this, people with access to Turing and Volta GPUs had a lot of advantage. A V100 can run one epoch of SeResNext50 with Augmentations under 6 mins (with validations) and the same took 12mins in a P100 provided by Colab. This is mostly due to the fact that the newer GPUs support FP16 mixed training while the older ones do not. I was too eager in training the models right away without a proper cv setup as a result I spent 2/3rds of my free gcp credits before I was training my best models. Do not do that. Start on colab, train and test there and finally move to GCP for the final training. I believe people with their own Turing GPUs had some advantage over others.\n\n-  Do not copy paste others inference kernels: While it looks like an easy way to top, it is not. It is merely a good starting point but you will have to train your own model and test in your own way to get the best results. I had lots of problems while making my own inference kernel and had to copy someone's kernel to able to submit. At first I used their model as a result, I got far higher score in public LB than I deserved. That score gave me motivation of what to beat which I finally did, after 3 days and my score finally became my own. But seriously, copy pasting blindly is not a good idea as you learn nothing. 300+ people above me in the LB had the same score with &lt;5 submissions was kind of disheartening and I slowly learnt to ignore that.\n\nOverall this competition gave me invaluable learning experience. I started at the 2nd Last position in the public LB, climbed to late 600, fell back to 800+ as I couldn't improve my score but was confident that a huge shakeup is going to happen and it finally did.",
      "votes": null
    },
    {
      "id": "775794",
      "postDate": "03/17/2020 01:03:19",
      "content": "<p>I really don't think local CV correlates with private LB in this competition.</p>",
      "rawMarkdown": "I really don't think local CV correlates with private LB in this competition.",
      "votes": null
    },
    {
      "id": "775802",
      "postDate": "03/17/2020 01:11:19",
      "content": "<p>The one made by Qishen Ha did correlate somewhat, more than the normal CV. It had unknown graphemes in the validation part. My ensembles were not tested with any of the CV setups. The once that were tested, didn't have unknown graphemes in the validation part.</p>",
      "rawMarkdown": "The one made by Qishen Ha did correlate somewhat, more than the normal CV. It had unknown graphemes in the validation part. My ensembles were not tested with any of the CV setups. The once that were tested, didn't have unknown graphemes in the validation part.",
      "votes": null
    },
    {
      "id": "776117",
      "postDate": "03/17/2020 06:08:39",
      "content": "<p>CV matters but not much in this competition. BTW, you can also train with P100 in FP16 mode.</p>",
      "rawMarkdown": "CV matters but not much in this competition. BTW, you can also train with P100 in FP16 mode.",
      "votes": null
    },
    {
      "id": "776119",
      "postDate": "03/17/2020 06:11:07",
      "content": "<p>You can, but the benefit is negligible. Correct me if I am wrong, the P100s do not have any tensor cores and the only benefit comes from reduced size of the computations. The real benefits can be seen on Volta and Turing architectures.</p>",
      "rawMarkdown": "You can, but the benefit is negligible. Correct me if I am wrong, the P100s do not have any tensor cores and the only benefit comes from reduced size of the computations. The real benefits can be seen on Volta and Turing architectures.",
      "votes": null
    },
    {
      "id": "776134",
      "postDate": "03/17/2020 06:22:31",
      "content": "<p>You're right. P100 with FP16 doesn't help accelerate training but it's beneficial when you train with larger image size.</p>",
      "rawMarkdown": "You're right. P100 with FP16 doesn't help accelerate training but it's beneficial when you train with larger image size.",
      "votes": null
    },
    {
      "id": "776202",
      "postDate": "03/17/2020 07:39:01",
      "content": "<p>You learned a lot of things as a beginner. Why were you sure of a shakeup?</p>\n\n<p>Was it due to the fact that EfficientNet Models tend to overfit without augmentation and after using the power of TPUs</p>",
      "rawMarkdown": "You learned a lot of things as a beginner. Why were you sure of a shakeup?\n\nWas it due to the fact that EfficientNet Models tend to overfit without augmentation and after using the power of TPUs",
      "votes": null
    },
    {
      "id": "776294",
      "postDate": "03/17/2020 09:17:16",
      "content": "<ul>\n<li>I was confident of the shakeup due to the fact that the private dataset was going to have unseen graphemes (This was written in the competition description as the primary objective of the competition) and also a lot of people were using single model for inference (educated guess). Ensembles kind of protected against overfitting.</li>\n<li>My efficientnet b4 (single model) actually performed much better on the private lb compared to other models and I used one in an ensemble with seresnext50 ohem loss. The seresnexts gave marginally better public lb scores.</li>\n<li>I believe all cnns are prone to overfitting and need to have a lot of augmentations to prevent it.</li>\n</ul>",
      "rawMarkdown": "I was confident of the shakeup due to the fact that the private dataset was going to have unseen graphemes (This was written in the competition description as the primary objective of the competition) and also a lot of people were using single model for inference (educated guess). Ensembles kind of protected against overfitting.\n- My efficientnet b4 (single model) actually performed much better on the private lb compared to other models and I used one in an ensemble with seresnext50 ohem loss. The seresnexts gave marginally better public lb scores.\n- I believe all cnns are prone to overfitting and need to have a lot of augmentations to prevent it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 775794,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "03/17/2020 01:03:19",
      "content": "<p>I really don't think local CV correlates with private LB in this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 775802,
          "author_name": "utsavnandi",
          "author_url": "",
          "post_date": "03/17/2020 01:11:19",
          "content": "<p>The one made by Qishen Ha did correlate somewhat, more than the normal CV. It had unknown graphemes in the validation part. My ensembles were not tested with any of the CV setups. The once that were tested, didn't have unknown graphemes in the validation part.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776117,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "03/17/2020 06:08:39",
      "content": "<p>CV matters but not much in this competition. BTW, you can also train with P100 in FP16 mode.</p>",
      "votes": null,
      "replies": [
        {
          "id": 776119,
          "author_name": "utsavnandi",
          "author_url": "",
          "post_date": "03/17/2020 06:11:07",
          "content": "<p>You can, but the benefit is negligible. Correct me if I am wrong, the P100s do not have any tensor cores and the only benefit comes from reduced size of the computations. The real benefits can be seen on Volta and Turing architectures.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 776134,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "03/17/2020 06:22:31",
          "content": "<p>You're right. P100 with FP16 doesn't help accelerate training but it's beneficial when you train with larger image size.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776202,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "03/17/2020 07:39:01",
      "content": "<p>You learned a lot of things as a beginner. Why were you sure of a shakeup?</p>\n\n<p>Was it due to the fact that EfficientNet Models tend to overfit without augmentation and after using the power of TPUs</p>",
      "votes": null,
      "replies": [
        {
          "id": 776294,
          "author_name": "utsavnandi",
          "author_url": "",
          "post_date": "03/17/2020 09:17:16",
          "content": "<ul>\n<li>I was confident of the shakeup due to the fact that the private dataset was going to have unseen graphemes (This was written in the competition description as the primary objective of the competition) and also a lot of people were using single model for inference (educated guess). Ensembles kind of protected against overfitting.</li>\n<li>My efficientnet b4 (single model) actually performed much better on the private lb compared to other models and I used one in an ensemble with seresnext50 ohem loss. The seresnexts gave marginally better public lb scores.</li>\n<li>I believe all cnns are prone to overfitting and need to have a lot of augmentations to prevent it.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "775789": "This was my first serious competition on Kaggle and these are what I learnt as a beginner:\n\n-  Trust thy CV: This has cost me my first bronze medal. It was well known (declared in competition description) that there is going to be unseen graphemes and this competition as a whole is a test of how well your model performs with unknown graphemes. Due to tight resource constraints, I was not able to test my models for a long time with a proper CV setup(something like what Qishen Ha provided). Turns out my best single model gave 0.9324 and best ensemble was giving 0.9326 either of them could have given a bronze medal but due to relying on Public LB, I got mislead and chose two non-optimal models.\n\n-  Discussions: The discussions forum is a gem. Almost all the tips and tricks needed to get a medal was in the discussions. The likes of @haqishen, @hengck23 and many others gave invaluable information to everyone. No wonder they are Discussions Master and Grandmaster.\n\n-  Augmentations: Augmentations is the King in this competition. This competition was in a way, Do Not Overfit type of competition. The CNNs with their humongus capacity to learn are prone to overfitting. The solution ? Heavy Augmentations. After trying ShiftScaleRotate, ElasticTransform, Cutout, Gridmask, Cutmix, Mixup, Gridmask, Augmix, what finally worked for me is a combination of ElasticTransfrom, ShiftScaleRotate and Cutmix with OHEM Loss.\n\n-  Compute resources are important: In competitions like this, people with access to Turing and Volta GPUs had a lot of advantage. A V100 can run one epoch of SeResNext50 with Augmentations under 6 mins (with validations) and the same took 12mins in a P100 provided by Colab. This is mostly due to the fact that the newer GPUs support FP16 mixed training while the older ones do not. I was too eager in training the models right away without a proper cv setup as a result I spent 2/3rds of my free gcp credits before I was training my best models. Do not do that. Start on colab, train and test there and finally move to GCP for the final training. I believe people with their own Turing GPUs had some advantage over others.\n\n-  Do not copy paste others inference kernels: While it looks like an easy way to top, it is not. It is merely a good starting point but you will have to train your own model and test in your own way to get the best results. I had lots of problems while making my own inference kernel and had to copy someone's kernel to able to submit. At first I used their model as a result, I got far higher score in public LB than I deserved. That score gave me motivation of what to beat which I finally did, after 3 days and my score finally became my own. But seriously, copy pasting blindly is not a good idea as you learn nothing. 300+ people above me in the LB had the same score with &lt;5 submissions was kind of disheartening and I slowly learnt to ignore that.\n\nOverall this competition gave me invaluable learning experience. I started at the 2nd Last position in the public LB, climbed to late 600, fell back to 800+ as I couldn't improve my score but was confident that a huge shakeup is going to happen and it finally did.",
    "775794": "I really don't think local CV correlates with private LB in this competition.",
    "775802": "The one made by Qishen Ha did correlate somewhat, more than the normal CV. It had unknown graphemes in the validation part. My ensembles were not tested with any of the CV setups. The once that were tested, didn't have unknown graphemes in the validation part.",
    "776117": "CV matters but not much in this competition. BTW, you can also train with P100 in FP16 mode.",
    "776119": "You can, but the benefit is negligible. Correct me if I am wrong, the P100s do not have any tensor cores and the only benefit comes from reduced size of the computations. The real benefits can be seen on Volta and Turing architectures.",
    "776134": "You're right. P100 with FP16 doesn't help accelerate training but it's beneficial when you train with larger image size.",
    "776202": "You learned a lot of things as a beginner. Why were you sure of a shakeup?\n\nWas it due to the fact that EfficientNet Models tend to overfit without augmentation and after using the power of TPUs",
    "776294": "I was confident of the shakeup due to the fact that the private dataset was going to have unseen graphemes (This was written in the competition description as the primary objective of the competition) and also a lot of people were using single model for inference (educated guess). Ensembles kind of protected against overfitting.\n- My efficientnet b4 (single model) actually performed much better on the private lb compared to other models and I used one in an ensemble with seresnext50 ohem loss. The seresnexts gave marginally better public lb scores.\n- I believe all cnns are prone to overfitting and need to have a lot of augmentations to prevent it."
  },
  "source": "meta"
}