{
  "id": 133296,
  "title": "Oversample the training dataset",
  "url": "/competitions/flower-classification-with-tpus/discussion/133296",
  "author_name": "",
  "post_date": "2020-03-02T00:36:20.575503800Z",
  "votes": 12,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I published a kernel <a href=\"https://www.kaggle.com/yihdarshieh/tutorial-oversample?scriptVersionId=29496139\">Tutorial -- Oversample</a> to demonstrate <code>oversampling</code> the training dataset.</p>\n\n<p>Reference for oversampling: <a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data\">Classification on imbalanced data</a>.</p>\n\n<p>Why oversampling instead of using sample weights or class weights: (from the above link)</p>\n\n<p>&gt; If the training process were considering the whole dataset on each gradient update, this oversampling would be basically identical to the class weighting.</p>\n\n<p>&gt; But when training the model batch-wise, as you did here, the oversampled data provides a smoother gradient signal: Instead of each positive example being shown in one batch with a large weight, they're shown in many different batches each time with a small weight.</p>\n\n<p>&gt; This smoother gradient signal makes it easier to train the model.</p>\n\n<p>A sample result from my own training (1 means no oversampling, 100 means each class occurs at least approximately 100 times):</p>\n\n<pre><code>{\n    \"EfficientNetB7\": {\n        \"oversampling\": {\n            \"1\": {\n                \"f1\": 0.9163497711078313,\n                \"recall\": 0.9205266060440362,\n                \"precision\": 0.919277486006819,\n                \"acc\": 0.9240301724137931\n            },\n            \"100\": {\n                \"f1\": 0.9306967499560146,\n                \"recall\": 0.9385304126936429,\n                \"precision\": 0.9277303112649578,\n                \"acc\": 0.9296875\n            },\n            \"300\": {\n                \"f1\": 0.9361016120193594,\n                \"recall\": 0.9394270473352216,\n                \"precision\": 0.939168918056952,\n                \"acc\": 0.9366918103448276\n            },\n            \"800\": {\n                \"f1\": 0.9404645585115528,\n                \"recall\": 0.9427588505810944,\n                \"precision\": 0.9440889947725051,\n                \"acc\": 0.9391163793103449\n            }\n        }\n    }\n}\n</code></pre>\n\n<p>You can use oversampling with more data augmentation to get better results.</p>",
  "messages": [
    {
      "id": "760951",
      "postDate": "03/02/2020 00:36:20",
      "content": "<p>I published a kernel <a href=\"https://www.kaggle.com/yihdarshieh/tutorial-oversample?scriptVersionId=29496139\">Tutorial -- Oversample</a> to demonstrate <code>oversampling</code> the training dataset.</p>\n\n<p>Reference for oversampling: <a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data\">Classification on imbalanced data</a>.</p>\n\n<p>Why oversampling instead of using sample weights or class weights: (from the above link)</p>\n\n<p>&gt; If the training process were considering the whole dataset on each gradient update, this oversampling would be basically identical to the class weighting.</p>\n\n<p>&gt; But when training the model batch-wise, as you did here, the oversampled data provides a smoother gradient signal: Instead of each positive example being shown in one batch with a large weight, they're shown in many different batches each time with a small weight.</p>\n\n<p>&gt; This smoother gradient signal makes it easier to train the model.</p>\n\n<p>A sample result from my own training (1 means no oversampling, 100 means each class occurs at least approximately 100 times):</p>\n\n<pre><code>{\n    \"EfficientNetB7\": {\n        \"oversampling\": {\n            \"1\": {\n                \"f1\": 0.9163497711078313,\n                \"recall\": 0.9205266060440362,\n                \"precision\": 0.919277486006819,\n                \"acc\": 0.9240301724137931\n            },\n            \"100\": {\n                \"f1\": 0.9306967499560146,\n                \"recall\": 0.9385304126936429,\n                \"precision\": 0.9277303112649578,\n                \"acc\": 0.9296875\n            },\n            \"300\": {\n                \"f1\": 0.9361016120193594,\n                \"recall\": 0.9394270473352216,\n                \"precision\": 0.939168918056952,\n                \"acc\": 0.9366918103448276\n            },\n            \"800\": {\n                \"f1\": 0.9404645585115528,\n                \"recall\": 0.9427588505810944,\n                \"precision\": 0.9440889947725051,\n                \"acc\": 0.9391163793103449\n            }\n        }\n    }\n}\n</code></pre>\n\n<p>You can use oversampling with more data augmentation to get better results.</p>",
      "rawMarkdown": "I published a kernel [Tutorial -- Oversample](https://www.kaggle.com/yihdarshieh/tutorial-oversample?scriptVersionId=29496139) to demonstrate `oversampling` the training dataset.\n\nReference for oversampling: [Classification on imbalanced data](https://www.tensorflow.org/tutorials/structured_data/imbalanced_data).\n\nWhy oversampling instead of using sample weights or class weights: (from the above link)\n\n   &gt; If the training process were considering the whole dataset on each gradient update, this oversampling would be basically identical to the class weighting.\n\n   &gt; But when training the model batch-wise, as you did here, the oversampled data provides a smoother gradient signal: Instead of each positive example being shown in one batch with a large weight, they're shown in many different batches each time with a small weight.\n\n   &gt; This smoother gradient signal makes it easier to train the model.\n\nA sample result from my own training (1 means no oversampling, 100 means each class occurs at least approximately 100 times):\n\n\n\n    {\n        \"EfficientNetB7\": {\n            \"oversampling\": {\n                \"1\": {\n                    \"f1\": 0.9163497711078313,\n                    \"recall\": 0.9205266060440362,\n                    \"precision\": 0.919277486006819,\n                    \"acc\": 0.9240301724137931\n                },\n                \"100\": {\n                    \"f1\": 0.9306967499560146,\n                    \"recall\": 0.9385304126936429,\n                    \"precision\": 0.9277303112649578,\n                    \"acc\": 0.9296875\n                },\n                \"300\": {\n                    \"f1\": 0.9361016120193594,\n                    \"recall\": 0.9394270473352216,\n                    \"precision\": 0.939168918056952,\n                    \"acc\": 0.9366918103448276\n                },\n                \"800\": {\n                    \"f1\": 0.9404645585115528,\n                    \"recall\": 0.9427588505810944,\n                    \"precision\": 0.9440889947725051,\n                    \"acc\": 0.9391163793103449\n                }\n            }\n        }\n    }\n\nYou can use oversampling with more data augmentation to get better results.",
      "votes": null
    },
    {
      "id": "761754",
      "postDate": "03/02/2020 22:35:54",
      "content": "<p>I was going to say \"cannot wait to see this in action with data augmentation while oversampling\" it it appears your Notebook already does that. Thank you for sharing. Amazing work.</p>",
      "rawMarkdown": "I was going to say \"cannot wait to see this in action with data augmentation while oversampling\" it it appears your Notebook already does that. Thank you for sharing. Amazing work.",
      "votes": null
    },
    {
      "id": "768261",
      "postDate": "03/10/2020 15:37:57",
      "content": "<p>Great post - did you work on this further and commit a kernel with this technique?</p>",
      "rawMarkdown": "Great post - did you work on this further and commit a kernel with this technique?",
      "votes": null
    },
    {
      "id": "768400",
      "postDate": "03/10/2020 18:42:42",
      "content": "<p>Yes. The link to my kernel is on the post.</p>",
      "rawMarkdown": "Yes. The link to my kernel is on the post.",
      "votes": null
    },
    {
      "id": "768521",
      "postDate": "03/10/2020 22:53:39",
      "content": "<p>This is very interesting, just have to wait for TPU quota to try 😄</p>",
      "rawMarkdown": "This is very interesting, just have to wait for TPU quota to try 😄",
      "votes": null
    },
    {
      "id": "768691",
      "postDate": "03/11/2020 05:05:35",
      "content": "<p>Thanks for your sharing. </p>",
      "rawMarkdown": "Thanks for your sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 761754,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/02/2020 22:35:54",
      "content": "<p>I was going to say \"cannot wait to see this in action with data augmentation while oversampling\" it it appears your Notebook already does that. Thank you for sharing. Amazing work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 768261,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "03/10/2020 15:37:57",
      "content": "<p>Great post - did you work on this further and commit a kernel with this technique?</p>",
      "votes": null,
      "replies": [
        {
          "id": 768400,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/10/2020 18:42:42",
          "content": "<p>Yes. The link to my kernel is on the post.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768521,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "03/10/2020 22:53:39",
      "content": "<p>This is very interesting, just have to wait for TPU quota to try 😄</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 768691,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "03/11/2020 05:05:35",
      "content": "<p>Thanks for your sharing. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "760951": "I published a kernel [Tutorial -- Oversample](https://www.kaggle.com/yihdarshieh/tutorial-oversample?scriptVersionId=29496139) to demonstrate `oversampling` the training dataset.\n\nReference for oversampling: [Classification on imbalanced data](https://www.tensorflow.org/tutorials/structured_data/imbalanced_data).\n\nWhy oversampling instead of using sample weights or class weights: (from the above link)\n\n   &gt; If the training process were considering the whole dataset on each gradient update, this oversampling would be basically identical to the class weighting.\n\n   &gt; But when training the model batch-wise, as you did here, the oversampled data provides a smoother gradient signal: Instead of each positive example being shown in one batch with a large weight, they're shown in many different batches each time with a small weight.\n\n   &gt; This smoother gradient signal makes it easier to train the model.\n\nA sample result from my own training (1 means no oversampling, 100 means each class occurs at least approximately 100 times):\n\n\n\n    {\n        \"EfficientNetB7\": {\n            \"oversampling\": {\n                \"1\": {\n                    \"f1\": 0.9163497711078313,\n                    \"recall\": 0.9205266060440362,\n                    \"precision\": 0.919277486006819,\n                    \"acc\": 0.9240301724137931\n                },\n                \"100\": {\n                    \"f1\": 0.9306967499560146,\n                    \"recall\": 0.9385304126936429,\n                    \"precision\": 0.9277303112649578,\n                    \"acc\": 0.9296875\n                },\n                \"300\": {\n                    \"f1\": 0.9361016120193594,\n                    \"recall\": 0.9394270473352216,\n                    \"precision\": 0.939168918056952,\n                    \"acc\": 0.9366918103448276\n                },\n                \"800\": {\n                    \"f1\": 0.9404645585115528,\n                    \"recall\": 0.9427588505810944,\n                    \"precision\": 0.9440889947725051,\n                    \"acc\": 0.9391163793103449\n                }\n            }\n        }\n    }\n\nYou can use oversampling with more data augmentation to get better results.",
    "761754": "I was going to say \"cannot wait to see this in action with data augmentation while oversampling\" it it appears your Notebook already does that. Thank you for sharing. Amazing work.",
    "768261": "Great post - did you work on this further and commit a kernel with this technique?",
    "768400": "Yes. The link to my kernel is on the post.",
    "768521": "This is very interesting, just have to wait for TPU quota to try 😄",
    "768691": "Thanks for your sharing."
  },
  "source": "meta"
}