{
  "id": 212182,
  "title": "A sample way to solve the balance problem",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/212182",
  "author_name": "",
  "post_date": "2021-01-17T21:29:44.194987700Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I find a way to solve the balance problem and to save quota GPU/TPU.</p>\n<p><strong>First step</strong><br>\nyou devise your data to : </p>\n<ul>\n<li>label3 (we have 13158 samples but you can use undersampling  to reduce this number to the same number Of the sum of other labels)</li>\n<li>not labl3 (2577 + 2386 + 2189 + 1087 = 8239 samples)</li>\n</ul>\n<p><strong>Second step</strong><br>\nyou train your model with this data and now we have a binary classification (it easy to get accuracy +0.96 ) </p>\n<ul>\n<li>if the output of this model is label3 we predict lablel3</li>\n<li>if not we go to the next step</li>\n</ul>\n<p><strong>Third step</strong><br>\nYou train another model (multiclass classification) with the other labels  </p>\n<ul>\n<li>label0 (1087 samples)</li>\n<li>label1 (2189 samples)</li>\n<li>label2 (2386 samples)</li>\n<li>label4 (2577 samples)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4246137%2Fc02eb716ac8fa9555c5b99e65bd37196%2FUntitled%20Diagram%20(7).png?generation=1610918966826804&amp;alt=media\" alt=\"\"></p>\n<p><strong>NB: With this method you can focus in the second model which use just 38.5% of the data and that how you can save your quota GPU when you train your model</strong></p>",
  "messages": [
    {
      "id": "1157393",
      "postDate": "01/17/2021 21:29:44",
      "content": "<p>I find a way to solve the balance problem and to save quota GPU/TPU.</p>\n<p><strong>First step</strong><br>\nyou devise your data to : </p>\n<ul>\n<li>label3 (we have 13158 samples but you can use undersampling  to reduce this number to the same number Of the sum of other labels)</li>\n<li>not labl3 (2577 + 2386 + 2189 + 1087 = 8239 samples)</li>\n</ul>\n<p><strong>Second step</strong><br>\nyou train your model with this data and now we have a binary classification (it easy to get accuracy +0.96 ) </p>\n<ul>\n<li>if the output of this model is label3 we predict lablel3</li>\n<li>if not we go to the next step</li>\n</ul>\n<p><strong>Third step</strong><br>\nYou train another model (multiclass classification) with the other labels  </p>\n<ul>\n<li>label0 (1087 samples)</li>\n<li>label1 (2189 samples)</li>\n<li>label2 (2386 samples)</li>\n<li>label4 (2577 samples)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4246137%2Fc02eb716ac8fa9555c5b99e65bd37196%2FUntitled%20Diagram%20(7).png?generation=1610918966826804&amp;alt=media\" alt=\"\"></p>\n<p><strong>NB: With this method you can focus in the second model which use just 38.5% of the data and that how you can save your quota GPU when you train your model</strong></p>",
      "rawMarkdown": "I find a way to solve the balance problem and to save quota GPU/TPU.\n\n**First step**\nyou devise your data to : \n\n- label3 (we have 13158 samples but you can use undersampling  to reduce this number to the same number Of the sum of other labels)\n- not labl3 (2577 + 2386 + 2189 + 1087 = 8239 samples)\n\n**Second step**\nyou train your model with this data and now we have a binary classification (it easy to get accuracy +0.96 ) \n- if the output of this model is label3 we predict lablel3\n- if not we go to the next step\n\n**Third step**\nYou train another model (multiclass classification) with the other labels  \n- label0 (1087 samples)\n- label1 (2189 samples)\n- label2 (2386 samples)\n- label4 (2577 samples)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4246137%2Fc02eb716ac8fa9555c5b99e65bd37196%2FUntitled%20Diagram%20(7).png?generation=1610918966826804&alt=media)\n\n**NB: With this method you can focus in the second model which use just 38.5% of the data and that how you can save your quota GPU when you train your model**",
      "votes": null
    },
    {
      "id": "1157537",
      "postDate": "01/18/2021 02:00:28",
      "content": "<p>Interesting implements. But, does it work well in LB? In previous multi-class competition, I tried some methods like this, but it did not get good LB score.</p>",
      "rawMarkdown": "Interesting implements. But, does it work well in LB? In previous multi-class competition, I tried some methods like this, but it did not get good LB score.",
      "votes": null
    },
    {
      "id": "1157549",
      "postDate": "01/18/2021 02:18:31",
      "content": "<p>I used a similar method, but I'm still doing experiments. I've done experiments before, and it can increase by about 0.2%, but the single model can't exceed 90%, so I'm still doing similar experiments…</p>",
      "rawMarkdown": "I used a similar method, but I'm still doing experiments. I've done experiments before, and it can increase by about 0.2%, but the single model can't exceed 90%, so I'm still doing similar experiments...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1157537,
      "author_name": "woshifym",
      "author_url": "",
      "post_date": "01/18/2021 02:00:28",
      "content": "<p>Interesting implements. But, does it work well in LB? In previous multi-class competition, I tried some methods like this, but it did not get good LB score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1157549,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "01/18/2021 02:18:31",
          "content": "<p>I used a similar method, but I'm still doing experiments. I've done experiments before, and it can increase by about 0.2%, but the single model can't exceed 90%, so I'm still doing similar experiments…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1157393": "I find a way to solve the balance problem and to save quota GPU/TPU.\n\n**First step**\nyou devise your data to : \n\n- label3 (we have 13158 samples but you can use undersampling  to reduce this number to the same number Of the sum of other labels)\n- not labl3 (2577 + 2386 + 2189 + 1087 = 8239 samples)\n\n**Second step**\nyou train your model with this data and now we have a binary classification (it easy to get accuracy +0.96 ) \n- if the output of this model is label3 we predict lablel3\n- if not we go to the next step\n\n**Third step**\nYou train another model (multiclass classification) with the other labels  \n- label0 (1087 samples)\n- label1 (2189 samples)\n- label2 (2386 samples)\n- label4 (2577 samples)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4246137%2Fc02eb716ac8fa9555c5b99e65bd37196%2FUntitled%20Diagram%20(7).png?generation=1610918966826804&alt=media)\n\n**NB: With this method you can focus in the second model which use just 38.5% of the data and that how you can save your quota GPU when you train your model**",
    "1157537": "Interesting implements. But, does it work well in LB? In previous multi-class competition, I tried some methods like this, but it did not get good LB score.",
    "1157549": "I used a similar method, but I'm still doing experiments. I've done experiments before, and it can increase by about 0.2%, but the single model can't exceed 90%, so I'm still doing similar experiments..."
  },
  "source": "meta"
}