{
  "id": 200855,
  "title": " Synthetic Minority Oversampling Technique (SMOTE)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200855",
  "author_name": "",
  "post_date": "2020-12-02T06:01:23.925225Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Has anyone tried SMOTE on the data , since using class_weights and kfolds is not helping much.</p>",
  "messages": [
    {
      "id": "1099156",
      "postDate": "12/02/2020 06:01:23",
      "content": "<p>Has anyone tried SMOTE on the data , since using class_weights and kfolds is not helping much.</p>",
      "rawMarkdown": "Has anyone tried SMOTE on the data , since using class_weights and kfolds is not helping much.",
      "votes": null
    },
    {
      "id": "1101431",
      "postDate": "12/03/2020 22:28:19",
      "content": "<p>Paper if anyone is interested: <a href=\"https://arxiv.org/pdf/1106.1813.pdf\" target=\"_blank\">https://arxiv.org/pdf/1106.1813.pdf</a></p>",
      "rawMarkdown": "Paper if anyone is interested: https://arxiv.org/pdf/1106.1813.pdf",
      "votes": null
    },
    {
      "id": "1101786",
      "postDate": "12/04/2020 09:01:19",
      "content": "<p>I only know some structured data could use SMOTE, like in some data mining competition.<br>\nIn CV field, how to use SMOTE is a question, maybe by some data augmentation? <br>\nHope someone else know that.</p>",
      "rawMarkdown": "I only know some structured data could use SMOTE, like in some data mining competition.\nIn CV field, how to use SMOTE is a question, maybe by some data augmentation? \nHope someone else know that.",
      "votes": null
    },
    {
      "id": "1101856",
      "postDate": "12/04/2020 10:37:07",
      "content": "<p>Since the distribution of the minority categories is similar, I plan to enhance the minority categories uniformly and ignore the majority categories.</p>\n<ol>\n<li>First classify the 5 categories of image data into 5 categories</li>\n<li>Images in the most frequent category are not considered for enhancement</li>\n<li>Unified enhancement of images in a few categories (because the number of other categories is similar)<br>\ncode show as below:<br>\nif class==3:<br>\npass<br>\nelse:<br>\nData enhancement<br>\nThis is my initial thoughts.</li>\n</ol>",
      "rawMarkdown": "Since the distribution of the minority categories is similar, I plan to enhance the minority categories uniformly and ignore the majority categories.\n\n1. First classify the 5 categories of image data into 5 categories\n2. Images in the most frequent category are not considered for enhancement\n3. Unified enhancement of images in a few categories (because the number of other categories is similar)\ncode show as below:\nif class==3:\n   pass\nelse:\n  Data enhancement\nThis is my initial thoughts.",
      "votes": null
    },
    {
      "id": "1101864",
      "postDate": "12/04/2020 10:44:21",
      "content": "<p>I think it can be used in CV , I don't know how. This is an interesting tutorial I came across <a href=\"https://medium.com/swlh/how-to-use-smote-for-dealing-with-imbalanced-image-dataset-for-solving-classification-problems-3aba7d2b9cad\" target=\"_blank\">https://medium.com/swlh/how-to-use-smote-for-dealing-with-imbalanced-image-dataset-for-solving-classification-problems-3aba7d2b9cad</a></p>",
      "rawMarkdown": "I think it can be used in CV , I don't know how. This is an interesting tutorial I came across https://medium.com/swlh/how-to-use-smote-for-dealing-with-imbalanced-image-dataset-for-solving-classification-problems-3aba7d2b9cad",
      "votes": null
    },
    {
      "id": "1101870",
      "postDate": "12/04/2020 10:53:55",
      "content": "<p>Very important information, learned</p>",
      "rawMarkdown": "Very important information, learned",
      "votes": null
    },
    {
      "id": "1101963",
      "postDate": "12/04/2020 12:59:44",
      "content": "<p>I admit I haven't tested it thoroughly yet. My guess is that upsample of minority classes plus augementations may help.</p>\n<pre><code>    train_images_ = df_train.loc[df_train['label']!=3, ['image_id', 'label']]\n    train_images_m = [os.path.join(training_data_path, i) for i in train_images_.image_id.values.tolist()]\n    train_targets_m = train_images_['label']    \n\n    for i in range(N_UPSAMPLE):\n        train_images = train_images + train_images_m\n        train_targets = np.concatenate([train_targets, train_targets_m])\n</code></pre>\n<p>So we add minority images to the train dataset and then apply augmentations to the whole train dataset, mimicing SMOTE in a way.</p>",
      "rawMarkdown": "I admit I haven't tested it thoroughly yet. My guess is that upsample of minority classes plus augementations may help.\n\n\n        train_images_ = df_train.loc[df_train['label']!=3, ['image_id', 'label']]\n        train_images_m = [os.path.join(training_data_path, i) for i in train_images_.image_id.values.tolist()]\n        train_targets_m = train_images_['label']    \n\n        for i in range(N_UPSAMPLE):\n            train_images = train_images + train_images_m\n            train_targets = np.concatenate([train_targets, train_targets_m])\n\n\n\nSo we add minority images to the train dataset and then apply augmentations to the whole train dataset, mimicing SMOTE in a way.",
      "votes": null
    },
    {
      "id": "1101981",
      "postDate": "12/04/2020 13:12:40",
      "content": "<p>Yes, this might work 👍</p>",
      "rawMarkdown": "Yes, this might work 👍",
      "votes": null
    },
    {
      "id": "1102023",
      "postDate": "12/04/2020 13:46:51",
      "content": "<p>I mean something like <a href=\"https://www.kaggle.com/dunklerwald/pytorch-efficientnet-with-tta-training\" target=\"_blank\">this</a>. In one of earlier versions I tried to turn on upsample (UPSAMPLE=True, N_UPSAMPLE = 1) but results got worse (higher overfitting).  I assume there was an issue with a model itself and upsample just made it worse. As soon as I stabilize my model, I want to try UPSAMPLE=True again. My gut feeling tells me it won't help and if I want to upsample minority, I need high-quality synthetic images (GAN etc) anyway. But who knows.</p>",
      "rawMarkdown": "I mean something like [this](https://www.kaggle.com/dunklerwald/pytorch-efficientnet-with-tta-training). In one of earlier versions I tried to turn on upsample (UPSAMPLE=True, N_UPSAMPLE = 1) but results got worse (higher overfitting).  I assume there was an issue with a model itself and upsample just made it worse. As soon as I stabilize my model, I want to try UPSAMPLE=True again. My gut feeling tells me it won't help and if I want to upsample minority, I need high-quality synthetic images (GAN etc) anyway. But who knows.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1101431,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "12/03/2020 22:28:19",
      "content": "<p>Paper if anyone is interested: <a href=\"https://arxiv.org/pdf/1106.1813.pdf\" target=\"_blank\">https://arxiv.org/pdf/1106.1813.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1101786,
      "author_name": "fire15",
      "author_url": "",
      "post_date": "12/04/2020 09:01:19",
      "content": "<p>I only know some structured data could use SMOTE, like in some data mining competition.<br>\nIn CV field, how to use SMOTE is a question, maybe by some data augmentation? <br>\nHope someone else know that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1101864,
          "author_name": "akshatdevve",
          "author_url": "",
          "post_date": "12/04/2020 10:44:21",
          "content": "<p>I think it can be used in CV , I don't know how. This is an interesting tutorial I came across <a href=\"https://medium.com/swlh/how-to-use-smote-for-dealing-with-imbalanced-image-dataset-for-solving-classification-problems-3aba7d2b9cad\" target=\"_blank\">https://medium.com/swlh/how-to-use-smote-for-dealing-with-imbalanced-image-dataset-for-solving-classification-problems-3aba7d2b9cad</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101870,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "12/04/2020 10:53:55",
          "content": "<p>Very important information, learned</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1101856,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/04/2020 10:37:07",
      "content": "<p>Since the distribution of the minority categories is similar, I plan to enhance the minority categories uniformly and ignore the majority categories.</p>\n<ol>\n<li>First classify the 5 categories of image data into 5 categories</li>\n<li>Images in the most frequent category are not considered for enhancement</li>\n<li>Unified enhancement of images in a few categories (because the number of other categories is similar)<br>\ncode show as below:<br>\nif class==3:<br>\npass<br>\nelse:<br>\nData enhancement<br>\nThis is my initial thoughts.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1101963,
      "author_name": "dunklerwald",
      "author_url": "",
      "post_date": "12/04/2020 12:59:44",
      "content": "<p>I admit I haven't tested it thoroughly yet. My guess is that upsample of minority classes plus augementations may help.</p>\n<pre><code>    train_images_ = df_train.loc[df_train['label']!=3, ['image_id', 'label']]\n    train_images_m = [os.path.join(training_data_path, i) for i in train_images_.image_id.values.tolist()]\n    train_targets_m = train_images_['label']    \n\n    for i in range(N_UPSAMPLE):\n        train_images = train_images + train_images_m\n        train_targets = np.concatenate([train_targets, train_targets_m])\n</code></pre>\n<p>So we add minority images to the train dataset and then apply augmentations to the whole train dataset, mimicing SMOTE in a way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1101981,
          "author_name": "akshatdevve",
          "author_url": "",
          "post_date": "12/04/2020 13:12:40",
          "content": "<p>Yes, this might work 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1102023,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "12/04/2020 13:46:51",
          "content": "<p>I mean something like <a href=\"https://www.kaggle.com/dunklerwald/pytorch-efficientnet-with-tta-training\" target=\"_blank\">this</a>. In one of earlier versions I tried to turn on upsample (UPSAMPLE=True, N_UPSAMPLE = 1) but results got worse (higher overfitting).  I assume there was an issue with a model itself and upsample just made it worse. As soon as I stabilize my model, I want to try UPSAMPLE=True again. My gut feeling tells me it won't help and if I want to upsample minority, I need high-quality synthetic images (GAN etc) anyway. But who knows.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1099156": "Has anyone tried SMOTE on the data , since using class_weights and kfolds is not helping much.",
    "1101431": "Paper if anyone is interested: https://arxiv.org/pdf/1106.1813.pdf",
    "1101786": "I only know some structured data could use SMOTE, like in some data mining competition.\nIn CV field, how to use SMOTE is a question, maybe by some data augmentation? \nHope someone else know that.",
    "1101856": "Since the distribution of the minority categories is similar, I plan to enhance the minority categories uniformly and ignore the majority categories.\n\n1. First classify the 5 categories of image data into 5 categories\n2. Images in the most frequent category are not considered for enhancement\n3. Unified enhancement of images in a few categories (because the number of other categories is similar)\ncode show as below:\nif class==3:\n   pass\nelse:\n  Data enhancement\nThis is my initial thoughts.",
    "1101864": "I think it can be used in CV , I don't know how. This is an interesting tutorial I came across https://medium.com/swlh/how-to-use-smote-for-dealing-with-imbalanced-image-dataset-for-solving-classification-problems-3aba7d2b9cad",
    "1101870": "Very important information, learned",
    "1101963": "I admit I haven't tested it thoroughly yet. My guess is that upsample of minority classes plus augementations may help.\n\n\n        train_images_ = df_train.loc[df_train['label']!=3, ['image_id', 'label']]\n        train_images_m = [os.path.join(training_data_path, i) for i in train_images_.image_id.values.tolist()]\n        train_targets_m = train_images_['label']    \n\n        for i in range(N_UPSAMPLE):\n            train_images = train_images + train_images_m\n            train_targets = np.concatenate([train_targets, train_targets_m])\n\n\n\nSo we add minority images to the train dataset and then apply augmentations to the whole train dataset, mimicing SMOTE in a way.",
    "1101981": "Yes, this might work 👍",
    "1102023": "I mean something like [this](https://www.kaggle.com/dunklerwald/pytorch-efficientnet-with-tta-training). In one of earlier versions I tried to turn on upsample (UPSAMPLE=True, N_UPSAMPLE = 1) but results got worse (higher overfitting).  I assume there was an issue with a model itself and upsample just made it worse. As soon as I stabilize my model, I want to try UPSAMPLE=True again. My gut feeling tells me it won't help and if I want to upsample minority, I need high-quality synthetic images (GAN etc) anyway. But who knows."
  },
  "source": "meta"
}