{
  "id": 100809,
  "title": "[Updated 7/22] How to deal with imbalanced data?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/100809",
  "author_name": "",
  "post_date": "2019-07-21T09:07:14.548140600Z",
  "votes": 7,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi all,\nAs you can see in <a href=\"https://www.kaggle.com/tanlikesmath/intro-aptos-diabetic-retinopathy-eda-starter\">Intro APTOS Diabetic Retinopathy (EDA &amp; Starter)</a> by <a href=\"https://www.kaggle.com/tanlikesmath\">ilovescience</a>, we have imbalanced data. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F4c7448973012103034f71567fa742108%2F__results___14_1.png?generation=1563698902563647&amp;alt=media\" alt=\"\"></p>\n\n<p>I'm new to deal with imbalanced data so let me know if I'm mistaken, but there are some ways.\n1. Use class weight (<a href=\"https://datascience.stackexchange.com/questions/13490/how-to-set-class-weights-for-imbalanced-classes-in-keras\">classification - How to set class weights for imbalanced classes in Keras?</a>)\n*<strong>the loss becomes a weighted average</strong>, where the weight of each sample is specified by <strong>class_weight</strong> and its corresponding class.*\n2. Use sample weight (<a href=\"https://stackoverflow.com/questions/43459317/keras-class-weight-vs-sample-weights-in-the-fit-generator\">tensorflow - Keras - class_weight vs sample_weights in the fit_generator</a>)</p>\n\n<p>*class_weight affects the relative weight of each class in the calculation of the objective function. sample_weights, as the name suggests, allows further control of the relative weight of samples hat belong to the same class*\n3. Focal loss\nTBD</p>\n\n<h2>Question 1</h2>\n\n<p>I'm using the multi label approach mentioned in <a href=\"https://www.kaggle.com/xhlulu/aptos-2019-densenet-keras-starter\">APTOS 2019: DenseNet Keras Starter </a> because of the class order. But I'm not sure what is the right way to deal with the imbalance data + multi label approach.\nCould someone chime in?</p>\n\n<h2>Question 2</h2>\n\n<p>Say we use sample_weights, how can we find the best weights? (I suppose there are some cases 'balanced' is not the best).</p>\n\n<h2>What I'm doing in Keras</h2>\n\n<p>But I'm not sure if I'm doing right...\n<code>\nsample_weights = class_weight.compute_sample_weight('balanced', y_train_multi)\n...snip...\nx_train, x_val, y_train, y_val, sample_weights_train, sample_weights_val = train_test_split(\n    x_train, y_train_multi, sample_weights,\n    test_size=gparams['test_split_size'], <br>\n    random_state=seed,\n)\n...snip...\ndata_generator = create_datagen().flow(x_train, y_train, batch_size=BATCH_SIZE, seed=seed, sample_weight=sample_weights_train)\n...snip...\nThen use data_generator in fit_generator\n</code></p>\n\n<h2>Updated:1 7/22/2019</h2>\n\n<p>I did the experiment above and used 'balanced' sample weights. I have two observations.\n- (1) It took more time to converge.\n- (2) I got lower CV score.</p>\n\n<p>(1) kinda makes sense. My model has to fetch and see more data to get balanced data.\nAbout (2), I'm not sure why this is happening. So I'd like your input here or answer to my Question 2 above.</p>\n\n<h2>Updated:2 7/22/2019 Real example of sample weights.</h2>\n\n<p>I have\ny_train_multi=\n<code>\n[1 1 1 0 0]\n [1 1 1 1 1]\n [1 1 0 0 0]\n [1 0 0 0 0]\n [1 0 0 0 0]\n [1 1 1 1 1]\n...snip...\n</code>\nCorresponding sample_weights is\n<code>\n[ 0.35618088 20.13115852  0.25792409  0.30278045  0.30278045 20.13115852 ...snip...]\n</code>\nYou can see 2nd and 6th element of y is  [1 1 1 1 1] which appears less in train set, so it has hihger sample_weight 20.13115852.</p>\n\n<h2>Updated:3 7/22/2019</h2>\n\n<p>I think <code>class_weight.compute_sample_weight('balanced', y_train_multi)</code> is wrong. I should have computed sample_weights before we do one hot encoding. Fixing and re-running my kernel.\nIt ended up 0.90355 which is a lot lower w/o sample weights.</p>\n\n<p>thanks in advance!</p>",
  "messages": [
    {
      "id": "581018",
      "postDate": "07/21/2019 09:07:14",
      "content": "<p>Hi all,\nAs you can see in <a href=\"https://www.kaggle.com/tanlikesmath/intro-aptos-diabetic-retinopathy-eda-starter\">Intro APTOS Diabetic Retinopathy (EDA &amp; Starter)</a> by <a href=\"https://www.kaggle.com/tanlikesmath\">ilovescience</a>, we have imbalanced data. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F4c7448973012103034f71567fa742108%2F__results___14_1.png?generation=1563698902563647&amp;alt=media\" alt=\"\"></p>\n\n<p>I'm new to deal with imbalanced data so let me know if I'm mistaken, but there are some ways.\n1. Use class weight (<a href=\"https://datascience.stackexchange.com/questions/13490/how-to-set-class-weights-for-imbalanced-classes-in-keras\">classification - How to set class weights for imbalanced classes in Keras?</a>)\n*<strong>the loss becomes a weighted average</strong>, where the weight of each sample is specified by <strong>class_weight</strong> and its corresponding class.*\n2. Use sample weight (<a href=\"https://stackoverflow.com/questions/43459317/keras-class-weight-vs-sample-weights-in-the-fit-generator\">tensorflow - Keras - class_weight vs sample_weights in the fit_generator</a>)</p>\n\n<p>*class_weight affects the relative weight of each class in the calculation of the objective function. sample_weights, as the name suggests, allows further control of the relative weight of samples hat belong to the same class*\n3. Focal loss\nTBD</p>\n\n<h2>Question 1</h2>\n\n<p>I'm using the multi label approach mentioned in <a href=\"https://www.kaggle.com/xhlulu/aptos-2019-densenet-keras-starter\">APTOS 2019: DenseNet Keras Starter </a> because of the class order. But I'm not sure what is the right way to deal with the imbalance data + multi label approach.\nCould someone chime in?</p>\n\n<h2>Question 2</h2>\n\n<p>Say we use sample_weights, how can we find the best weights? (I suppose there are some cases 'balanced' is not the best).</p>\n\n<h2>What I'm doing in Keras</h2>\n\n<p>But I'm not sure if I'm doing right...\n<code>\nsample_weights = class_weight.compute_sample_weight('balanced', y_train_multi)\n...snip...\nx_train, x_val, y_train, y_val, sample_weights_train, sample_weights_val = train_test_split(\n    x_train, y_train_multi, sample_weights,\n    test_size=gparams['test_split_size'], <br>\n    random_state=seed,\n)\n...snip...\ndata_generator = create_datagen().flow(x_train, y_train, batch_size=BATCH_SIZE, seed=seed, sample_weight=sample_weights_train)\n...snip...\nThen use data_generator in fit_generator\n</code></p>\n\n<h2>Updated:1 7/22/2019</h2>\n\n<p>I did the experiment above and used 'balanced' sample weights. I have two observations.\n- (1) It took more time to converge.\n- (2) I got lower CV score.</p>\n\n<p>(1) kinda makes sense. My model has to fetch and see more data to get balanced data.\nAbout (2), I'm not sure why this is happening. So I'd like your input here or answer to my Question 2 above.</p>\n\n<h2>Updated:2 7/22/2019 Real example of sample weights.</h2>\n\n<p>I have\ny_train_multi=\n<code>\n[1 1 1 0 0]\n [1 1 1 1 1]\n [1 1 0 0 0]\n [1 0 0 0 0]\n [1 0 0 0 0]\n [1 1 1 1 1]\n...snip...\n</code>\nCorresponding sample_weights is\n<code>\n[ 0.35618088 20.13115852  0.25792409  0.30278045  0.30278045 20.13115852 ...snip...]\n</code>\nYou can see 2nd and 6th element of y is  [1 1 1 1 1] which appears less in train set, so it has hihger sample_weight 20.13115852.</p>\n\n<h2>Updated:3 7/22/2019</h2>\n\n<p>I think <code>class_weight.compute_sample_weight('balanced', y_train_multi)</code> is wrong. I should have computed sample_weights before we do one hot encoding. Fixing and re-running my kernel.\nIt ended up 0.90355 which is a lot lower w/o sample weights.</p>\n\n<p>thanks in advance!</p>",
      "rawMarkdown": "Hi all,\nAs you can see in [Intro APTOS Diabetic Retinopathy (EDA &amp; Starter)](https://www.kaggle.com/tanlikesmath/intro-aptos-diabetic-retinopathy-eda-starter) by [ilovescience](https://www.kaggle.com/tanlikesmath), we have imbalanced data. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F4c7448973012103034f71567fa742108%2F__results___14_1.png?generation=1563698902563647&amp;alt=media)\n\nI'm new to deal with imbalanced data so let me know if I'm mistaken, but there are some ways.\n1. Use class weight ([classification - How to set class weights for imbalanced classes in Keras?](https://datascience.stackexchange.com/questions/13490/how-to-set-class-weights-for-imbalanced-classes-in-keras))\n*__the loss becomes a weighted average__, where the weight of each sample is specified by __class_weight__ and its corresponding class.*\n2. Use sample weight ([tensorflow - Keras - class_weight vs sample_weights in the fit_generator](https://stackoverflow.com/questions/43459317/keras-class-weight-vs-sample-weights-in-the-fit-generator))\n\n*class_weight affects the relative weight of each class in the calculation of the objective function. sample_weights, as the name suggests, allows further control of the relative weight of samples hat belong to the same class*\n3. Focal loss\nTBD\n\n## Question 1\nI'm using the multi label approach mentioned in [APTOS 2019: DenseNet Keras Starter ](https://www.kaggle.com/xhlulu/aptos-2019-densenet-keras-starter) because of the class order. But I'm not sure what is the right way to deal with the imbalance data + multi label approach.\nCould someone chime in?\n\n## Question 2\nSay we use sample_weights, how can we find the best weights? (I suppose there are some cases 'balanced' is not the best).\n\n## What I'm doing in Keras\nBut I'm not sure if I'm doing right...\n```\nsample_weights = class_weight.compute_sample_weight('balanced', y_train_multi)\n...snip...\nx_train, x_val, y_train, y_val, sample_weights_train, sample_weights_val = train_test_split(\n    x_train, y_train_multi, sample_weights,\n    test_size=gparams['test_split_size'],  \n    random_state=seed,\n)\n...snip...\ndata_generator = create_datagen().flow(x_train, y_train, batch_size=BATCH_SIZE, seed=seed, sample_weight=sample_weights_train)\n...snip...\nThen use data_generator in fit_generator\n```\n## Updated:1 7/22/2019\nI did the experiment above and used 'balanced' sample weights. I have two observations.\n- (1) It took more time to converge.\n- (2) I got lower CV score.\n\n(1) kinda makes sense. My model has to fetch and see more data to get balanced data.\nAbout (2), I'm not sure why this is happening. So I'd like your input here or answer to my Question 2 above.\n\n## Updated:2 7/22/2019 Real example of sample weights.\nI have\ny_train_multi=\n```\n[1 1 1 0 0]\n [1 1 1 1 1]\n [1 1 0 0 0]\n [1 0 0 0 0]\n [1 0 0 0 0]\n [1 1 1 1 1]\n...snip...\n```\nCorresponding sample_weights is\n```\n[ 0.35618088 20.13115852  0.25792409  0.30278045  0.30278045 20.13115852 ...snip...]\n```\nYou can see 2nd and 6th element of y is  [1 1 1 1 1] which appears less in train set, so it has hihger sample_weight 20.13115852.\n\n## Updated:3 7/22/2019\nI think ``class_weight.compute_sample_weight('balanced', y_train_multi)`` is wrong. I should have computed sample_weights before we do one hot encoding. Fixing and re-running my kernel.\nIt ended up 0.90355 which is a lot lower w/o sample weights.\n\nthanks in advance!",
      "votes": null
    },
    {
      "id": "581304",
      "postDate": "07/21/2019 18:27:39",
      "content": "<p>dear <a href=\"/higepon\">@higepon</a>  you can try oversampling</p>",
      "rawMarkdown": "dear @higepon  you can try oversampling",
      "votes": null
    },
    {
      "id": "581418",
      "postDate": "07/21/2019 23:20:08",
      "content": "<p>Hi, <a href=\"/mobassir\">@mobassir</a> \nThanks for your comment.</p>\n\n<p>Correct me if I'm wrong, but did you mean oversampling of low frequency classes using sample_weights?</p>\n\n<p>I tried the following which I assume samples all the classes balanced. But it didn't improve my CV.\n<code>sample_weights = class_weight.compute_sample_weight('balanced', y_train_multi)</code></p>\n\n<p>I appreciate your thoughts.</p>",
      "rawMarkdown": "Hi, @mobassir \nThanks for your comment.\n\nCorrect me if I'm wrong, but did you mean oversampling of low frequency classes using sample_weights?\n\nI tried the following which I assume samples all the classes balanced. But it didn't improve my CV.\n`sample_weights = class_weight.compute_sample_weight('balanced', y_train_multi)`\n\nI appreciate your thoughts.",
      "votes": null
    },
    {
      "id": "581821",
      "postDate": "07/22/2019 12:51:12",
      "content": "<p>Don't panic! Test set is super imbalanced, so train is good!!! (0, 1 class super small)</p>",
      "rawMarkdown": "Don't panic! Test set is super imbalanced, so train is good!!! (0, 1 class super small)",
      "votes": null
    },
    {
      "id": "581899",
      "postDate": "07/22/2019 14:11:13",
      "content": "<p>Thank you for your comment.\nCould you elaborate a bit more?</p>",
      "rawMarkdown": "Thank you for your comment.\nCould you elaborate a bit more?",
      "votes": null
    },
    {
      "id": "582195",
      "postDate": "07/22/2019 21:59:18",
      "content": "<p>We don't known proportion in private set, so can't do with that anything. i just use penalty like in metric qwk for predictions, validation qwk stable but low (</p>",
      "rawMarkdown": "We don't known proportion in private set, so can't do with that anything. i just use penalty like in metric qwk for predictions, validation qwk stable but low (",
      "votes": null
    },
    {
      "id": "582199",
      "postDate": "07/22/2019 22:08:05",
      "content": "<p>Thank you for your reply! I see your point.</p>",
      "rawMarkdown": "Thank you for your reply! I see your point.",
      "votes": null
    },
    {
      "id": "582278",
      "postDate": "07/23/2019 02:18:22",
      "content": "<p><a href=\"/higepon\">@higepon</a> My hypothesis is that the private/public test data should have small number of class 1/3 also. This is due to the human nature of 5-level estimation. Class 0/2/4 are easy to assign (no problem/ intermediate / severe). For class 1/3 You have to estimate “something between” (e.g. class 1 is between no problem and intermediate) and it is quite difficult for human to consistently justify that. In my confusion matrix analysis, the most difficult classes are classes 1 and 3 as well (tend to mis-classify to nearby classes).</p>\n\n<p>So my conclusion is not to worry too much about class imbalance. Just my hypothesis though :D</p>",
      "rawMarkdown": "higepon My hypothesis is that the private/public test data should have small number of class 1/3 also. This is due to the human nature of 5-level estimation. Class 0/2/4 are easy to assign (no problem/ intermediate / severe). For class 1/3 You have to estimate “something between” (e.g. class 1 is between no problem and intermediate) and it is quite difficult for human to consistently justify that. In my confusion matrix analysis, the most difficult classes are classes 1 and 3 as well (tend to mis-classify to nearby classes).\n\nSo my conclusion is not to worry too much about class imbalance. Just my hypothesis though :D",
      "votes": null
    },
    {
      "id": "582282",
      "postDate": "07/23/2019 02:25:48",
      "content": "<p>update ;</p>\n\n<p>after i read <a href=\"/juliencs\">@juliencs</a> 's reply, i think i make some mistakes about the imbalance problem, if i have some idea i will update later</p>",
      "rawMarkdown": "update ;\n\nafter i read @juliencs 's reply, i think i make some mistakes about the imbalance problem, if i have some idea i will update later",
      "votes": null
    },
    {
      "id": "582330",
      "postDate": "07/23/2019 04:18:33",
      "content": "<p>Neat</p>",
      "rawMarkdown": "Neat",
      "votes": null
    },
    {
      "id": "582343",
      "postDate": "07/23/2019 04:53:27",
      "content": "<p>Thanks for your comment!\nI thought it would be better if my model has more opportunities to see the most severe class so that it can predict better when testing.</p>",
      "rawMarkdown": "Thanks for your comment!\nI thought it would be better if my model has more opportunities to see the most severe class so that it can predict better when testing.",
      "votes": null
    },
    {
      "id": "582346",
      "postDate": "07/23/2019 04:54:59",
      "content": "<p>Thank you for your comment. It's really interesting to see your point of view and it makes sense.\nBtw I'm still learning to deal with imbalance and I thought it would be better if my model has more opportunities to see the most severe class so that it can predict better when testing. Do you think it's not worth considering?</p>",
      "rawMarkdown": "Thank you for your comment. It's really interesting to see your point of view and it makes sense.\nBtw I'm still learning to deal with imbalance and I thought it would be better if my model has more opportunities to see the most severe class so that it can predict better when testing. Do you think it's not worth considering?",
      "votes": null
    },
    {
      "id": "582362",
      "postDate": "07/23/2019 05:24:37",
      "content": "<p>Hi <a href=\"/higepon\">@higepon</a> , </p>\n\n<p>At the bottom line, for most severity examples [class 3&amp;4], I think it’s worth to adjust the weights. </p>\n\n<p>In fact, I did do a small adjustment on this case : I give more weights on class 3&amp;4 in validation test (since class 4 which is very important also contains small number of examples by nature [a lot more of healthy people than severely injured people])</p>\n\n<p>However, I cannot find any solid evidence to support that this method really benefits my performance or not [the benefit magnitude may lower than Keras fluctuation]</p>\n\n<p>BTW, what i did is just simply copy more class3&amp;4 in validation set.</p>",
      "rawMarkdown": "Hi @higepon , \n\nAt the bottom line, for most severity examples [class 3&amp;4], I think it’s worth to adjust the weights. \n\nIn fact, I did do a small adjustment on this case : I give more weights on class 3&amp;4 in validation test (since class 4 which is very important also contains small number of examples by nature [a lot more of healthy people than severely injured people])\n\nHowever, I cannot find any solid evidence to support that this method really benefits my performance or not [the benefit magnitude may lower than Keras fluctuation]\n\nBTW, what i did is just simply copy more class3&amp;4 in validation set.",
      "votes": null
    },
    {
      "id": "582509",
      "postDate": "07/23/2019 09:12:17",
      "content": "<p>Thanks for your great insights. I'm glad to know that your thoughts are aligned with mine.</p>\n\n<blockquote>\n  <p>BTW, what i did is just simply copy more class3&amp;4 in validation set.</p>\n</blockquote>\n\n<p>I had no idea about this. But it makes sense. Trying on my kernel now.</p>",
      "rawMarkdown": "Thanks for your great insights. I'm glad to know that your thoughts are aligned with mine.\n\n&gt; BTW, what i did is just simply copy more class3&amp;4 in validation set.\n\nI had no idea about this. But it makes sense. Trying on my kernel now.",
      "votes": null
    },
    {
      "id": "582999",
      "postDate": "07/23/2019 21:37:52",
      "content": "<p>I might note that oversampling decreased my public LB score significantly.</p>",
      "rawMarkdown": "I might note that oversampling decreased my public LB score significantly.",
      "votes": null
    },
    {
      "id": "583260",
      "postDate": "07/24/2019 08:34:45",
      "content": "<p>FYI\nI made class3&amp;4 repeat 4 times in val set and result was very poor. I'll experiment more.</p>",
      "rawMarkdown": "FYI\nI made class3&amp;4 repeat 4 times in val set and result was very poor. I'll experiment more.",
      "votes": null
    },
    {
      "id": "583275",
      "postDate": "07/24/2019 09:05:23",
      "content": "<p>It seems to me you may be mixing things up. The distribution of test data is not pertinent here.</p>\n\n<p>When your model learns over severely imbalanced data, it is more likely to try to optimize the most dominant class. Think of the gradient descent, trying to go in the best direction at each step, and that direction will be mostly defined by the dominant class, since your model will see dominant class data points most of the time. That's the reason some algorithms have a \"class_weight\" parameter : to make bigger steps when your model sees a minority class data point and hopefully minimize this way the effect of the dominant class. </p>\n\n<p>Whether these \"class_weight\" parameters make sense is another topic. I wish there was a similar trick for regression tasks.</p>\n\n<p>So here, no matter the distribution of the test set data, you're going to predict  with a model that focused on predicting the dominant class correctly, and will for this reason probably struggle with the other classes more than it should.</p>",
      "rawMarkdown": "It seems to me you may be mixing things up. The distribution of test data is not pertinent here.\n\nWhen your model learns over severely imbalanced data, it is more likely to try to optimize the most dominant class. Think of the gradient descent, trying to go in the best direction at each step, and that direction will be mostly defined by the dominant class, since your model will see dominant class data points most of the time. That's the reason some algorithms have a \"class_weight\" parameter : to make bigger steps when your model sees a minority class data point and hopefully minimize this way the effect of the dominant class. \n\nWhether these \"class_weight\" parameters make sense is another topic. I wish there was a similar trick for regression tasks.\n\nSo here, no matter the distribution of the test set data, you're going to predict  with a model that focused on predicting the dominant class correctly, and will for this reason probably struggle with the other classes more than it should.",
      "votes": null
    },
    {
      "id": "583291",
      "postDate": "07/24/2019 09:38:53",
      "content": "<p>Thank you so much for your comment. Now I have better understanding and I'm on the same age with what you described above.</p>\n\n<p>Surprisingly there are many ways to deal with imbalance.\n- class weight\n- sample weight\n- direct oversampling / undersampling</p>",
      "rawMarkdown": "Thank you so much for your comment. Now I have better understanding and I'm on the same age with what you described above.\n\nSurprisingly there are many ways to deal with imbalance.\n- class weight\n- sample weight\n- direct oversampling / undersampling",
      "votes": null
    },
    {
      "id": "584590",
      "postDate": "07/26/2019 07:44:41",
      "content": "<p>When you have big enough sample from each class, you will not worry about class imbalance. My suggestion is to try pretraining on previous competition data.</p>",
      "rawMarkdown": "When you have big enough sample from each class, you will not worry about class imbalance. My suggestion is to try pretraining on previous competition data.",
      "votes": null
    },
    {
      "id": "614577",
      "postDate": "08/31/2019 17:44:12",
      "content": "<p>Hey! you were stuck on 370th for a long while. I saw your posts where you had been struggling to improve beyond 0.79. What did you to jump to 0.81ish?\nI am trying to improve my single model with new competition data only to hit 0.79ish but so far not getting beyond 0.75. The reason being that any pre training with old competition data doesn't give me more than 0.01 improvement.</p>",
      "rawMarkdown": "Hey! you were stuck on 370th for a long while. I saw your posts where you had been struggling to improve beyond 0.79. What did you to jump to 0.81ish?\nI am trying to improve my single model with new competition data only to hit 0.79ish but so far not getting beyond 0.75. The reason being that any pre training with old competition data doesn't give me more than 0.01 improvement.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 581304,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "07/21/2019 18:27:39",
      "content": "<p>dear <a href=\"/higepon\">@higepon</a>  you can try oversampling</p>",
      "votes": null,
      "replies": [
        {
          "id": 581418,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/21/2019 23:20:08",
          "content": "<p>Hi, <a href=\"/mobassir\">@mobassir</a> \nThanks for your comment.</p>\n\n<p>Correct me if I'm wrong, but did you mean oversampling of low frequency classes using sample_weights?</p>\n\n<p>I tried the following which I assume samples all the classes balanced. But it didn't improve my CV.\n<code>sample_weights = class_weight.compute_sample_weight('balanced', y_train_multi)</code></p>\n\n<p>I appreciate your thoughts.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 581821,
      "author_name": "leighplt",
      "author_url": "",
      "post_date": "07/22/2019 12:51:12",
      "content": "<p>Don't panic! Test set is super imbalanced, so train is good!!! (0, 1 class super small)</p>",
      "votes": null,
      "replies": [
        {
          "id": 581899,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/22/2019 14:11:13",
          "content": "<p>Thank you for your comment.\nCould you elaborate a bit more?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 582195,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "07/22/2019 21:59:18",
          "content": "<p>We don't known proportion in private set, so can't do with that anything. i just use penalty like in metric qwk for predictions, validation qwk stable but low (</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 582199,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/22/2019 22:08:05",
          "content": "<p>Thank you for your reply! I see your point.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 582278,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "07/23/2019 02:18:22",
      "content": "<p><a href=\"/higepon\">@higepon</a> My hypothesis is that the private/public test data should have small number of class 1/3 also. This is due to the human nature of 5-level estimation. Class 0/2/4 are easy to assign (no problem/ intermediate / severe). For class 1/3 You have to estimate “something between” (e.g. class 1 is between no problem and intermediate) and it is quite difficult for human to consistently justify that. In my confusion matrix analysis, the most difficult classes are classes 1 and 3 as well (tend to mis-classify to nearby classes).</p>\n\n<p>So my conclusion is not to worry too much about class imbalance. Just my hypothesis though :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 582346,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/23/2019 04:54:59",
          "content": "<p>Thank you for your comment. It's really interesting to see your point of view and it makes sense.\nBtw I'm still learning to deal with imbalance and I thought it would be better if my model has more opportunities to see the most severe class so that it can predict better when testing. Do you think it's not worth considering?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 582362,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "07/23/2019 05:24:37",
          "content": "<p>Hi <a href=\"/higepon\">@higepon</a> , </p>\n\n<p>At the bottom line, for most severity examples [class 3&amp;4], I think it’s worth to adjust the weights. </p>\n\n<p>In fact, I did do a small adjustment on this case : I give more weights on class 3&amp;4 in validation test (since class 4 which is very important also contains small number of examples by nature [a lot more of healthy people than severely injured people])</p>\n\n<p>However, I cannot find any solid evidence to support that this method really benefits my performance or not [the benefit magnitude may lower than Keras fluctuation]</p>\n\n<p>BTW, what i did is just simply copy more class3&amp;4 in validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 582509,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/23/2019 09:12:17",
          "content": "<p>Thanks for your great insights. I'm glad to know that your thoughts are aligned with mine.</p>\n\n<blockquote>\n  <p>BTW, what i did is just simply copy more class3&amp;4 in validation set.</p>\n</blockquote>\n\n<p>I had no idea about this. But it makes sense. Trying on my kernel now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583260,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/24/2019 08:34:45",
          "content": "<p>FYI\nI made class3&amp;4 repeat 4 times in val set and result was very poor. I'll experiment more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584590,
          "author_name": "yohalf",
          "author_url": "",
          "post_date": "07/26/2019 07:44:41",
          "content": "<p>When you have big enough sample from each class, you will not worry about class imbalance. My suggestion is to try pretraining on previous competition data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 582282,
      "author_name": "suzhengpeng",
      "author_url": "",
      "post_date": "07/23/2019 02:25:48",
      "content": "<p>update ;</p>\n\n<p>after i read <a href=\"/juliencs\">@juliencs</a> 's reply, i think i make some mistakes about the imbalance problem, if i have some idea i will update later</p>",
      "votes": null,
      "replies": [
        {
          "id": 582343,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/23/2019 04:53:27",
          "content": "<p>Thanks for your comment!\nI thought it would be better if my model has more opportunities to see the most severe class so that it can predict better when testing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583275,
          "author_name": "juliencs",
          "author_url": "",
          "post_date": "07/24/2019 09:05:23",
          "content": "<p>It seems to me you may be mixing things up. The distribution of test data is not pertinent here.</p>\n\n<p>When your model learns over severely imbalanced data, it is more likely to try to optimize the most dominant class. Think of the gradient descent, trying to go in the best direction at each step, and that direction will be mostly defined by the dominant class, since your model will see dominant class data points most of the time. That's the reason some algorithms have a \"class_weight\" parameter : to make bigger steps when your model sees a minority class data point and hopefully minimize this way the effect of the dominant class. </p>\n\n<p>Whether these \"class_weight\" parameters make sense is another topic. I wish there was a similar trick for regression tasks.</p>\n\n<p>So here, no matter the distribution of the test set data, you're going to predict  with a model that focused on predicting the dominant class correctly, and will for this reason probably struggle with the other classes more than it should.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583291,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/24/2019 09:38:53",
          "content": "<p>Thank you so much for your comment. Now I have better understanding and I'm on the same age with what you described above.</p>\n\n<p>Surprisingly there are many ways to deal with imbalance.\n- class weight\n- sample weight\n- direct oversampling / undersampling</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 582330,
      "author_name": "joycevarg",
      "author_url": "",
      "post_date": "07/23/2019 04:18:33",
      "content": "<p>Neat</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 582999,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "07/23/2019 21:37:52",
      "content": "<p>I might note that oversampling decreased my public LB score significantly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 614577,
          "author_name": "basitsheikh",
          "author_url": "",
          "post_date": "08/31/2019 17:44:12",
          "content": "<p>Hey! you were stuck on 370th for a long while. I saw your posts where you had been struggling to improve beyond 0.79. What did you to jump to 0.81ish?\nI am trying to improve my single model with new competition data only to hit 0.79ish but so far not getting beyond 0.75. The reason being that any pre training with old competition data doesn't give me more than 0.01 improvement.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "581018": "Hi all,\nAs you can see in [Intro APTOS Diabetic Retinopathy (EDA &amp; Starter)](https://www.kaggle.com/tanlikesmath/intro-aptos-diabetic-retinopathy-eda-starter) by [ilovescience](https://www.kaggle.com/tanlikesmath), we have imbalanced data. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F4c7448973012103034f71567fa742108%2F__results___14_1.png?generation=1563698902563647&amp;alt=media)\n\nI'm new to deal with imbalanced data so let me know if I'm mistaken, but there are some ways.\n1. Use class weight ([classification - How to set class weights for imbalanced classes in Keras?](https://datascience.stackexchange.com/questions/13490/how-to-set-class-weights-for-imbalanced-classes-in-keras))\n*__the loss becomes a weighted average__, where the weight of each sample is specified by __class_weight__ and its corresponding class.*\n2. Use sample weight ([tensorflow - Keras - class_weight vs sample_weights in the fit_generator](https://stackoverflow.com/questions/43459317/keras-class-weight-vs-sample-weights-in-the-fit-generator))\n\n*class_weight affects the relative weight of each class in the calculation of the objective function. sample_weights, as the name suggests, allows further control of the relative weight of samples hat belong to the same class*\n3. Focal loss\nTBD\n\n## Question 1\nI'm using the multi label approach mentioned in [APTOS 2019: DenseNet Keras Starter ](https://www.kaggle.com/xhlulu/aptos-2019-densenet-keras-starter) because of the class order. But I'm not sure what is the right way to deal with the imbalance data + multi label approach.\nCould someone chime in?\n\n## Question 2\nSay we use sample_weights, how can we find the best weights? (I suppose there are some cases 'balanced' is not the best).\n\n## What I'm doing in Keras\nBut I'm not sure if I'm doing right...\n```\nsample_weights = class_weight.compute_sample_weight('balanced', y_train_multi)\n...snip...\nx_train, x_val, y_train, y_val, sample_weights_train, sample_weights_val = train_test_split(\n    x_train, y_train_multi, sample_weights,\n    test_size=gparams['test_split_size'],  \n    random_state=seed,\n)\n...snip...\ndata_generator = create_datagen().flow(x_train, y_train, batch_size=BATCH_SIZE, seed=seed, sample_weight=sample_weights_train)\n...snip...\nThen use data_generator in fit_generator\n```\n## Updated:1 7/22/2019\nI did the experiment above and used 'balanced' sample weights. I have two observations.\n- (1) It took more time to converge.\n- (2) I got lower CV score.\n\n(1) kinda makes sense. My model has to fetch and see more data to get balanced data.\nAbout (2), I'm not sure why this is happening. So I'd like your input here or answer to my Question 2 above.\n\n## Updated:2 7/22/2019 Real example of sample weights.\nI have\ny_train_multi=\n```\n[1 1 1 0 0]\n [1 1 1 1 1]\n [1 1 0 0 0]\n [1 0 0 0 0]\n [1 0 0 0 0]\n [1 1 1 1 1]\n...snip...\n```\nCorresponding sample_weights is\n```\n[ 0.35618088 20.13115852  0.25792409  0.30278045  0.30278045 20.13115852 ...snip...]\n```\nYou can see 2nd and 6th element of y is  [1 1 1 1 1] which appears less in train set, so it has hihger sample_weight 20.13115852.\n\n## Updated:3 7/22/2019\nI think ``class_weight.compute_sample_weight('balanced', y_train_multi)`` is wrong. I should have computed sample_weights before we do one hot encoding. Fixing and re-running my kernel.\nIt ended up 0.90355 which is a lot lower w/o sample weights.\n\nthanks in advance!",
    "581304": "dear @higepon  you can try oversampling",
    "581418": "Hi, @mobassir \nThanks for your comment.\n\nCorrect me if I'm wrong, but did you mean oversampling of low frequency classes using sample_weights?\n\nI tried the following which I assume samples all the classes balanced. But it didn't improve my CV.\n`sample_weights = class_weight.compute_sample_weight('balanced', y_train_multi)`\n\nI appreciate your thoughts.",
    "581821": "Don't panic! Test set is super imbalanced, so train is good!!! (0, 1 class super small)",
    "581899": "Thank you for your comment.\nCould you elaborate a bit more?",
    "582195": "We don't known proportion in private set, so can't do with that anything. i just use penalty like in metric qwk for predictions, validation qwk stable but low (",
    "582199": "Thank you for your reply! I see your point.",
    "582278": "higepon My hypothesis is that the private/public test data should have small number of class 1/3 also. This is due to the human nature of 5-level estimation. Class 0/2/4 are easy to assign (no problem/ intermediate / severe). For class 1/3 You have to estimate “something between” (e.g. class 1 is between no problem and intermediate) and it is quite difficult for human to consistently justify that. In my confusion matrix analysis, the most difficult classes are classes 1 and 3 as well (tend to mis-classify to nearby classes).\n\nSo my conclusion is not to worry too much about class imbalance. Just my hypothesis though :D",
    "582282": "update ;\n\nafter i read @juliencs 's reply, i think i make some mistakes about the imbalance problem, if i have some idea i will update later",
    "582330": "Neat",
    "582343": "Thanks for your comment!\nI thought it would be better if my model has more opportunities to see the most severe class so that it can predict better when testing.",
    "582346": "Thank you for your comment. It's really interesting to see your point of view and it makes sense.\nBtw I'm still learning to deal with imbalance and I thought it would be better if my model has more opportunities to see the most severe class so that it can predict better when testing. Do you think it's not worth considering?",
    "582362": "Hi @higepon , \n\nAt the bottom line, for most severity examples [class 3&amp;4], I think it’s worth to adjust the weights. \n\nIn fact, I did do a small adjustment on this case : I give more weights on class 3&amp;4 in validation test (since class 4 which is very important also contains small number of examples by nature [a lot more of healthy people than severely injured people])\n\nHowever, I cannot find any solid evidence to support that this method really benefits my performance or not [the benefit magnitude may lower than Keras fluctuation]\n\nBTW, what i did is just simply copy more class3&amp;4 in validation set.",
    "582509": "Thanks for your great insights. I'm glad to know that your thoughts are aligned with mine.\n\n&gt; BTW, what i did is just simply copy more class3&amp;4 in validation set.\n\nI had no idea about this. But it makes sense. Trying on my kernel now.",
    "582999": "I might note that oversampling decreased my public LB score significantly.",
    "583260": "FYI\nI made class3&amp;4 repeat 4 times in val set and result was very poor. I'll experiment more.",
    "583275": "It seems to me you may be mixing things up. The distribution of test data is not pertinent here.\n\nWhen your model learns over severely imbalanced data, it is more likely to try to optimize the most dominant class. Think of the gradient descent, trying to go in the best direction at each step, and that direction will be mostly defined by the dominant class, since your model will see dominant class data points most of the time. That's the reason some algorithms have a \"class_weight\" parameter : to make bigger steps when your model sees a minority class data point and hopefully minimize this way the effect of the dominant class. \n\nWhether these \"class_weight\" parameters make sense is another topic. I wish there was a similar trick for regression tasks.\n\nSo here, no matter the distribution of the test set data, you're going to predict  with a model that focused on predicting the dominant class correctly, and will for this reason probably struggle with the other classes more than it should.",
    "583291": "Thank you so much for your comment. Now I have better understanding and I'm on the same age with what you described above.\n\nSurprisingly there are many ways to deal with imbalance.\n- class weight\n- sample weight\n- direct oversampling / undersampling",
    "584590": "When you have big enough sample from each class, you will not worry about class imbalance. My suggestion is to try pretraining on previous competition data.",
    "614577": "Hey! you were stuck on 370th for a long while. I saw your posts where you had been struggling to improve beyond 0.79. What did you to jump to 0.81ish?\nI am trying to improve my single model with new competition data only to hit 0.79ish but so far not getting beyond 0.75. The reason being that any pre training with old competition data doesn't give me more than 0.01 improvement."
  },
  "source": "meta"
}