{
  "id": 142107,
  "title": "Dataset is imbalanced",
  "url": "/competitions/flower-classification-with-tpus/discussion/142107",
  "author_name": "ⵀⴰⵜⴻⵎ ⵎ'ⵀⴰⵎⴻⴷ ⴰⵎⵉⵏⴻ",
  "post_date": "2020-04-09T00:15:22.332000",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>you may have noticed that the <strong>dataset is imbalanced</strong> , and to deal with it you my use the <strong>class_weight ** parameter in model.fit() \nmodel.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=3, **class_weight=class_weights</strong>)</p>\n\n<p>it can be used to assigned different weights to different classes, to deal with the problem of imbalanced data , when you have more sample for one class then another (exemple you have 400 image of class 1 and 20 image of class 2 and 10 image of class3 etc )</p>\n\n<p>to calculate the class_weights i use this code\nit my be interesting for you , and <strong>do not forget to upvote</strong>\n<em><strong>you may also fork my public kernel\n<a href=\"https://www.kaggle.com/hatemamine/flowertpuwin\">https://www.kaggle.com/hatemamine/flowertpuwin</a></strong></em>\n`\nlabels_train = training_dataset.map(lambda image, label: label).unbatch()\nT_correct_labels = next(iter(labels_train.batch(NUM_TRAINING_IMAGES))).numpy() </p>\n\n<h1>get everything as one batch</h1>\n\n<p>tmps=[]</p>\n\n<h1>for each labele calculate how much it occur</h1>\n\n<p>for x in range (len(CLASSES)):\n    tmp=0\n    for i in range(len(T_correct_labels)):\n        if x == T_correct_labels[i]:\n            tmp= tmp+1\n    tmps.append(tmp)</p>\n\n<p>tmps = np.array(tmps)\nprint(tmps.min())\nprint(tmps.max())\nprint(CLASSES[tmps.argmax()])\nprint(tmps.argmax())\nprint(CLASSES[tmps.argmin()])\nplt.plot(tmps) # plotting by columns\nplt.show()   </p>\n\n<h1>on keras you have to pass a dictionary to classweight parameter in model.fit()</h1>\n\n<h1>example classweights ={0:50, 1:0.5, 2:20}</h1>\n\n<p>class_weights = { i : ((tmps.max()+1)-tmps[i]) for i in range(0, len(tmps) ) }\nprint(class_weights)</p>\n\n<h1>but on TPU your have to pass flat list of weights (not a dictionary as written in the Keras docs)</h1>\n\n<h1>so you have just to convert your dictionary to a flat list</h1>\n\n<p>class_weights =[item for k in class_weights for item in (k, class_weights[k])]\nprint(class_weights)`</p>",
  "messages": [
    {
      "id": 801927,
      "postDate": "2020-04-09T00:15:22.333Z",
      "content": "<p>you may have noticed that the <strong>dataset is imbalanced</strong> , and to deal with it you my use the <strong>class_weight ** parameter in model.fit() \nmodel.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=3, **class_weight=class_weights</strong>)</p>\n\n<p>it can be used to assigned different weights to different classes, to deal with the problem of imbalanced data , when you have more sample for one class then another (exemple you have 400 image of class 1 and 20 image of class 2 and 10 image of class3 etc )</p>\n\n<p>to calculate the class_weights i use this code\nit my be interesting for you , and <strong>do not forget to upvote</strong>\n<em><strong>you may also fork my public kernel\n<a href=\"https://www.kaggle.com/hatemamine/flowertpuwin\">https://www.kaggle.com/hatemamine/flowertpuwin</a></strong></em>\n`\nlabels_train = training_dataset.map(lambda image, label: label).unbatch()\nT_correct_labels = next(iter(labels_train.batch(NUM_TRAINING_IMAGES))).numpy() </p>\n\n<h1>get everything as one batch</h1>\n\n<p>tmps=[]</p>\n\n<h1>for each labele calculate how much it occur</h1>\n\n<p>for x in range (len(CLASSES)):\n    tmp=0\n    for i in range(len(T_correct_labels)):\n        if x == T_correct_labels[i]:\n            tmp= tmp+1\n    tmps.append(tmp)</p>\n\n<p>tmps = np.array(tmps)\nprint(tmps.min())\nprint(tmps.max())\nprint(CLASSES[tmps.argmax()])\nprint(tmps.argmax())\nprint(CLASSES[tmps.argmin()])\nplt.plot(tmps) # plotting by columns\nplt.show()   </p>\n\n<h1>on keras you have to pass a dictionary to classweight parameter in model.fit()</h1>\n\n<h1>example classweights ={0:50, 1:0.5, 2:20}</h1>\n\n<p>class_weights = { i : ((tmps.max()+1)-tmps[i]) for i in range(0, len(tmps) ) }\nprint(class_weights)</p>\n\n<h1>but on TPU your have to pass flat list of weights (not a dictionary as written in the Keras docs)</h1>\n\n<h1>so you have just to convert your dictionary to a flat list</h1>\n\n<p>class_weights =[item for k in class_weights for item in (k, class_weights[k])]\nprint(class_weights)`</p>",
      "rawMarkdown": "you may have noticed that the **dataset is imbalanced** , and to deal with it you my use the **class_weight ** parameter in model.fit() \nmodel.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=3, **class_weight=class_weights**)\n\nit can be used to assigned different weights to different classes, to deal with the problem of imbalanced data , when you have more sample for one class then another (exemple you have 400 image of class 1 and 20 image of class 2 and 10 image of class3 etc )\n\nto calculate the class_weights i use this code\nit my be interesting for you , and **do not forget to upvote**\n***you may also fork my public kernel\n[https://www.kaggle.com/hatemamine/flowertpuwin](https://www.kaggle.com/hatemamine/flowertpuwin)***\n`\nlabels_train = training_dataset.map(lambda image, label: label).unbatch()\nT_correct_labels = next(iter(labels_train.batch(NUM_TRAINING_IMAGES))).numpy() \n# get everything as one batch\n\ntmps=[]\n# for each labele calculate how much it occur\nfor x in range (len(CLASSES)):\n    tmp=0\n    for i in range(len(T_correct_labels)):\n        if x == T_correct_labels[i]:\n            tmp= tmp+1\n    tmps.append(tmp)\n    \ntmps = np.array(tmps)\nprint(tmps.min())\nprint(tmps.max())\nprint(CLASSES[tmps.argmax()])\nprint(tmps.argmax())\nprint(CLASSES[tmps.argmin()])\nplt.plot(tmps) # plotting by columns\nplt.show()   \n\n#on keras you have to pass a dictionary to classweight parameter in model.fit() \n#example classweights ={0:50, 1:0.5, 2:20}\nclass_weights = { i : ((tmps.max()+1)-tmps[i]) for i in range(0, len(tmps) ) }\nprint(class_weights)\n\n#but on TPU your have to pass flat list of weights (not a dictionary as written in the Keras docs)\n#so you have just to convert your dictionary to a flat list \nclass_weights =[item for k in class_weights for item in (k, class_weights[k])]\nprint(class_weights)`",
      "votes": 8
    },
    {
      "id": 829712,
      "postDate": "2020-05-02T03:18:01.720Z",
      "content": "<p>i try adding class_weight, but it not improved..or improved very little</p>",
      "rawMarkdown": "i try adding class_weight, but it not improved..or improved very little",
      "replies": [
        {
          "id": 829720,
          "postDate": "2020-05-02T03:40:25.753Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 829640,
      "postDate": "2020-05-02T01:37:03.273Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 830259,
          "postDate": "2020-05-02T12:50:47.493Z",
          "content": "<p>I have tried and there is NO improve but you can try by yourself you may be lucky</p>",
          "rawMarkdown": "I have tried and there is NO improve but you can try by yourself you may be lucky"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 829712,
      "author_name": "astzls",
      "author_url": "",
      "post_date": "2020-05-02T03:18:01.720000",
      "content": "<p>i try adding class_weight, but it not improved..or improved very little</p>",
      "votes": 0,
      "replies": [
        {
          "id": 829720,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-02T03:40:25.753000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 829640,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-02T01:37:03.273000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 830259,
          "author_name": "ⵀⴰⵜⴻⵎ ⵎ'ⵀⴰⵎⴻⴷ ⴰⵎⵉⵏⴻ",
          "author_url": "",
          "post_date": "2020-05-02T12:50:47.493000",
          "content": "<p>I have tried and there is NO improve but you can try by yourself you may be lucky</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "801927": "you may have noticed that the **dataset is imbalanced** , and to deal with it you my use the **class_weight ** parameter in model.fit() \nmodel.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=3, **class_weight=class_weights**)\n\nit can be used to assigned different weights to different classes, to deal with the problem of imbalanced data , when you have more sample for one class then another (exemple you have 400 image of class 1 and 20 image of class 2 and 10 image of class3 etc )\n\nto calculate the class_weights i use this code\nit my be interesting for you , and **do not forget to upvote**\n***you may also fork my public kernel\n[https://www.kaggle.com/hatemamine/flowertpuwin](https://www.kaggle.com/hatemamine/flowertpuwin)***\n`\nlabels_train = training_dataset.map(lambda image, label: label).unbatch()\nT_correct_labels = next(iter(labels_train.batch(NUM_TRAINING_IMAGES))).numpy() \n# get everything as one batch\n\ntmps=[]\n# for each labele calculate how much it occur\nfor x in range (len(CLASSES)):\n    tmp=0\n    for i in range(len(T_correct_labels)):\n        if x == T_correct_labels[i]:\n            tmp= tmp+1\n    tmps.append(tmp)\n    \ntmps = np.array(tmps)\nprint(tmps.min())\nprint(tmps.max())\nprint(CLASSES[tmps.argmax()])\nprint(tmps.argmax())\nprint(CLASSES[tmps.argmin()])\nplt.plot(tmps) # plotting by columns\nplt.show()   \n\n#on keras you have to pass a dictionary to classweight parameter in model.fit() \n#example classweights ={0:50, 1:0.5, 2:20}\nclass_weights = { i : ((tmps.max()+1)-tmps[i]) for i in range(0, len(tmps) ) }\nprint(class_weights)\n\n#but on TPU your have to pass flat list of weights (not a dictionary as written in the Keras docs)\n#so you have just to convert your dictionary to a flat list \nclass_weights =[item for k in class_weights for item in (k, class_weights[k])]\nprint(class_weights)`",
    "829712": "i try adding class_weight, but it not improved..or improved very little",
    "829640": ""
  }
}