{
  "id": 155317,
  "title": "Unbalanced classes and CNN",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155317",
  "author_name": "Gabi",
  "post_date": "2020-06-01T07:46:21.073000",
  "votes": 0,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My first attemp was to make a CNN, it's VERY slow, and I trained it only 3 epochs, but after training (with a high accuracy but this means nothing with such an unbalanced set, I know) if I try to predict, all images range fom 0.42 to 0.47 so foir a threshold of 0.5 all images are predicted as \"benign\"</p>\n\n<p>I believe it's a problem with such unbalanced data, because I've applied CNN 3 times so far (I'm a newbie) and first time it worked very well, 2nd time in a very unbalanced problem it did'nt worked and now with this unbalanced problem too it does no seems to work very well again.</p>\n\n<p>Any suggestion about the use of CNN with highly unbalanced data?\nIs using a CNN a correct technique for this problem?</p>\n\n<p>I'm not expecting a great result on my first try but at least some \"malign\" detected :-p</p>\n\n<p>Code used: I'm not linking a notebook because the train was done in my home computer (it was more than 6 hours...) but basically it's like:</p>\n\n<p>train_datagen = ImageDataGenerator(\n    rescale = 1./255,\n    zoom_range = 0.2,\n    horizontal_flip=True)\nval_datagen=train_datagen = ImageDataGenerator(rescale = 1./255)\ntest_datagen = ImageDataGenerator(rescale = 1./255)\ntrain_generator = train_datagen.flow_from_dataframe(\n    X_train,\n    x_col= 'image_fullpath',\n    y_col='benign_malignant',\n    target_size=(224, 224),\n    class_mode='binary',\n    batch_size=32\n)\nvalidation_generator = val_datagen.flow_from_dataframe(\n    X_val,\n    x_col='image_fullpath',\n    y_col='benign_malignant',\n    target_size=(224, 224),\n    class_mode='binary',\n    batch_size=32\n)\ntest_generatorFULL = test_datagen.flow_from_dataframe(\n    test,\n    x_col='image_fullpath',\n    y_col=None,\n    target_size=(224, 224),\n    class_mode=None\n )</p>\n\n<h1>Creation of the CNN</h1>\n\n<p>classifier = Sequential() <br>\nclassifier.add(Conv2D(32, (3, 3), input_shape = (224, 224, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Conv2D(32, (3, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Conv2D(32, (3, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Flatten())\nclassifier.add(Dense(units = 64, activation = 'relu'))\nclassifier.add(Dropout(0.4))\nclassifier.add(Dense(units = 1, activation = 'sigmoid'))\nclassifier.compile(optimizer = 'adam', loss = 'binary_crossentropy', metrics = ['accuracy'])\ntrain_size = X_train.shape[0]/32\nvalidation_size = X_val.shape[0]/32\nepochs = 3\nhistory = classifier.fit_generator(train_generator,\n                         steps_per_epoch = train_size, \n                         epochs = epochs,\n                         validation_data = validation_generator,\n                         validation_steps = validation_size)\nprediction = classifier.predict(test_generatorFULL)</p>",
  "messages": [
    {
      "id": 869684,
      "postDate": "2020-06-01T07:46:21.073Z",
      "content": "<p>My first attemp was to make a CNN, it's VERY slow, and I trained it only 3 epochs, but after training (with a high accuracy but this means nothing with such an unbalanced set, I know) if I try to predict, all images range fom 0.42 to 0.47 so foir a threshold of 0.5 all images are predicted as \"benign\"</p>\n\n<p>I believe it's a problem with such unbalanced data, because I've applied CNN 3 times so far (I'm a newbie) and first time it worked very well, 2nd time in a very unbalanced problem it did'nt worked and now with this unbalanced problem too it does no seems to work very well again.</p>\n\n<p>Any suggestion about the use of CNN with highly unbalanced data?\nIs using a CNN a correct technique for this problem?</p>\n\n<p>I'm not expecting a great result on my first try but at least some \"malign\" detected :-p</p>\n\n<p>Code used: I'm not linking a notebook because the train was done in my home computer (it was more than 6 hours...) but basically it's like:</p>\n\n<p>train_datagen = ImageDataGenerator(\n    rescale = 1./255,\n    zoom_range = 0.2,\n    horizontal_flip=True)\nval_datagen=train_datagen = ImageDataGenerator(rescale = 1./255)\ntest_datagen = ImageDataGenerator(rescale = 1./255)\ntrain_generator = train_datagen.flow_from_dataframe(\n    X_train,\n    x_col= 'image_fullpath',\n    y_col='benign_malignant',\n    target_size=(224, 224),\n    class_mode='binary',\n    batch_size=32\n)\nvalidation_generator = val_datagen.flow_from_dataframe(\n    X_val,\n    x_col='image_fullpath',\n    y_col='benign_malignant',\n    target_size=(224, 224),\n    class_mode='binary',\n    batch_size=32\n)\ntest_generatorFULL = test_datagen.flow_from_dataframe(\n    test,\n    x_col='image_fullpath',\n    y_col=None,\n    target_size=(224, 224),\n    class_mode=None\n )</p>\n\n<h1>Creation of the CNN</h1>\n\n<p>classifier = Sequential() <br>\nclassifier.add(Conv2D(32, (3, 3), input_shape = (224, 224, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Conv2D(32, (3, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Conv2D(32, (3, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Flatten())\nclassifier.add(Dense(units = 64, activation = 'relu'))\nclassifier.add(Dropout(0.4))\nclassifier.add(Dense(units = 1, activation = 'sigmoid'))\nclassifier.compile(optimizer = 'adam', loss = 'binary_crossentropy', metrics = ['accuracy'])\ntrain_size = X_train.shape[0]/32\nvalidation_size = X_val.shape[0]/32\nepochs = 3\nhistory = classifier.fit_generator(train_generator,\n                         steps_per_epoch = train_size, \n                         epochs = epochs,\n                         validation_data = validation_generator,\n                         validation_steps = validation_size)\nprediction = classifier.predict(test_generatorFULL)</p>",
      "rawMarkdown": "My first attemp was to make a CNN, it's VERY slow, and I trained it only 3 epochs, but after training (with a high accuracy but this means nothing with such an unbalanced set, I know) if I try to predict, all images range fom 0.42 to 0.47 so foir a threshold of 0.5 all images are predicted as \"benign\"\n\nI believe it's a problem with such unbalanced data, because I've applied CNN 3 times so far (I'm a newbie) and first time it worked very well, 2nd time in a very unbalanced problem it did'nt worked and now with this unbalanced problem too it does no seems to work very well again.\n\nAny suggestion about the use of CNN with highly unbalanced data?\nIs using a CNN a correct technique for this problem?\n\nI'm not expecting a great result on my first try but at least some \"malign\" detected :-p\n\nCode used: I'm not linking a notebook because the train was done in my home computer (it was more than 6 hours...) but basically it's like:\n\ntrain_datagen = ImageDataGenerator(\n    rescale = 1./255,\n    zoom_range = 0.2,\n    horizontal_flip=True)\nval_datagen=train_datagen = ImageDataGenerator(rescale = 1./255)\ntest_datagen = ImageDataGenerator(rescale = 1./255)\ntrain_generator = train_datagen.flow_from_dataframe(\n    X_train,\n    x_col= 'image_fullpath',\n    y_col='benign_malignant',\n    target_size=(224, 224),\n    class_mode='binary',\n    batch_size=32\n)\nvalidation_generator = val_datagen.flow_from_dataframe(\n    X_val,\n    x_col='image_fullpath',\n    y_col='benign_malignant',\n    target_size=(224, 224),\n    class_mode='binary',\n    batch_size=32\n)\ntest_generatorFULL = test_datagen.flow_from_dataframe(\n    test,\n    x_col='image_fullpath',\n    y_col=None,\n    target_size=(224, 224),\n    class_mode=None\n )\n#Creation of the CNN\nclassifier = Sequential()                                                                           \nclassifier.add(Conv2D(32, (3, 3), input_shape = (224, 224, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Conv2D(32, (3, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Conv2D(32, (3, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Flatten())\nclassifier.add(Dense(units = 64, activation = 'relu'))\nclassifier.add(Dropout(0.4))\nclassifier.add(Dense(units = 1, activation = 'sigmoid'))\nclassifier.compile(optimizer = 'adam', loss = 'binary_crossentropy', metrics = ['accuracy'])\ntrain_size = X_train.shape[0]/32\nvalidation_size = X_val.shape[0]/32\nepochs = 3\nhistory = classifier.fit_generator(train_generator,\n                         steps_per_epoch = train_size, \n                         epochs = epochs,\n                         validation_data = validation_generator,\n                         validation_steps = validation_size)\nprediction = classifier.predict(test_generatorFULL)"
    },
    {
      "id": 869847,
      "postDate": "2020-06-01T10:13:07.753Z",
      "content": "<p>Instead of training a CNN from scratch try using a pre-trained model. As a baseline,  I created a small balanced dataset with 1000 images and used a pre-trained model. I found out that DenseNet works well with this small dataset (Leader board score 0.738 after 5 epochs).\nWhile using the entire dataset you can try oversampling the data with minority class to make it balanced and use the metadata from the CSV files\nBTW I'm a newbie too :)</p>",
      "rawMarkdown": "Instead of training a CNN from scratch try using a pre-trained model. As a baseline,  I created a small balanced dataset with 1000 images and used a pre-trained model. I found out that DenseNet works well with this small dataset (Leader board score 0.738 after 5 epochs).\nWhile using the entire dataset you can try oversampling the data with minority class to make it balanced and use the metadata from the CSV files\nBTW I'm a newbie too :)",
      "replies": [
        {
          "id": 873474,
          "postDate": "2020-06-04T07:17:07.557Z",
          "content": "<p>Thanks Safiuddin, I tried to use a smaller balanced dataset removing proportionally data from train and validation set, I want to try with different NN like DenseNet but as far I only know the theory behing CNN generic networks I'm waiting until have time to read about another networks like DenseNet and others, and not making cut&amp;paste without knowing them, this is the reason I stil have not tried other networks.\nI'll try to have some time the weekend to open the notebook again and doing some work.\nIn the link below there is a very interesting discussion about how to deal with unbalanced classes not so simple as my first approximation \"remove in a proportional way data from training to go from 98%-2% (or similar) proportion to 85%-15% or similar...\"</p>",
          "rawMarkdown": "Thanks Safiuddin, I tried to use a smaller balanced dataset removing proportionally data from train and validation set, I want to try with different NN like DenseNet but as far I only know the theory behing CNN generic networks I'm waiting until have time to read about another networks like DenseNet and others, and not making cut&amp;paste without knowing them, this is the reason I stil have not tried other networks.\nI'll try to have some time the weekend to open the notebook again and doing some work.\nIn the link below there is a very interesting discussion about how to deal with unbalanced classes not so simple as my first approximation \"remove in a proportional way data from training to go from 98%-2% (or similar) proportion to 85%-15% or similar...\"\n"
        }
      ]
    },
    {
      "id": 869829,
      "postDate": "2020-06-01T10:04:27.623Z",
      "content": "<p>Sorry I found in the same discussion forum this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154791\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154791</a></p>\n\n<p>But the question about, is it \"normal\" to get 0 samples of test as maliagnt? or I'm missing something else??? something clearly wrong in my model/code?\nAapart of trying the suggestions on that thread...</p>",
      "rawMarkdown": "Sorry I found in the same discussion forum this https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154791\n\nBut the question about, is it \"normal\" to get 0 samples of test as maliagnt? or I'm missing something else??? something clearly wrong in my model/code?\nAapart of trying the suggestions on that thread...",
      "replies": [
        {
          "id": 886557,
          "postDate": "2020-06-15T06:04:36.017Z",
          "content": "<p>Answer to myself to not confuse anyone reading this thread, I was wrong the submit has the probabilities not the target classification 0 or 1, so this part of my comment had no sense.</p>",
          "rawMarkdown": "Answer to myself to not confuse anyone reading this thread, I was wrong the submit has the probabilities not the target classification 0 or 1, so this part of my comment had no sense."
        }
      ]
    },
    {
      "id": 869842,
      "postDate": "2020-06-01T10:12:05.493Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 869847,
      "author_name": "Safiuddin Mohammad",
      "author_url": "",
      "post_date": "2020-06-01T10:13:07.753000",
      "content": "<p>Instead of training a CNN from scratch try using a pre-trained model. As a baseline,  I created a small balanced dataset with 1000 images and used a pre-trained model. I found out that DenseNet works well with this small dataset (Leader board score 0.738 after 5 epochs).\nWhile using the entire dataset you can try oversampling the data with minority class to make it balanced and use the metadata from the CSV files\nBTW I'm a newbie too :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 873474,
          "author_name": "Gabi",
          "author_url": "",
          "post_date": "2020-06-04T07:17:07.557000",
          "content": "<p>Thanks Safiuddin, I tried to use a smaller balanced dataset removing proportionally data from train and validation set, I want to try with different NN like DenseNet but as far I only know the theory behing CNN generic networks I'm waiting until have time to read about another networks like DenseNet and others, and not making cut&amp;paste without knowing them, this is the reason I stil have not tried other networks.\nI'll try to have some time the weekend to open the notebook again and doing some work.\nIn the link below there is a very interesting discussion about how to deal with unbalanced classes not so simple as my first approximation \"remove in a proportional way data from training to go from 98%-2% (or similar) proportion to 85%-15% or similar...\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 869829,
      "author_name": "Gabi",
      "author_url": "",
      "post_date": "2020-06-01T10:04:27.623000",
      "content": "<p>Sorry I found in the same discussion forum this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154791\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154791</a></p>\n\n<p>But the question about, is it \"normal\" to get 0 samples of test as maliagnt? or I'm missing something else??? something clearly wrong in my model/code?\nAapart of trying the suggestions on that thread...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 886557,
          "author_name": "Gabi",
          "author_url": "",
          "post_date": "2020-06-15T06:04:36.017000",
          "content": "<p>Answer to myself to not confuse anyone reading this thread, I was wrong the submit has the probabilities not the target classification 0 or 1, so this part of my comment had no sense.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 869842,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-01T10:12:05.493000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "869684": "My first attemp was to make a CNN, it's VERY slow, and I trained it only 3 epochs, but after training (with a high accuracy but this means nothing with such an unbalanced set, I know) if I try to predict, all images range fom 0.42 to 0.47 so foir a threshold of 0.5 all images are predicted as \"benign\"\n\nI believe it's a problem with such unbalanced data, because I've applied CNN 3 times so far (I'm a newbie) and first time it worked very well, 2nd time in a very unbalanced problem it did'nt worked and now with this unbalanced problem too it does no seems to work very well again.\n\nAny suggestion about the use of CNN with highly unbalanced data?\nIs using a CNN a correct technique for this problem?\n\nI'm not expecting a great result on my first try but at least some \"malign\" detected :-p\n\nCode used: I'm not linking a notebook because the train was done in my home computer (it was more than 6 hours...) but basically it's like:\n\ntrain_datagen = ImageDataGenerator(\n    rescale = 1./255,\n    zoom_range = 0.2,\n    horizontal_flip=True)\nval_datagen=train_datagen = ImageDataGenerator(rescale = 1./255)\ntest_datagen = ImageDataGenerator(rescale = 1./255)\ntrain_generator = train_datagen.flow_from_dataframe(\n    X_train,\n    x_col= 'image_fullpath',\n    y_col='benign_malignant',\n    target_size=(224, 224),\n    class_mode='binary',\n    batch_size=32\n)\nvalidation_generator = val_datagen.flow_from_dataframe(\n    X_val,\n    x_col='image_fullpath',\n    y_col='benign_malignant',\n    target_size=(224, 224),\n    class_mode='binary',\n    batch_size=32\n)\ntest_generatorFULL = test_datagen.flow_from_dataframe(\n    test,\n    x_col='image_fullpath',\n    y_col=None,\n    target_size=(224, 224),\n    class_mode=None\n )\n#Creation of the CNN\nclassifier = Sequential()                                                                           \nclassifier.add(Conv2D(32, (3, 3), input_shape = (224, 224, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Conv2D(32, (3, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Conv2D(32, (3, 3), activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Flatten())\nclassifier.add(Dense(units = 64, activation = 'relu'))\nclassifier.add(Dropout(0.4))\nclassifier.add(Dense(units = 1, activation = 'sigmoid'))\nclassifier.compile(optimizer = 'adam', loss = 'binary_crossentropy', metrics = ['accuracy'])\ntrain_size = X_train.shape[0]/32\nvalidation_size = X_val.shape[0]/32\nepochs = 3\nhistory = classifier.fit_generator(train_generator,\n                         steps_per_epoch = train_size, \n                         epochs = epochs,\n                         validation_data = validation_generator,\n                         validation_steps = validation_size)\nprediction = classifier.predict(test_generatorFULL)",
    "869847": "Instead of training a CNN from scratch try using a pre-trained model. As a baseline,  I created a small balanced dataset with 1000 images and used a pre-trained model. I found out that DenseNet works well with this small dataset (Leader board score 0.738 after 5 epochs).\nWhile using the entire dataset you can try oversampling the data with minority class to make it balanced and use the metadata from the CSV files\nBTW I'm a newbie too :)",
    "869829": "Sorry I found in the same discussion forum this https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154791\n\nBut the question about, is it \"normal\" to get 0 samples of test as maliagnt? or I'm missing something else??? something clearly wrong in my model/code?\nAapart of trying the suggestions on that thread...",
    "869842": ""
  }
}