{
  "id": 172863,
  "title": "can any one tell me how can i balance the data-set with tfrec files ???",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172863",
  "author_name": "",
  "post_date": "2020-08-06T19:13:07.726018900Z",
  "votes": 4,
  "comment_count": 15,
  "views": 0,
  "content": "<p><code>\nGCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\nTRAINING_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nTRAINING_FILENAMES =TRAINING_FILENAMES[:14]\nVALIDATION_FILENAMES= np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nVALIDATION_FILENAMES=VALIDATION_FILENAMES[14:]\nTEST_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/test*.tfrec'))\nCLASSES = [0,1] <br>\n</code>\nas u can see i have the TRAINING_FILENAMES and the VALIDATION_FILENAMES but with tfrec files i dont really understand how can i balance the data set  </p>\n\n<p>any help will be very grateful </p>",
  "messages": [
    {
      "id": "960902",
      "postDate": "08/06/2020 19:13:07",
      "content": "<p><code>\nGCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\nTRAINING_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nTRAINING_FILENAMES =TRAINING_FILENAMES[:14]\nVALIDATION_FILENAMES= np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nVALIDATION_FILENAMES=VALIDATION_FILENAMES[14:]\nTEST_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/test*.tfrec'))\nCLASSES = [0,1] <br>\n</code>\nas u can see i have the TRAINING_FILENAMES and the VALIDATION_FILENAMES but with tfrec files i dont really understand how can i balance the data set  </p>\n\n<p>any help will be very grateful </p>",
      "rawMarkdown": "```\nGCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\nTRAINING_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nTRAINING_FILENAMES =TRAINING_FILENAMES[:14]\nVALIDATION_FILENAMES= np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nVALIDATION_FILENAMES=VALIDATION_FILENAMES[14:]\nTEST_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/test*.tfrec'))\nCLASSES = [0,1]   \n```\nas u can see i have the TRAINING_FILENAMES and the VALIDATION_FILENAMES but with tfrec files i dont really understand how can i balance the data set  \n\nany help will be very grateful",
      "votes": null
    },
    {
      "id": "960929",
      "postDate": "08/06/2020 19:48:17",
      "content": "<p>You should understand how these tfrec files were written. For example when we use TFRecz from <a href=\"/cdeotte\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">https://www.kaggle.com/cdeotte/melanoma-768x768</a>, we know that in every TFRec equal part of melanomas among all TFRecs and all patient's images in one TFRec. So taken this into account we can easily make Stratified Group KFold or make any divisions of Traininig Set.\nBut if you want to divide TrainSet exoticaly, for example to have in every TFRec equal parts of image resolution or place, then you need to write your own TFRec</p>",
      "rawMarkdown": "You should understand how these tfrec files were written. For example when we use TFRecz from @cdeotte https://www.kaggle.com/cdeotte/melanoma-768x768, we know that in every TFRec equal part of melanomas among all TFRecs and all patient's images in one TFRec. So taken this into account we can easily make Stratified Group KFold or make any divisions of Traininig Set.\nBut if you want to divide TrainSet exoticaly, for example to have in every TFRec equal parts of image resolution or place, then you need to write your own TFRec",
      "votes": null
    },
    {
      "id": "960960",
      "postDate": "08/06/2020 20:15:44",
      "content": "<p>What do you mean \"balance the dataset\"? The TFRecords are already balanced inside (with regard to malignant proportion). So what you have here is already balanced.</p>",
      "rawMarkdown": "What do you mean \"balance the dataset\"? The TFRecords are already balanced inside (with regard to malignant proportion). So what you have here is already balanced.",
      "votes": null
    },
    {
      "id": "960974",
      "postDate": "08/06/2020 20:31:56",
      "content": "<p>the data-set is imbalanced there are more normal cases then the melanoma cases </p>",
      "rawMarkdown": "the data-set is imbalanced there are more normal cases then the melanoma cases",
      "votes": null
    },
    {
      "id": "960989",
      "postDate": "08/06/2020 20:46:00",
      "content": "<p>i know how its works and how creat a tfrec \nthe problem is like i have accuracy with 0.9836 with valid_acc =0.9851 , but when i try to calculate  score, precision, recall i got , score  : 0.49, precision: 0.51, recall :0.50, which is incorrect ,so can any when tell me if there is any error in this code or if there is any other way to calculate score, precision, recall (<strong>to specify i want to calculate the sensitivity and specificity</strong>)</p>",
      "rawMarkdown": "i know how its works and how creat a tfrec \nthe problem is like i have accuracy with 0.9836 with valid_acc =0.9851 , but when i try to calculate  score, precision, recall i got , score  : 0.49, precision: 0.51, recall :0.50, which is incorrect ,so can any when tell me if there is any error in this code or if there is any other way to calculate score, precision, recall (**to specify i want to calculate the sensitivity and specificity**)",
      "votes": null
    },
    {
      "id": "960991",
      "postDate": "08/06/2020 20:46:52",
      "content": "<p>This could maybe work <a href=\"https://www.tensorflow.org/guide/data#resampling\">https://www.tensorflow.org/guide/data#resampling</a></p>",
      "rawMarkdown": "This could maybe work https://www.tensorflow.org/guide/data#resampling",
      "votes": null
    },
    {
      "id": "960992",
      "postDate": "08/06/2020 20:48:09",
      "content": "<p>Since the data is imbalanced as you mention, accuracy is a very bad metric. A naive model that just ALWAYS predicts negative will obtain similar metric scores as the ones you report.</p>",
      "rawMarkdown": "Since the data is imbalanced as you mention, accuracy is a very bad metric. A naive model that just ALWAYS predicts negative will obtain similar metric scores as the ones you report.",
      "votes": null
    },
    {
      "id": "960998",
      "postDate": "08/06/2020 20:52:02",
      "content": "<p>the data-set is imbalanced there are more normal cases then the melanoma cases  for this reaisent i got score : 0.49, precision: 0.51, recall :0.50 so how can i fix this issue </p>",
      "rawMarkdown": "the data-set is imbalanced there are more normal cases then the melanoma cases  for this reaisent i got score : 0.49, precision: 0.51, recall :0.50 so how can i fix this issue",
      "votes": null
    },
    {
      "id": "961000",
      "postDate": "08/06/2020 20:52:44",
      "content": "<p>thnaks bro ur really very helpful</p>",
      "rawMarkdown": "thnaks bro ur really very helpful",
      "votes": null
    },
    {
      "id": "961007",
      "postDate": "08/06/2020 20:59:03",
      "content": "<p>i will write an article im a student so i really need to calculate  the sensitivity and specificity ,my problem is the imbalanced data , as u know if the data is imbalanced the sensitivity and specificity will be very less </p>",
      "rawMarkdown": "i will write an article im a student so i really need to calculate  the sensitivity and specificity ,my problem is the imbalanced data , as u know if the data is imbalanced the sensitivity and specificity will be very less",
      "votes": null
    },
    {
      "id": "961011",
      "postDate": "08/06/2020 21:00:23",
      "content": "<p>Most welcome <a href=\"/tikoboss\">@tikoboss</a>! Glad that I could be of any help :)</p>",
      "rawMarkdown": "Most welcome @tikoboss! Glad that I could be of any help :)",
      "votes": null
    },
    {
      "id": "961015",
      "postDate": "08/06/2020 21:01:58",
      "content": "<p>There are many ways to go about fixing it. The real problem is in the fact that you use an objective function that maximizes the \"global\" error. Since 98% of the data is negative/healthy, the network starts focusing only on that data, as it can get the largest optimization there.</p>\n\n<p>You need to make more explicit to the network that you are interesting in the positive samples. This can be done by over-, under-sampling or balanced sampling, by using a custom loss function or by giving more weight in the loss to the positive samples (see, for example, <code>class_weights</code> for <code>keras.model.fit()</code>). There are definitely more possible solutions as well.</p>",
      "rawMarkdown": "There are many ways to go about fixing it. The real problem is in the fact that you use an objective function that maximizes the \"global\" error. Since 98% of the data is negative/healthy, the network starts focusing only on that data, as it can get the largest optimization there.\n\nYou need to make more explicit to the network that you are interesting in the positive samples. This can be done by over-, under-sampling or balanced sampling, by using a custom loss function or by giving more weight in the loss to the positive samples (see, for example, `class_weights` for `keras.model.fit()`). There are definitely more possible solutions as well.",
      "votes": null
    },
    {
      "id": "961020",
      "postDate": "08/06/2020 21:07:00",
      "content": "<p>sir just one question i know how to make  over-, under-sampling or balanced sampling ,the problem here is the tfrec , i know to work with over-, under-sampling  with images but with tfrec files i dont have any idea can u give me a suggetion or a notbook  that expalin how i will make this ,**thanks *</p>",
      "rawMarkdown": "sir just one question i know how to make  over-, under-sampling or balanced sampling ,the problem here is the tfrec , i know to work with over-, under-sampling  with images but with tfrec files i dont have any idea can u give me a suggetion or a notbook  that expalin how i will make this ,**thanks *",
      "votes": null
    },
    {
      "id": "961024",
      "postDate": "08/06/2020 21:12:53",
      "content": "<p>Ok. I explain how to upsample <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a>. I have made more TFRecords with just the malignant images. So if you want more malignant, you can add these additional TFRecords.</p>\n<p>My malignant TFRecords numbered 0-15 contain the same malignant from my 2020 TFRecords numbered 0-15. So in your case, you have this so far</p>\n<pre><code>GCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\nTRAINING_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nTRAINING_FILENAMES =TRAINING_FILENAMES[:14]\n</code></pre>\n<p>Now first remove <code>np.array()</code> from above like this</p>\n<pre><code>GCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\nTRAINING_FILENAMES =TRAINING_FILENAMES[:14]\n</code></pre>\n<p>Add more malignant with this</p>\n<pre><code>GCS_PATH2 = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\nMALIGNANT_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec'))\nTRAINING_FILENAMES += MALIGNANT_FILENAMES[:14]\n</code></pre>\n<p>Since you are training with <code>0, 1, 2, ..., 12, 13</code> you can add malignant TFRecords <code>0, 1, 2, ..., 12, 13</code> without causing leakage.</p>",
      "rawMarkdown": "Ok. I explain how to upsample [here][1]. I have made more TFRecords with just the malignant images. So if you want more malignant, you can add these additional TFRecords.\n\nMy malignant TFRecords numbered 0-15 contain the same malignant from my 2020 TFRecords numbered 0-15. So in your case, you have this so far\n\n    GCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\n    TRAINING_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\n    TRAINING_FILENAMES =TRAINING_FILENAMES[:14]\n\nNow first remove `np.array()` from above like this\n\n    GCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\n    TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n    TRAINING_FILENAMES =TRAINING_FILENAMES[:14]\n\nAdd more malignant with this\n\n    GCS_PATH2 = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\n    MALIGNANT_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec'))\n    TRAINING_FILENAMES += MALIGNANT_FILENAMES[:14]\n\nSince you are training with `0, 1, 2, ..., 12, 13` you can add malignant TFRecords `0, 1, 2, ..., 12, 13` without causing leakage.\n\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139",
      "votes": null
    },
    {
      "id": "961026",
      "postDate": "08/06/2020 21:15:56",
      "content": "<p>thank you sir ur very helpful , merci ,gracias </p>",
      "rawMarkdown": "thank you sir ur very helpful , merci ,gracias",
      "votes": null
    },
    {
      "id": "961029",
      "postDate": "08/06/2020 21:21:24",
      "content": "<p>No problem. I also posted last years competition data <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">here</a>. So you can include more data as follows:</p>\n<pre><code>GCS_PATH3 = KaggleDatasets().get_gcs_path('isic2019-384x384')\nMORE_FILENAMES = tf.io.gfile.glob(GCS_PATH3 + '/train*.tfrec'))\nTRAINING_FILENAMES += MORE_FILENAMES\n</code></pre>\n<p>And afterward, you can include extra malignant from last year's data too</p>\n<pre><code>GCS_PATH2 = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\nMALIGNANT_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec'))\nTRAINING_FILENAMES += MALIGNANT_FILENAMES[15:]\n</code></pre>\n<p>Note, since you have TFRecord 13 and 14 in your validation fold, don't add MALIGNANT_FILENAMES[13] nor [14]. Here was add <code>[15:]</code>.</p>\n<p>In total, my malignant dataset has 60 TFRecords containing 4000 malignant images from this year 2020, last years 2019 2018 2017 and scraped from ISIC-archive website!</p>",
      "rawMarkdown": "No problem. I also posted last years competition data [here][1]. So you can include more data as follows:\n\n    GCS_PATH3 = KaggleDatasets().get_gcs_path('isic2019-384x384')\n    MORE_FILENAMES = tf.io.gfile.glob(GCS_PATH3 + '/train*.tfrec'))\n    TRAINING_FILENAMES += MORE_FILENAMES\n\nAnd afterward, you can include extra malignant from last year's data too\n\n    GCS_PATH2 = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\n    MALIGNANT_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec'))\n    TRAINING_FILENAMES += MALIGNANT_FILENAMES[15:]\n\nNote, since you have TFRecord 13 and 14 in your validation fold, don't add MALIGNANT_FILENAMES[13] nor [14]. Here was add `[15:]`.\n\nIn total, my malignant dataset has 60 TFRecords containing 4000 malignant images from this year 2020, last years 2019 2018 2017 and scraped from ISIC-archive website!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 960960,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/06/2020 20:15:44",
      "content": "<p>What do you mean \"balance the dataset\"? The TFRecords are already balanced inside (with regard to malignant proportion). So what you have here is already balanced.</p>",
      "votes": null,
      "replies": [
        {
          "id": 960974,
          "author_name": "tikoboss",
          "author_url": "",
          "post_date": "08/06/2020 20:31:56",
          "content": "<p>the data-set is imbalanced there are more normal cases then the melanoma cases </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961024,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/06/2020 21:12:53",
          "content": "<p>Ok. I explain how to upsample <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a>. I have made more TFRecords with just the malignant images. So if you want more malignant, you can add these additional TFRecords.</p>\n<p>My malignant TFRecords numbered 0-15 contain the same malignant from my 2020 TFRecords numbered 0-15. So in your case, you have this so far</p>\n<pre><code>GCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\nTRAINING_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nTRAINING_FILENAMES =TRAINING_FILENAMES[:14]\n</code></pre>\n<p>Now first remove <code>np.array()</code> from above like this</p>\n<pre><code>GCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\nTRAINING_FILENAMES =TRAINING_FILENAMES[:14]\n</code></pre>\n<p>Add more malignant with this</p>\n<pre><code>GCS_PATH2 = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\nMALIGNANT_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec'))\nTRAINING_FILENAMES += MALIGNANT_FILENAMES[:14]\n</code></pre>\n<p>Since you are training with <code>0, 1, 2, ..., 12, 13</code> you can add malignant TFRecords <code>0, 1, 2, ..., 12, 13</code> without causing leakage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961026,
          "author_name": "tikoboss",
          "author_url": "",
          "post_date": "08/06/2020 21:15:56",
          "content": "<p>thank you sir ur very helpful , merci ,gracias </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961029,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/06/2020 21:21:24",
          "content": "<p>No problem. I also posted last years competition data <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">here</a>. So you can include more data as follows:</p>\n<pre><code>GCS_PATH3 = KaggleDatasets().get_gcs_path('isic2019-384x384')\nMORE_FILENAMES = tf.io.gfile.glob(GCS_PATH3 + '/train*.tfrec'))\nTRAINING_FILENAMES += MORE_FILENAMES\n</code></pre>\n<p>And afterward, you can include extra malignant from last year's data too</p>\n<pre><code>GCS_PATH2 = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\nMALIGNANT_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec'))\nTRAINING_FILENAMES += MALIGNANT_FILENAMES[15:]\n</code></pre>\n<p>Note, since you have TFRecord 13 and 14 in your validation fold, don't add MALIGNANT_FILENAMES[13] nor [14]. Here was add <code>[15:]</code>.</p>\n<p>In total, my malignant dataset has 60 TFRecords containing 4000 malignant images from this year 2020, last years 2019 2018 2017 and scraped from ISIC-archive website!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 960929,
      "author_name": "aybatov",
      "author_url": "",
      "post_date": "08/06/2020 19:48:17",
      "content": "<p>You should understand how these tfrec files were written. For example when we use TFRecz from <a href=\"/cdeotte\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">https://www.kaggle.com/cdeotte/melanoma-768x768</a>, we know that in every TFRec equal part of melanomas among all TFRecs and all patient's images in one TFRec. So taken this into account we can easily make Stratified Group KFold or make any divisions of Traininig Set.\nBut if you want to divide TrainSet exoticaly, for example to have in every TFRec equal parts of image resolution or place, then you need to write your own TFRec</p>",
      "votes": null,
      "replies": [
        {
          "id": 960989,
          "author_name": "tikoboss",
          "author_url": "",
          "post_date": "08/06/2020 20:46:00",
          "content": "<p>i know how its works and how creat a tfrec \nthe problem is like i have accuracy with 0.9836 with valid_acc =0.9851 , but when i try to calculate  score, precision, recall i got , score  : 0.49, precision: 0.51, recall :0.50, which is incorrect ,so can any when tell me if there is any error in this code or if there is any other way to calculate score, precision, recall (<strong>to specify i want to calculate the sensitivity and specificity</strong>)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960992,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/06/2020 20:48:09",
          "content": "<p>Since the data is imbalanced as you mention, accuracy is a very bad metric. A naive model that just ALWAYS predicts negative will obtain similar metric scores as the ones you report.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960998,
          "author_name": "tikoboss",
          "author_url": "",
          "post_date": "08/06/2020 20:52:02",
          "content": "<p>the data-set is imbalanced there are more normal cases then the melanoma cases  for this reaisent i got score : 0.49, precision: 0.51, recall :0.50 so how can i fix this issue </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961007,
          "author_name": "tikoboss",
          "author_url": "",
          "post_date": "08/06/2020 20:59:03",
          "content": "<p>i will write an article im a student so i really need to calculate  the sensitivity and specificity ,my problem is the imbalanced data , as u know if the data is imbalanced the sensitivity and specificity will be very less </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961015,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/06/2020 21:01:58",
          "content": "<p>There are many ways to go about fixing it. The real problem is in the fact that you use an objective function that maximizes the \"global\" error. Since 98% of the data is negative/healthy, the network starts focusing only on that data, as it can get the largest optimization there.</p>\n\n<p>You need to make more explicit to the network that you are interesting in the positive samples. This can be done by over-, under-sampling or balanced sampling, by using a custom loss function or by giving more weight in the loss to the positive samples (see, for example, <code>class_weights</code> for <code>keras.model.fit()</code>). There are definitely more possible solutions as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961020,
          "author_name": "tikoboss",
          "author_url": "",
          "post_date": "08/06/2020 21:07:00",
          "content": "<p>sir just one question i know how to make  over-, under-sampling or balanced sampling ,the problem here is the tfrec , i know to work with over-, under-sampling  with images but with tfrec files i dont have any idea can u give me a suggetion or a notbook  that expalin how i will make this ,**thanks *</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 960991,
      "author_name": "group16",
      "author_url": "",
      "post_date": "08/06/2020 20:46:52",
      "content": "<p>This could maybe work <a href=\"https://www.tensorflow.org/guide/data#resampling\">https://www.tensorflow.org/guide/data#resampling</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 961000,
          "author_name": "tikoboss",
          "author_url": "",
          "post_date": "08/06/2020 20:52:44",
          "content": "<p>thnaks bro ur really very helpful</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961011,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/06/2020 21:00:23",
          "content": "<p>Most welcome <a href=\"/tikoboss\">@tikoboss</a>! Glad that I could be of any help :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "960902": "```\nGCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\nTRAINING_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nTRAINING_FILENAMES =TRAINING_FILENAMES[:14]\nVALIDATION_FILENAMES= np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\nVALIDATION_FILENAMES=VALIDATION_FILENAMES[14:]\nTEST_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/test*.tfrec'))\nCLASSES = [0,1]   \n```\nas u can see i have the TRAINING_FILENAMES and the VALIDATION_FILENAMES but with tfrec files i dont really understand how can i balance the data set  \n\nany help will be very grateful",
    "960929": "You should understand how these tfrec files were written. For example when we use TFRecz from @cdeotte https://www.kaggle.com/cdeotte/melanoma-768x768, we know that in every TFRec equal part of melanomas among all TFRecs and all patient's images in one TFRec. So taken this into account we can easily make Stratified Group KFold or make any divisions of Traininig Set.\nBut if you want to divide TrainSet exoticaly, for example to have in every TFRec equal parts of image resolution or place, then you need to write your own TFRec",
    "960960": "What do you mean \"balance the dataset\"? The TFRecords are already balanced inside (with regard to malignant proportion). So what you have here is already balanced.",
    "960974": "the data-set is imbalanced there are more normal cases then the melanoma cases",
    "960989": "i know how its works and how creat a tfrec \nthe problem is like i have accuracy with 0.9836 with valid_acc =0.9851 , but when i try to calculate  score, precision, recall i got , score  : 0.49, precision: 0.51, recall :0.50, which is incorrect ,so can any when tell me if there is any error in this code or if there is any other way to calculate score, precision, recall (**to specify i want to calculate the sensitivity and specificity**)",
    "960991": "This could maybe work https://www.tensorflow.org/guide/data#resampling",
    "960992": "Since the data is imbalanced as you mention, accuracy is a very bad metric. A naive model that just ALWAYS predicts negative will obtain similar metric scores as the ones you report.",
    "960998": "the data-set is imbalanced there are more normal cases then the melanoma cases  for this reaisent i got score : 0.49, precision: 0.51, recall :0.50 so how can i fix this issue",
    "961000": "thnaks bro ur really very helpful",
    "961007": "i will write an article im a student so i really need to calculate  the sensitivity and specificity ,my problem is the imbalanced data , as u know if the data is imbalanced the sensitivity and specificity will be very less",
    "961011": "Most welcome @tikoboss! Glad that I could be of any help :)",
    "961015": "There are many ways to go about fixing it. The real problem is in the fact that you use an objective function that maximizes the \"global\" error. Since 98% of the data is negative/healthy, the network starts focusing only on that data, as it can get the largest optimization there.\n\nYou need to make more explicit to the network that you are interesting in the positive samples. This can be done by over-, under-sampling or balanced sampling, by using a custom loss function or by giving more weight in the loss to the positive samples (see, for example, `class_weights` for `keras.model.fit()`). There are definitely more possible solutions as well.",
    "961020": "sir just one question i know how to make  over-, under-sampling or balanced sampling ,the problem here is the tfrec , i know to work with over-, under-sampling  with images but with tfrec files i dont have any idea can u give me a suggetion or a notbook  that expalin how i will make this ,**thanks *",
    "961024": "Ok. I explain how to upsample [here][1]. I have made more TFRecords with just the malignant images. So if you want more malignant, you can add these additional TFRecords.\n\nMy malignant TFRecords numbered 0-15 contain the same malignant from my 2020 TFRecords numbered 0-15. So in your case, you have this so far\n\n    GCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\n    TRAINING_FILENAMES = np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec'))\n    TRAINING_FILENAMES =TRAINING_FILENAMES[:14]\n\nNow first remove `np.array()` from above like this\n\n    GCS_PATH = KaggleDatasets().get_gcs_path('melanoma-384x384')\n    TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n    TRAINING_FILENAMES =TRAINING_FILENAMES[:14]\n\nAdd more malignant with this\n\n    GCS_PATH2 = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\n    MALIGNANT_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec'))\n    TRAINING_FILENAMES += MALIGNANT_FILENAMES[:14]\n\nSince you are training with `0, 1, 2, ..., 12, 13` you can add malignant TFRecords `0, 1, 2, ..., 12, 13` without causing leakage.\n\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139",
    "961026": "thank you sir ur very helpful , merci ,gracias",
    "961029": "No problem. I also posted last years competition data [here][1]. So you can include more data as follows:\n\n    GCS_PATH3 = KaggleDatasets().get_gcs_path('isic2019-384x384')\n    MORE_FILENAMES = tf.io.gfile.glob(GCS_PATH3 + '/train*.tfrec'))\n    TRAINING_FILENAMES += MORE_FILENAMES\n\nAnd afterward, you can include extra malignant from last year's data too\n\n    GCS_PATH2 = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\n    MALIGNANT_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec'))\n    TRAINING_FILENAMES += MALIGNANT_FILENAMES[15:]\n\nNote, since you have TFRecord 13 and 14 in your validation fold, don't add MALIGNANT_FILENAMES[13] nor [14]. Here was add `[15:]`.\n\nIn total, my malignant dataset has 60 TFRecords containing 4000 malignant images from this year 2020, last years 2019 2018 2017 and scraped from ISIC-archive website!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910"
  },
  "source": "meta"
}