{
  "id": 164524,
  "title": "Help on AUC score Value Error",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164524",
  "author_name": "",
  "post_date": "2020-07-06T16:00:03.443521600Z",
  "votes": null,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Sometimes the true value is all 0 in a batch and rocaucscore gives a <code>ValueError: Only one class present in y_true. ROC AUC score is not defined in that case.</code> \nWhat is the best way to solve this problem? - I am getting this error while calculating auc score for validation dataset. </p>",
  "messages": [
    {
      "id": "917571",
      "postDate": "07/06/2020 16:00:03",
      "content": "<p>Sometimes the true value is all 0 in a batch and rocaucscore gives a <code>ValueError: Only one class present in y_true. ROC AUC score is not defined in that case.</code> \nWhat is the best way to solve this problem? - I am getting this error while calculating auc score for validation dataset. </p>",
      "rawMarkdown": "Sometimes the true value is all 0 in a batch and rocaucscore gives a `ValueError: Only one class present in y_true. ROC AUC score is not defined in that case.` \nWhat is the best way to solve this problem? - I am getting this error while calculating auc score for validation dataset.",
      "votes": null
    },
    {
      "id": "917585",
      "postDate": "07/06/2020 16:11:21",
      "content": "<p>Hi, do you have the code in a public notebook so I can take a look?</p>",
      "rawMarkdown": "Hi, do you have the code in a public notebook so I can take a look?",
      "votes": null
    },
    {
      "id": "917608",
      "postDate": "07/06/2020 16:23:43",
      "content": "<p>That means all target values are \"0\" (only one class); somehow your selected fold (validation) does not contain any Malignant images.</p>",
      "rawMarkdown": "That means all target values are \"0\" (only one class); somehow your selected fold (validation) does not contain any Malignant images.",
      "votes": null
    },
    {
      "id": "917614",
      "postDate": "07/06/2020 16:28:32",
      "content": "<p>I would recommend you to consider making your validation data set more diverse. IMHO, it is not a good idea to have only one class represented in either of your validation folds. Read, for example, this discussion topic about building a reasonable validation strategy:\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156002\">Implementation of Stratified Group K-Folds</a>\nAnd here is the public kernel illustrating the implementation of this strategy:\n<a href=\"https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds\">SIIM Stratified GroupKFold 5-folds</a></p>",
      "rawMarkdown": "I would recommend you to consider making your validation data set more diverse. IMHO, it is not a good idea to have only one class represented in either of your validation folds. Read, for example, this discussion topic about building a reasonable validation strategy:\n[Implementation of Stratified Group K-Folds](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156002)\nAnd here is the public kernel illustrating the implementation of this strategy:\n[SIIM Stratified GroupKFold 5-folds](https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds)",
      "votes": null
    },
    {
      "id": "917646",
      "postDate": "07/06/2020 17:01:10",
      "content": "<p>Instead of using K-fold CV, I have just performed a single train-test (stratified) split by maintaining the same class ratio in both the train &amp; valid sets and then I have created train &amp; validation data loaders with help of <code>DataLoader</code> function. While calculating auc score on validation dataset, on some of the batches I am facing the mentioned value  Error. Are there any strategies to perform a better train-test split for this kind of imbalance problem. ?</p>\n\n<p>Anyhow, I will try to implement stratified k-fold as per the mentioned notebooks.  Thanks for the response</p>",
      "rawMarkdown": "Instead of using K-fold CV, I have just performed a single train-test (stratified) split by maintaining the same class ratio in both the train &amp; valid sets and then I have created train &amp; validation data loaders with help of `DataLoader` function. While calculating auc score on validation dataset, on some of the batches I am facing the mentioned value  Error. Are there any strategies to perform a better train-test split for this kind of imbalance problem. ?\n\nAnyhow, I will try to implement stratified k-fold as per the mentioned notebooks.  Thanks for the response",
      "votes": null
    },
    {
      "id": "917647",
      "postDate": "07/06/2020 17:01:54",
      "content": "<p>Hey, You can find my notebook at <a href=\"https://jovian.ml/vineel369/class-project-skin-cancer\">https://jovian.ml/vineel369/class-project-skin-cancer</a>. </p>\n\n<p>Please provide me your feedback on it</p>",
      "rawMarkdown": "Hey, You can find my notebook at https://jovian.ml/vineel369/class-project-skin-cancer. \n\nPlease provide me your feedback on it",
      "votes": null
    },
    {
      "id": "917651",
      "postDate": "07/06/2020 17:05:54",
      "content": "<p>Hi Sirish, </p>\n\n<p>You are correct...  I have performed a single train-test (stratified) split by maintaining the same class ratio in both the train &amp; valid sets and I have checked the same. \n<code>\nTrain set : ~98% (0 - class) &amp; ~2 % (1 - class)\nValidation set : ~98% (0 - class) &amp; ~2 % (1 - class)\n</code>\nI have created train &amp; validation data loaders with help of DataLoader function. \nWhile calculating auc score on validation dataset, on some of the batches I am facing the mentioned value Error. Are there any strategies to perform a better train-test split for this kind of imbalance problem. ?</p>",
      "rawMarkdown": "Hi Sirish, \n\nYou are correct...  I have performed a single train-test (stratified) split by maintaining the same class ratio in both the train &amp; valid sets and I have checked the same. \n```\nTrain set : ~98% (0 - class) &amp; ~2 % (1 - class)\nValidation set : ~98% (0 - class) &amp; ~2 % (1 - class)\n```\nI have created train &amp; validation data loaders with help of DataLoader function. \nWhile calculating auc score on validation dataset, on some of the batches I am facing the mentioned value Error. Are there any strategies to perform a better train-test split for this kind of imbalance problem. ?",
      "votes": null
    },
    {
      "id": "917658",
      "postDate": "07/06/2020 17:09:56",
      "content": "<p>Hmm, so the problem happens on the level of batches. Interesting -- I never experience anything like that. I can monitor the changing value of AUC  in Keras but it never caused any problems. </p>",
      "rawMarkdown": "Hmm, so the problem happens on the level of batches. Interesting -- I never experience anything like that. I can monitor the changing value of AUC  in Keras but it never caused any problems.",
      "votes": null
    },
    {
      "id": "917673",
      "postDate": "07/06/2020 17:22:21",
      "content": "<p><code>My best guess is that for some reason this did not work as intended because you ended up with no positive examples in your validation set.</code>\nIn both train &amp; validation sets I have maintained same ratio of classes (0 &amp; 1). Please do check my code &amp; output below</p>\n\n<p>```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]</p>\n\n<h1>print(\"X_columns :\", X.columns.tolist())</h1>\n\n<h1>print(\"y_column :\", y.name)</h1>\n\n<p>X_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)</p>\n\n<p>train_df = X_train.copy()\ntrain_df['target'] = y_train\ntrain_df = train_df.reset_index()</p>\n\n<p>val_df = X_test.copy()\nval_df['target'] = y_test\nval_df = val_df.reset_index()</p>\n\n<p>print(\"Training set\")\ndisplay(train_df['target'].value_counts(normalize = True))\nprint(\"Validation set\")\ndisplay(val_df['target'].value_counts(normalize  = True))</p>\n\n<p>Output:\nTraining set\n0    0.982362\n1    0.017638\nName: target, dtype: float64\nValidation set\n0    0.982391\n1    0.017609\nName: target, dtype: float64\n```\nI was using batch size of 128 for validation set and in some of the batches I am facing this issue.</p>",
      "rawMarkdown": "`My best guess is that for some reason this did not work as intended because you ended up with no positive examples in your validation set.`\nIn both train &amp; validation sets I have maintained same ratio of classes (0 &amp; 1). Please do check my code &amp; output below\n\n```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]\n\n# print(\"X_columns :\", X.columns.tolist())\n# print(\"y_column :\", y.name)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)\n\ntrain_df = X_train.copy()\ntrain_df['target'] = y_train\ntrain_df = train_df.reset_index()\n\nval_df = X_test.copy()\nval_df['target'] = y_test\nval_df = val_df.reset_index()\n\nprint(\"Training set\")\ndisplay(train_df['target'].value_counts(normalize = True))\nprint(\"Validation set\")\ndisplay(val_df['target'].value_counts(normalize  = True))\n\nOutput:\nTraining set\n0    0.982362\n1    0.017638\nName: target, dtype: float64\nValidation set\n0    0.982391\n1    0.017609\nName: target, dtype: float64\n```\nI was using batch size of 128 for validation set and in some of the batches I am facing this issue.",
      "votes": null
    },
    {
      "id": "917679",
      "postDate": "07/06/2020 17:26:50",
      "content": "<p>Please do check my code here -<a href=\"https://jovian.ml/vineel369/class-project-skin-cancer\">https://jovian.ml/vineel369/class-project-skin-cancer</a> and please let me know your feedback on it so I can correct it</p>",
      "rawMarkdown": "Please do check my code here -https://jovian.ml/vineel369/class-project-skin-cancer and please let me know your feedback on it so I can correct it",
      "votes": null
    },
    {
      "id": "918360",
      "postDate": "07/07/2020 07:53:25",
      "content": "<p><strong>Recommended way</strong></p>\n\n<p>You can \"fix\" it by creating a list of val-pred and val-targets, save your current values each iteration, concat them to an array and calculate the roc_score, both outside of the iteration. </p>\n\n<p>Code: </p>\n\n<p>```\nval_preds = []\nval_targets = []</p>\n\n<p>...</p>\n\n<h1>in the validation loop add</h1>\n\n<p>val_preds.append(outputs_sig.detach().cpu().numpy()) #your outputs\nval_targets.append(labels.detach().cpu().numpy()) #your true labels</p>\n\n<h1>outside of validation loop add</h1>\n\n<p>val_preds = np.concatenate(val_preds)\nval_targets = np.concatenate(val_targets)\nroc_score =  roc_auc_score(val_targets, val_preds)</p>\n\n<p>```</p>\n\n<p>You should not get the error anymore, now you make sure, that the y_true has to have the complete real targets. </p>\n\n<p><strong>Also possible but can lead to same error</strong></p>\n\n<p>You can also use StatifiedKFold to split the data better OR you can increase the batch-size of the valid-dataloader to make sure to have more y_true labels</p>",
      "rawMarkdown": "**Recommended way**\n\nYou can \"fix\" it by creating a list of val-pred and val-targets, save your current values each iteration, concat them to an array and calculate the roc_score, both outside of the iteration. \n\nCode: \n\n```\nval_preds = []\nval_targets = []\n\n...\n# in the validation loop add\nval_preds.append(outputs_sig.detach().cpu().numpy()) #your outputs\nval_targets.append(labels.detach().cpu().numpy()) #your true labels\n\n#outside of validation loop add\nval_preds = np.concatenate(val_preds)\nval_targets = np.concatenate(val_targets)\nroc_score =  roc_auc_score(val_targets, val_preds)\n\n```\n\nYou should not get the error anymore, now you make sure, that the y_true has to have the complete real targets. \n\n**Also possible but can lead to same error**\n\nYou can also use StatifiedKFold to split the data better OR you can increase the batch-size of the valid-dataloader to make sure to have more y_true labels",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 917585,
      "author_name": "santiviquez",
      "author_url": "",
      "post_date": "07/06/2020 16:11:21",
      "content": "<p>Hi, do you have the code in a public notebook so I can take a look?</p>",
      "votes": null,
      "replies": [
        {
          "id": 917647,
          "author_name": "vineel369",
          "author_url": "",
          "post_date": "07/06/2020 17:01:54",
          "content": "<p>Hey, You can find my notebook at <a href=\"https://jovian.ml/vineel369/class-project-skin-cancer\">https://jovian.ml/vineel369/class-project-skin-cancer</a>. </p>\n\n<p>Please provide me your feedback on it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 917608,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "07/06/2020 16:23:43",
      "content": "<p>That means all target values are \"0\" (only one class); somehow your selected fold (validation) does not contain any Malignant images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 917651,
          "author_name": "vineel369",
          "author_url": "",
          "post_date": "07/06/2020 17:05:54",
          "content": "<p>Hi Sirish, </p>\n\n<p>You are correct...  I have performed a single train-test (stratified) split by maintaining the same class ratio in both the train &amp; valid sets and I have checked the same. \n<code>\nTrain set : ~98% (0 - class) &amp; ~2 % (1 - class)\nValidation set : ~98% (0 - class) &amp; ~2 % (1 - class)\n</code>\nI have created train &amp; validation data loaders with help of DataLoader function. \nWhile calculating auc score on validation dataset, on some of the batches I am facing the mentioned value Error. Are there any strategies to perform a better train-test split for this kind of imbalance problem. ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 917614,
      "author_name": "graf10a",
      "author_url": "",
      "post_date": "07/06/2020 16:28:32",
      "content": "<p>I would recommend you to consider making your validation data set more diverse. IMHO, it is not a good idea to have only one class represented in either of your validation folds. Read, for example, this discussion topic about building a reasonable validation strategy:\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156002\">Implementation of Stratified Group K-Folds</a>\nAnd here is the public kernel illustrating the implementation of this strategy:\n<a href=\"https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds\">SIIM Stratified GroupKFold 5-folds</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 917646,
          "author_name": "vineel369",
          "author_url": "",
          "post_date": "07/06/2020 17:01:10",
          "content": "<p>Instead of using K-fold CV, I have just performed a single train-test (stratified) split by maintaining the same class ratio in both the train &amp; valid sets and then I have created train &amp; validation data loaders with help of <code>DataLoader</code> function. While calculating auc score on validation dataset, on some of the batches I am facing the mentioned value  Error. Are there any strategies to perform a better train-test split for this kind of imbalance problem. ?</p>\n\n<p>Anyhow, I will try to implement stratified k-fold as per the mentioned notebooks.  Thanks for the response</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917658,
          "author_name": "graf10a",
          "author_url": "",
          "post_date": "07/06/2020 17:09:56",
          "content": "<p>Hmm, so the problem happens on the level of batches. Interesting -- I never experience anything like that. I can monitor the changing value of AUC  in Keras but it never caused any problems. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917673,
          "author_name": "vineel369",
          "author_url": "",
          "post_date": "07/06/2020 17:22:21",
          "content": "<p><code>My best guess is that for some reason this did not work as intended because you ended up with no positive examples in your validation set.</code>\nIn both train &amp; validation sets I have maintained same ratio of classes (0 &amp; 1). Please do check my code &amp; output below</p>\n\n<p>```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]</p>\n\n<h1>print(\"X_columns :\", X.columns.tolist())</h1>\n\n<h1>print(\"y_column :\", y.name)</h1>\n\n<p>X_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)</p>\n\n<p>train_df = X_train.copy()\ntrain_df['target'] = y_train\ntrain_df = train_df.reset_index()</p>\n\n<p>val_df = X_test.copy()\nval_df['target'] = y_test\nval_df = val_df.reset_index()</p>\n\n<p>print(\"Training set\")\ndisplay(train_df['target'].value_counts(normalize = True))\nprint(\"Validation set\")\ndisplay(val_df['target'].value_counts(normalize  = True))</p>\n\n<p>Output:\nTraining set\n0    0.982362\n1    0.017638\nName: target, dtype: float64\nValidation set\n0    0.982391\n1    0.017609\nName: target, dtype: float64\n```\nI was using batch size of 128 for validation set and in some of the batches I am facing this issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917679,
          "author_name": "vineel369",
          "author_url": "",
          "post_date": "07/06/2020 17:26:50",
          "content": "<p>Please do check my code here -<a href=\"https://jovian.ml/vineel369/class-project-skin-cancer\">https://jovian.ml/vineel369/class-project-skin-cancer</a> and please let me know your feedback on it so I can correct it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 918360,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "07/07/2020 07:53:25",
      "content": "<p><strong>Recommended way</strong></p>\n\n<p>You can \"fix\" it by creating a list of val-pred and val-targets, save your current values each iteration, concat them to an array and calculate the roc_score, both outside of the iteration. </p>\n\n<p>Code: </p>\n\n<p>```\nval_preds = []\nval_targets = []</p>\n\n<p>...</p>\n\n<h1>in the validation loop add</h1>\n\n<p>val_preds.append(outputs_sig.detach().cpu().numpy()) #your outputs\nval_targets.append(labels.detach().cpu().numpy()) #your true labels</p>\n\n<h1>outside of validation loop add</h1>\n\n<p>val_preds = np.concatenate(val_preds)\nval_targets = np.concatenate(val_targets)\nroc_score =  roc_auc_score(val_targets, val_preds)</p>\n\n<p>```</p>\n\n<p>You should not get the error anymore, now you make sure, that the y_true has to have the complete real targets. </p>\n\n<p><strong>Also possible but can lead to same error</strong></p>\n\n<p>You can also use StatifiedKFold to split the data better OR you can increase the batch-size of the valid-dataloader to make sure to have more y_true labels</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "917571": "Sometimes the true value is all 0 in a batch and rocaucscore gives a `ValueError: Only one class present in y_true. ROC AUC score is not defined in that case.` \nWhat is the best way to solve this problem? - I am getting this error while calculating auc score for validation dataset.",
    "917585": "Hi, do you have the code in a public notebook so I can take a look?",
    "917608": "That means all target values are \"0\" (only one class); somehow your selected fold (validation) does not contain any Malignant images.",
    "917614": "I would recommend you to consider making your validation data set more diverse. IMHO, it is not a good idea to have only one class represented in either of your validation folds. Read, for example, this discussion topic about building a reasonable validation strategy:\n[Implementation of Stratified Group K-Folds](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156002)\nAnd here is the public kernel illustrating the implementation of this strategy:\n[SIIM Stratified GroupKFold 5-folds](https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds)",
    "917646": "Instead of using K-fold CV, I have just performed a single train-test (stratified) split by maintaining the same class ratio in both the train &amp; valid sets and then I have created train &amp; validation data loaders with help of `DataLoader` function. While calculating auc score on validation dataset, on some of the batches I am facing the mentioned value  Error. Are there any strategies to perform a better train-test split for this kind of imbalance problem. ?\n\nAnyhow, I will try to implement stratified k-fold as per the mentioned notebooks.  Thanks for the response",
    "917647": "Hey, You can find my notebook at https://jovian.ml/vineel369/class-project-skin-cancer. \n\nPlease provide me your feedback on it",
    "917651": "Hi Sirish, \n\nYou are correct...  I have performed a single train-test (stratified) split by maintaining the same class ratio in both the train &amp; valid sets and I have checked the same. \n```\nTrain set : ~98% (0 - class) &amp; ~2 % (1 - class)\nValidation set : ~98% (0 - class) &amp; ~2 % (1 - class)\n```\nI have created train &amp; validation data loaders with help of DataLoader function. \nWhile calculating auc score on validation dataset, on some of the batches I am facing the mentioned value Error. Are there any strategies to perform a better train-test split for this kind of imbalance problem. ?",
    "917658": "Hmm, so the problem happens on the level of batches. Interesting -- I never experience anything like that. I can monitor the changing value of AUC  in Keras but it never caused any problems.",
    "917673": "`My best guess is that for some reason this did not work as intended because you ended up with no positive examples in your validation set.`\nIn both train &amp; validation sets I have maintained same ratio of classes (0 &amp; 1). Please do check my code &amp; output below\n\n```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]\n\n# print(\"X_columns :\", X.columns.tolist())\n# print(\"y_column :\", y.name)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)\n\ntrain_df = X_train.copy()\ntrain_df['target'] = y_train\ntrain_df = train_df.reset_index()\n\nval_df = X_test.copy()\nval_df['target'] = y_test\nval_df = val_df.reset_index()\n\nprint(\"Training set\")\ndisplay(train_df['target'].value_counts(normalize = True))\nprint(\"Validation set\")\ndisplay(val_df['target'].value_counts(normalize  = True))\n\nOutput:\nTraining set\n0    0.982362\n1    0.017638\nName: target, dtype: float64\nValidation set\n0    0.982391\n1    0.017609\nName: target, dtype: float64\n```\nI was using batch size of 128 for validation set and in some of the batches I am facing this issue.",
    "917679": "Please do check my code here -https://jovian.ml/vineel369/class-project-skin-cancer and please let me know your feedback on it so I can correct it",
    "918360": "**Recommended way**\n\nYou can \"fix\" it by creating a list of val-pred and val-targets, save your current values each iteration, concat them to an array and calculate the roc_score, both outside of the iteration. \n\nCode: \n\n```\nval_preds = []\nval_targets = []\n\n...\n# in the validation loop add\nval_preds.append(outputs_sig.detach().cpu().numpy()) #your outputs\nval_targets.append(labels.detach().cpu().numpy()) #your true labels\n\n#outside of validation loop add\nval_preds = np.concatenate(val_preds)\nval_targets = np.concatenate(val_targets)\nroc_score =  roc_auc_score(val_targets, val_preds)\n\n```\n\nYou should not get the error anymore, now you make sure, that the y_true has to have the complete real targets. \n\n**Also possible but can lead to same error**\n\nYou can also use StatifiedKFold to split the data better OR you can increase the batch-size of the valid-dataloader to make sure to have more y_true labels"
  },
  "source": "meta"
}