{
  "id": 173316,
  "title": "Need Help: AUC is always less that 0.5. What might be wrong with my model?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173316",
  "author_name": "Krishna Kishor Kammaje",
  "post_date": "2020-08-08T17:11:55.013000",
  "votes": 0,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I use a Pytorch Lightning and EfficientNet model. However, I always end up having less than 0.5 AUC scores, though CV scores look good (around 0.9). <br>\nHere is my <a href=\"https://www.kaggle.com/krisho007/melanoma-with-pylightning\" target=\"_blank\">notebook</a>, and here are the highlights.</p>\n<ul>\n<li>Triple Stratified 256*256 JPEG images from Chris Deotte</li>\n<li>simple augmentations </li>\n<li>Uses Efficientnet</li>\n<li>Uses BinaryCrossEntropyWithLogits as the Loss function</li>\n<li>5 fold CV</li>\n<li>AdamW optimizer with ReduceLROnPlateau scheduler</li>\n<li>GPU training with 12 epochs per fold, lr=1e-4</li>\n<li>Simple average of 5 fold output for Submission</li>\n</ul>\n<p>Grateful to any hints. </p>",
  "messages": [
    {
      "id": 964323,
      "postDate": "2020-08-09T18:32:33.783Z",
      "content": "<p>Your code has a few issues.</p>\n\n<p>First, you are shuffling your train data. Which is fine this is what you want to do.  But. you can't just take each output and add to  a list, and then run that through test.  Everything will be out of order.  You need to create a datastructure that is of the exact size of test, and then populate it item by item, into the exact index it belongs.</p>\n\n<p>Also you are doing binary cross entropy with logits loss, which is fine, but I see you are additionally running that through a sigmoid, which basically means you are doing double sigmoid.</p>",
      "rawMarkdown": "Your code has a few issues.\n\nFirst, you are shuffling your train data. Which is fine this is what you want to do.  But. you can't just take each output and add to  a list, and then run that through test.  Everything will be out of order.  You need to create a datastructure that is of the exact size of test, and then populate it item by item, into the exact index it belongs.\n\nAlso you are doing binary cross entropy with logits loss, which is fine, but I see you are additionally running that through a sigmoid, which basically means you are doing double sigmoid.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 964995,
          "postDate": "2020-08-10T10:07:57.473Z",
          "content": "<p>Thanks to you, I have solved the problem. The problem was due to the extra sigmoid. I removed the extra sigmoid (for test data) and now I have scores in the range of .85.<br>\nBut, since AUC only cares for the rank, extra sigmoid should not have caused any issue, isn't it? </p>\n<p>Anyway, lots of thanks to you, <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a>, and others for helping me through this. </p>",
          "rawMarkdown": "Thanks to you, I have solved the problem. The problem was due to the extra sigmoid. I removed the extra sigmoid (for test data) and now I have scores in the range of .85.\nBut, since AUC only cares for the rank, extra sigmoid should not have caused any issue, isn't it? \n\nAnyway, lots of thanks to you, @group16, and others for helping me through this. "
        },
        {
          "id": 965030,
          "postDate": "2020-08-10T10:39:16.913Z",
          "content": "<p><a href=\"/krisho007\">@krisho007</a> I am glad your problem with the low AUC is solved.  Yes, you are correct that ROC only cares about rank.  So off the top of my head, I don't have an answer for you, but perhaps someone does</p>",
          "rawMarkdown": "@krisho007 I am glad your problem with the low AUC is solved.  Yes, you are correct that ROC only cares about rank.  So off the top of my head, I don't have an answer for you, but perhaps someone does"
        }
      ]
    },
    {
      "id": 963124,
      "postDate": "2020-08-08T17:18:26.543Z",
      "content": "<p>Are you sure that your predictions are in the right order or are aligned with the corresponding images?</p>",
      "rawMarkdown": "Are you sure that your predictions are in the right order or are aligned with the corresponding images?",
      "votes": 1,
      "replies": [
        {
          "id": 963140,
          "postDate": "2020-08-08T17:31:25.577Z",
          "content": "<p>You beat me to it! This is exactly what I would check first. 😃 </p>",
          "rawMarkdown": "You beat me to it! This is exactly what I would check first. 😃 ",
          "votes": 1
        },
        {
          "id": 963150,
          "postDate": "2020-08-08T17:43:09.243Z",
          "content": "<p>Great minds think alike, Alexey ;)</p>",
          "rawMarkdown": "Great minds think alike, Alexey ;)",
          "votes": 1
        },
        {
          "id": 963873,
          "postDate": "2020-08-09T11:28:34.123Z",
          "content": "<p>Yeah, checked it. No problems there. They are in the right order. </p>",
          "rawMarkdown": "Yeah, checked it. No problems there. They are in the right order. "
        },
        {
          "id": 963917,
          "postDate": "2020-08-09T12:17:49.217Z",
          "content": "<p>Ok that's weird. Could you show me a histogram of your predictions? Are all your predictions the same?</p>\n\n<p>Perhaps the pre-processing of test and train images are different?</p>",
          "rawMarkdown": "Ok that's weird. Could you show me a histogram of your predictions? Are all your predictions the same?\n\nPerhaps the pre-processing of test and train images are different?",
          "votes": 1
        },
        {
          "id": 964215,
          "postDate": "2020-08-09T17:03:35.457Z",
          "content": "<p>Here is the histogram.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134317%2Fdaa7b27d708669c3646fa1ce4b22a92a%2Fdownload.png?generation=1596992159380433&amp;alt=media\" alt=\"\"></p>\n<p>Here is a line plot (x-axis is indices)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134317%2F9f83cd0c5c7ff01db8011545164e6d8c%2Fdownload%20(1).png?generation=1596992609934818&amp;alt=media\" alt=\"\"></p>\n<p>Here is my augmentations</p>\n<pre><code>def get_train_transforms():\n    return A.Compose([\n            A.HorizontalFlip(p=0.5),\n            A.VerticalFlip(p=0.5),\n            A.GaussianBlur(p=0.3),\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n\ndef get_valid_transforms():\n    return A.Compose([\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n\ndef get_tta_transforms():\n    return A.Compose([\n            A.HorizontalFlip(p=0.5),\n            A.VerticalFlip(p=0.5),\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n</code></pre>\n<p>A complete notebook is <a href=\"https://www.kaggle.com/krisho007/melanoma-with-pylightning\" target=\"_blank\">here</a>. </p>",
          "rawMarkdown": "Here is the histogram.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134317%2Fdaa7b27d708669c3646fa1ce4b22a92a%2Fdownload.png?generation=1596992159380433&amp;alt=media)\n\nHere is a line plot (x-axis is indices)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134317%2F9f83cd0c5c7ff01db8011545164e6d8c%2Fdownload%20(1).png?generation=1596992609934818&amp;alt=media)\n\nHere is my augmentations\n```\ndef get_train_transforms():\n    return A.Compose([\n            A.HorizontalFlip(p=0.5),\n            A.VerticalFlip(p=0.5),\n            A.GaussianBlur(p=0.3),\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n\ndef get_valid_transforms():\n    return A.Compose([\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n\ndef get_tta_transforms():\n    return A.Compose([\n            A.HorizontalFlip(p=0.5),\n            A.VerticalFlip(p=0.5),\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n```\n\nA complete notebook is [here](https://www.kaggle.com/krisho007/melanoma-with-pylightning). \n"
        }
      ]
    },
    {
      "id": 963926,
      "postDate": "2020-08-09T12:29:17.617Z",
      "content": "<p>i would pick one fold, make a submission using just that fold, and then compare the test and validation auc. if there's no big difference, the issue must be in aggregating test predictions from multiple folds. </p>",
      "rawMarkdown": "i would pick one fold, make a submission using just that fold, and then compare the test and validation auc. if there's no big difference, the issue must be in aggregating test predictions from multiple folds. \n"
    },
    {
      "id": 963449,
      "postDate": "2020-08-09T03:38:19.873Z",
      "content": "<p>just post your code snippet, where you grab preds then calculate AOC/ROC the entire loop</p>",
      "rawMarkdown": "just post your code snippet, where you grab preds then calculate AOC/ROC the entire loop"
    },
    {
      "id": 963421,
      "postDate": "2020-08-09T02:40:29.803Z",
      "content": "<p>Is your train AUC high but validation AUC low? If so, it might be because your validation set is shuffled too. I made the same silly mistake which took me a day to figure out.</p>",
      "rawMarkdown": "Is your train AUC high but validation AUC low? If so, it might be because your validation set is shuffled too. I made the same silly mistake which took me a day to figure out.",
      "replies": [
        {
          "id": 963875,
          "postDate": "2020-08-09T11:29:46.277Z",
          "content": "<p>Both train and validation AUCs are around 0.9. The issue is only with the test AUC. </p>",
          "rawMarkdown": "Both train and validation AUCs are around 0.9. The issue is only with the test AUC. "
        }
      ]
    },
    {
      "id": 963112,
      "postDate": "2020-08-08T17:11:55.013Z",
      "content": "<p>I use a Pytorch Lightning and EfficientNet model. However, I always end up having less than 0.5 AUC scores, though CV scores look good (around 0.9). <br>\nHere is my <a href=\"https://www.kaggle.com/krisho007/melanoma-with-pylightning\" target=\"_blank\">notebook</a>, and here are the highlights.</p>\n<ul>\n<li>Triple Stratified 256*256 JPEG images from Chris Deotte</li>\n<li>simple augmentations </li>\n<li>Uses Efficientnet</li>\n<li>Uses BinaryCrossEntropyWithLogits as the Loss function</li>\n<li>5 fold CV</li>\n<li>AdamW optimizer with ReduceLROnPlateau scheduler</li>\n<li>GPU training with 12 epochs per fold, lr=1e-4</li>\n<li>Simple average of 5 fold output for Submission</li>\n</ul>\n<p>Grateful to any hints. </p>",
      "rawMarkdown": "I use a Pytorch Lightning and EfficientNet model. However, I always end up having less than 0.5 AUC scores, though CV scores look good (around 0.9). \nHere is my [notebook](https://www.kaggle.com/krisho007/melanoma-with-pylightning), and here are the highlights.\n\n- Triple Stratified 256*256 JPEG images from Chris Deotte\n- simple augmentations \n- Uses Efficientnet\n- Uses BinaryCrossEntropyWithLogits as the Loss function\n- 5 fold CV\n- AdamW optimizer with ReduceLROnPlateau scheduler\n- GPU training with 12 epochs per fold, lr=1e-4\n- Simple average of 5 fold output for Submission\n\nGrateful to any hints. "
    },
    {
      "id": 963423,
      "postDate": "2020-08-09T02:40:42.007Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 964323,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-08-09T18:32:33.783000",
      "content": "<p>Your code has a few issues.</p>\n\n<p>First, you are shuffling your train data. Which is fine this is what you want to do.  But. you can't just take each output and add to  a list, and then run that through test.  Everything will be out of order.  You need to create a datastructure that is of the exact size of test, and then populate it item by item, into the exact index it belongs.</p>\n\n<p>Also you are doing binary cross entropy with logits loss, which is fine, but I see you are additionally running that through a sigmoid, which basically means you are doing double sigmoid.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 964995,
          "author_name": "Krishna Kishor Kammaje",
          "author_url": "",
          "post_date": "2020-08-10T10:07:57.473000",
          "content": "<p>Thanks to you, I have solved the problem. The problem was due to the extra sigmoid. I removed the extra sigmoid (for test data) and now I have scores in the range of .85.<br>\nBut, since AUC only cares for the rank, extra sigmoid should not have caused any issue, isn't it? </p>\n<p>Anyway, lots of thanks to you, <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a>, and others for helping me through this. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 965030,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-10T10:39:16.913000",
          "content": "<p><a href=\"/krisho007\">@krisho007</a> I am glad your problem with the low AUC is solved.  Yes, you are correct that ROC only cares about rank.  So off the top of my head, I don't have an answer for you, but perhaps someone does</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 963124,
      "author_name": "Gilles Vandewiele",
      "author_url": "",
      "post_date": "2020-08-08T17:18:26.543000",
      "content": "<p>Are you sure that your predictions are in the right order or are aligned with the corresponding images?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 963140,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-08T17:31:25.577000",
          "content": "<p>You beat me to it! This is exactly what I would check first. 😃 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 963150,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-08T17:43:09.243000",
          "content": "<p>Great minds think alike, Alexey ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 963873,
          "author_name": "Krishna Kishor Kammaje",
          "author_url": "",
          "post_date": "2020-08-09T11:28:34.123000",
          "content": "<p>Yeah, checked it. No problems there. They are in the right order. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 963917,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-09T12:17:49.217000",
          "content": "<p>Ok that's weird. Could you show me a histogram of your predictions? Are all your predictions the same?</p>\n\n<p>Perhaps the pre-processing of test and train images are different?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 964215,
          "author_name": "Krishna Kishor Kammaje",
          "author_url": "",
          "post_date": "2020-08-09T17:03:35.457000",
          "content": "<p>Here is the histogram.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134317%2Fdaa7b27d708669c3646fa1ce4b22a92a%2Fdownload.png?generation=1596992159380433&amp;alt=media\" alt=\"\"></p>\n<p>Here is a line plot (x-axis is indices)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134317%2F9f83cd0c5c7ff01db8011545164e6d8c%2Fdownload%20(1).png?generation=1596992609934818&amp;alt=media\" alt=\"\"></p>\n<p>Here is my augmentations</p>\n<pre><code>def get_train_transforms():\n    return A.Compose([\n            A.HorizontalFlip(p=0.5),\n            A.VerticalFlip(p=0.5),\n            A.GaussianBlur(p=0.3),\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n\ndef get_valid_transforms():\n    return A.Compose([\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n\ndef get_tta_transforms():\n    return A.Compose([\n            A.HorizontalFlip(p=0.5),\n            A.VerticalFlip(p=0.5),\n            A.Normalize(mean, std, max_pixel_value=255, always_apply=True),\n            ToTensorV2(),\n        ], p=1.0)\n</code></pre>\n<p>A complete notebook is <a href=\"https://www.kaggle.com/krisho007/melanoma-with-pylightning\" target=\"_blank\">here</a>. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 963926,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "2020-08-09T12:29:17.617000",
      "content": "<p>i would pick one fold, make a submission using just that fold, and then compare the test and validation auc. if there's no big difference, the issue must be in aggregating test predictions from multiple folds. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 963449,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-08-09T03:38:19.873000",
      "content": "<p>just post your code snippet, where you grab preds then calculate AOC/ROC the entire loop</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 963421,
      "author_name": "TeYang Lau",
      "author_url": "",
      "post_date": "2020-08-09T02:40:29.803000",
      "content": "<p>Is your train AUC high but validation AUC low? If so, it might be because your validation set is shuffled too. I made the same silly mistake which took me a day to figure out.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 963875,
          "author_name": "Krishna Kishor Kammaje",
          "author_url": "",
          "post_date": "2020-08-09T11:29:46.277000",
          "content": "<p>Both train and validation AUCs are around 0.9. The issue is only with the test AUC. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 963423,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-09T02:40:42.007000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "964323": "Your code has a few issues.\n\nFirst, you are shuffling your train data. Which is fine this is what you want to do.  But. you can't just take each output and add to  a list, and then run that through test.  Everything will be out of order.  You need to create a datastructure that is of the exact size of test, and then populate it item by item, into the exact index it belongs.\n\nAlso you are doing binary cross entropy with logits loss, which is fine, but I see you are additionally running that through a sigmoid, which basically means you are doing double sigmoid.\n\n",
    "963124": "Are you sure that your predictions are in the right order or are aligned with the corresponding images?",
    "963926": "i would pick one fold, make a submission using just that fold, and then compare the test and validation auc. if there's no big difference, the issue must be in aggregating test predictions from multiple folds. \n",
    "963449": "just post your code snippet, where you grab preds then calculate AOC/ROC the entire loop",
    "963421": "Is your train AUC high but validation AUC low? If so, it might be because your validation set is shuffled too. I made the same silly mistake which took me a day to figure out.",
    "963112": "I use a Pytorch Lightning and EfficientNet model. However, I always end up having less than 0.5 AUC scores, though CV scores look good (around 0.9). \nHere is my [notebook](https://www.kaggle.com/krisho007/melanoma-with-pylightning), and here are the highlights.\n\n- Triple Stratified 256*256 JPEG images from Chris Deotte\n- simple augmentations \n- Uses Efficientnet\n- Uses BinaryCrossEntropyWithLogits as the Loss function\n- 5 fold CV\n- AdamW optimizer with ReduceLROnPlateau scheduler\n- GPU training with 12 epochs per fold, lr=1e-4\n- Simple average of 5 fold output for Submission\n\nGrateful to any hints. ",
    "963423": ""
  }
}