{
  "id": 250023,
  "title": "Roc Auc is always 0.5",
  "url": "/competitions/seti-breakthrough-listen/discussion/250023",
  "author_name": "Vadim",
  "post_date": "2021-06-30T21:09:47.762000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Why do I always get roc_auc_score equal to 0.5? I used data augmentation and normalization, several types of models, but the score does not rise. There are no errors in the order of values, I checked. I use CNN (also tried vgg), and in all predictions, I always have the number 0.50… My test loss doesn't decrease at all. Does anyone have ideas?</p>",
  "messages": [
    {
      "id": 1372063,
      "postDate": "2021-07-01T11:52:33.700Z",
      "content": "<p>0.5 means that your predictions are absolutely random (or constant). The simplest way to get 0.5 in ROC AUC is to mess with IDs and corresponding predictions - double check that your samples are not shuffled during inference. On the other hand, you've mentioned that your test loss does not decrease - it usually means that something is completely wrong with your model or training procedure. Check that your model has correct output (single neuron with sigmoid activation), you are using appropriate loss function (binary cross entropy in this case for the simplest model), your learning rate is not too high (no reason to use much higher than 0.001), shuffle is turned off for validation and inference, data preprocessing is the same for train and validation (however, augmentation is not necessary for validation data), and labels are properly feed to the model and actually contain zeroes and ones. For debugging purposes I am usually trying to drop everything that is not that necessary for succesful model training, like complex preprocessing, augmentations, custom loss functions; use the fastest and the lightest model (e.g. efficientnet_b0 or resnet18) and debug the most crucial parts of my code.</p>",
      "rawMarkdown": "0.5 means that your predictions are absolutely random (or constant). The simplest way to get 0.5 in ROC AUC is to mess with IDs and corresponding predictions - double check that your samples are not shuffled during inference. On the other hand, you've mentioned that your test loss does not decrease - it usually means that something is completely wrong with your model or training procedure. Check that your model has correct output (single neuron with sigmoid activation), you are using appropriate loss function (binary cross entropy in this case for the simplest model), your learning rate is not too high (no reason to use much higher than 0.001), shuffle is turned off for validation and inference, data preprocessing is the same for train and validation (however, augmentation is not necessary for validation data), and labels are properly feed to the model and actually contain zeroes and ones. For debugging purposes I am usually trying to drop everything that is not that necessary for succesful model training, like complex preprocessing, augmentations, custom loss functions; use the fastest and the lightest model (e.g. efficientnet_b0 or resnet18) and debug the most crucial parts of my code.",
      "votes": 3,
      "replies": [
        {
          "id": 1373065,
          "postDate": "2021-07-02T08:01:25.980Z",
          "content": "<p>Thank you, I'll try to check more carefully</p>",
          "rawMarkdown": "Thank you, I'll try to check more carefully"
        }
      ]
    },
    {
      "id": 1371301,
      "postDate": "2021-06-30T21:09:47.763Z",
      "content": "<p>Why do I always get roc_auc_score equal to 0.5? I used data augmentation and normalization, several types of models, but the score does not rise. There are no errors in the order of values, I checked. I use CNN (also tried vgg), and in all predictions, I always have the number 0.50… My test loss doesn't decrease at all. Does anyone have ideas?</p>",
      "rawMarkdown": "Why do I always get roc_auc_score equal to 0.5? I used data augmentation and normalization, several types of models, but the score does not rise. There are no errors in the order of values, I checked. I use CNN (also tried vgg), and in all predictions, I always have the number 0.50... My test loss doesn't decrease at all. Does anyone have ideas?",
      "votes": 1
    },
    {
      "id": 1371866,
      "postDate": "2021-07-01T09:27:20.323Z",
      "content": "<p>Do you submit probabilities or do you round to 0 or 1?</p>",
      "rawMarkdown": "Do you submit probabilities or do you round to 0 or 1?",
      "votes": 2,
      "replies": [
        {
          "id": 1373067,
          "postDate": "2021-07-02T08:04:20.553Z",
          "content": "<p>Probabilities. My model always predicts numbers like 0.48594.., 0.50345…, 0.52435.. etc.</p>",
          "rawMarkdown": "Probabilities. My model always predicts numbers like 0.48594.., 0.50345..., 0.52435.. etc.",
          "votes": 1
        },
        {
          "id": 1373151,
          "postDate": "2021-07-02T08:58:39.713Z",
          "content": "<p>Then, as suggested by <a href=\"https://www.kaggle.com/opanichev\" target=\"_blank\">@opanichev</a> check that your indexing is right.  It is the most common mistake: rows for your predicitons are not in the same order as the id in sample submission.</p>",
          "rawMarkdown": "Then, as suggested by @opanichev check that your indexing is right.  It is the most common mistake: rows for your predicitons are not in the same order as the id in sample submission.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1372063,
      "author_name": "Oleg Panichev",
      "author_url": "",
      "post_date": "2021-07-01T11:52:33.700000",
      "content": "<p>0.5 means that your predictions are absolutely random (or constant). The simplest way to get 0.5 in ROC AUC is to mess with IDs and corresponding predictions - double check that your samples are not shuffled during inference. On the other hand, you've mentioned that your test loss does not decrease - it usually means that something is completely wrong with your model or training procedure. Check that your model has correct output (single neuron with sigmoid activation), you are using appropriate loss function (binary cross entropy in this case for the simplest model), your learning rate is not too high (no reason to use much higher than 0.001), shuffle is turned off for validation and inference, data preprocessing is the same for train and validation (however, augmentation is not necessary for validation data), and labels are properly feed to the model and actually contain zeroes and ones. For debugging purposes I am usually trying to drop everything that is not that necessary for succesful model training, like complex preprocessing, augmentations, custom loss functions; use the fastest and the lightest model (e.g. efficientnet_b0 or resnet18) and debug the most crucial parts of my code.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1373065,
          "author_name": "Vadim",
          "author_url": "",
          "post_date": "2021-07-02T08:01:25.980000",
          "content": "<p>Thank you, I'll try to check more carefully</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1371866,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-07-01T09:27:20.323000",
      "content": "<p>Do you submit probabilities or do you round to 0 or 1?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1373067,
          "author_name": "Vadim",
          "author_url": "",
          "post_date": "2021-07-02T08:04:20.553000",
          "content": "<p>Probabilities. My model always predicts numbers like 0.48594.., 0.50345…, 0.52435.. etc.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1373151,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-07-02T08:58:39.713000",
          "content": "<p>Then, as suggested by <a href=\"https://www.kaggle.com/opanichev\" target=\"_blank\">@opanichev</a> check that your indexing is right.  It is the most common mistake: rows for your predicitons are not in the same order as the id in sample submission.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1372063": "0.5 means that your predictions are absolutely random (or constant). The simplest way to get 0.5 in ROC AUC is to mess with IDs and corresponding predictions - double check that your samples are not shuffled during inference. On the other hand, you've mentioned that your test loss does not decrease - it usually means that something is completely wrong with your model or training procedure. Check that your model has correct output (single neuron with sigmoid activation), you are using appropriate loss function (binary cross entropy in this case for the simplest model), your learning rate is not too high (no reason to use much higher than 0.001), shuffle is turned off for validation and inference, data preprocessing is the same for train and validation (however, augmentation is not necessary for validation data), and labels are properly feed to the model and actually contain zeroes and ones. For debugging purposes I am usually trying to drop everything that is not that necessary for succesful model training, like complex preprocessing, augmentations, custom loss functions; use the fastest and the lightest model (e.g. efficientnet_b0 or resnet18) and debug the most crucial parts of my code.",
    "1371301": "Why do I always get roc_auc_score equal to 0.5? I used data augmentation and normalization, several types of models, but the score does not rise. There are no errors in the order of values, I checked. I use CNN (also tried vgg), and in all predictions, I always have the number 0.50... My test loss doesn't decrease at all. Does anyone have ideas?",
    "1371866": "Do you submit probabilities or do you round to 0 or 1?"
  }
}