{
  "id": 242204,
  "title": "AUC seems too different with score (solved)",
  "url": "/competitions/seti-breakthrough-listen/discussion/242204",
  "author_name": "assign",
  "post_date": "2021-05-28T00:46:26.679000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello,<br>\nThis is my first competition for AI in my life!</p>\n<p>But I have a trouble with different expected score.</p>\n<p>train dataset shows well divided distribution. (BCE + label smoothing 0.1) loss = 0.1996 auc = 0.9900<br>\nvalidation dataset output seems gamma distribution. loss = 0.4244. but auc = 0.9900<br>\ntest dataset outputs are make same distribution with validation dataset output (top1% output value = 0.252)</p>\n<p>train : valid = 96:4 (4% of data is validation)</p>\n<p>But, I can't understand why submission score is 0.5<br>\nAlso I know BCE loss = 0.4244 is very bad. But auc score can 0.99 if model make weak positive almost time.</p>\n<p>I want to post with distribution plot, but image cannot upload as base64 format (how other user upload image?)</p>\n<p>+EDIT1: with seabon.distplot, 0.5 is reasonable. but I cant understand why 0.9900 in val auc<br>\n+problem cuase: wrong ordering of input and target. but still I cant understand why 0.9900 in val auc<br>\n+I understand keras AUC has bug </p>",
  "messages": [
    {
      "id": 1326779,
      "postDate": "2021-05-28T17:22:24.137Z",
      "content": "<p>0.99 is a very high score!!!… possibly because of overfitting in the training process or because you are using the valid samples for train the network… Are you mixing the valid dataset with training data in the whole process of training? or is the valid dataset keep apart from training from the begining and used only at the end for the score validation?</p>",
      "rawMarkdown": "0.99 is a very high score!!!... possibly because of overfitting in the training process or because you are using the valid samples for train the network... Are you mixing the valid dataset with training data in the whole process of training? or is the valid dataset keep apart from training from the begining and used only at the end for the score validation?",
      "votes": 1,
      "replies": [
        {
          "id": 1327119,
          "postDate": "2021-05-28T23:39:24.863Z",
          "content": "<p>yes I use them definitely separated</p>\n<p>I think it was related to keras bug, their AUC is not real AUC. <br>\nI check with sklearn, validation AUC is 0.5 (not .99)</p>\n<p>also, with new corrected dataset, keras validation AUC = 0.97<br>\nbut, sklearn AUC = .855 (and submission score is 0.87)<br>\ndon't believe keras auc.</p>\n<p>I'm puzzled by the overfitting of deep learning on random labelling<br>\nbecause I can't reach 0.1996 training BCE loss with same model</p>",
          "rawMarkdown": "yes I use them definitely separated\n\nI think it was related to keras bug, their AUC is not real AUC. \nI check with sklearn, validation AUC is 0.5 (not .99)\n\nalso, with new corrected dataset, keras validation AUC = 0.97\nbut, sklearn AUC = .855 (and submission score is 0.87)\ndon't believe keras auc.\n\n\nI'm puzzled by the overfitting of deep learning on random labelling\nbecause I can't reach 0.1996 training BCE loss with same model"
        }
      ]
    },
    {
      "id": 1326774,
      "postDate": "2021-05-28T17:17:14.663Z",
      "content": "<p>I had to face a similar problem with my first model … I had AUC about 0.90 in 10 epochs but 0.50 in LB submissions</p>\n<p>My problem was that in the Dataset clas that I had implemented, I performed a 'shuffle' of the data… for all datasets, train, validation and test datasets</p>\n<p>Randomize the dataset for train dataset is a good idea, but innecesary for valid and test data, and if you do it,  you must be very carefoull from where you take the target ID after to build the submission file… My problem: I was tooking the target ID from the original submission sample file (in the original order), while the predictions had been randomly mixed by the 'shuffle' command, so they didn't match</p>\n<p>I hope it helps you </p>",
      "rawMarkdown": "I had to face a similar problem with my first model ... I had AUC about 0.90 in 10 epochs but 0.50 in LB submissions\n\nMy problem was that in the Dataset clas that I had implemented, I performed a 'shuffle' of the data... for all datasets, train, validation and test datasets\n\nRandomize the dataset for train dataset is a good idea, but innecesary for valid and test data, and if you do it,  you must be very carefoull from where you take the target ID after to build the submission file... My problem: I was tooking the target ID from the original submission sample file (in the original order), while the predictions had been randomly mixed by the 'shuffle' command, so they didn't match\n\nI hope it helps you ",
      "votes": 2
    },
    {
      "id": 1325707,
      "postDate": "2021-05-28T00:46:26.680Z",
      "content": "<p>Hello,<br>\nThis is my first competition for AI in my life!</p>\n<p>But I have a trouble with different expected score.</p>\n<p>train dataset shows well divided distribution. (BCE + label smoothing 0.1) loss = 0.1996 auc = 0.9900<br>\nvalidation dataset output seems gamma distribution. loss = 0.4244. but auc = 0.9900<br>\ntest dataset outputs are make same distribution with validation dataset output (top1% output value = 0.252)</p>\n<p>train : valid = 96:4 (4% of data is validation)</p>\n<p>But, I can't understand why submission score is 0.5<br>\nAlso I know BCE loss = 0.4244 is very bad. But auc score can 0.99 if model make weak positive almost time.</p>\n<p>I want to post with distribution plot, but image cannot upload as base64 format (how other user upload image?)</p>\n<p>+EDIT1: with seabon.distplot, 0.5 is reasonable. but I cant understand why 0.9900 in val auc<br>\n+problem cuase: wrong ordering of input and target. but still I cant understand why 0.9900 in val auc<br>\n+I understand keras AUC has bug </p>",
      "rawMarkdown": "Hello,\nThis is my first competition for AI in my life!\n\nBut I have a trouble with different expected score.\n\ntrain dataset shows well divided distribution. (BCE + label smoothing 0.1) loss = 0.1996 auc = 0.9900\nvalidation dataset output seems gamma distribution. loss = 0.4244. but auc = 0.9900\ntest dataset outputs are make same distribution with validation dataset output (top1% output value = 0.252)\n\ntrain : valid = 96:4 (4% of data is validation)\n\nBut, I can't understand why submission score is 0.5\nAlso I know BCE loss = 0.4244 is very bad. But auc score can 0.99 if model make weak positive almost time.\n\nI want to post with distribution plot, but image cannot upload as base64 format (how other user upload image?)\n\n+EDIT1: with seabon.distplot, 0.5 is reasonable. but I cant understand why 0.9900 in val auc\n+problem cuase: wrong ordering of input and target. but still I cant understand why 0.9900 in val auc\n+I understand keras AUC has bug ",
      "votes": 2
    },
    {
      "id": 1326704,
      "postDate": "2021-05-28T16:08:15.683Z",
      "content": "<p>AUC Score of 0.5 also suggests that you might have an issue with the creation of your submission file as this value indicates a random distribution of target values.    </p>\n<p>If you still hitting 0.5 after you increase the size of your validation set than look in the submission creation part of your code. </p>",
      "rawMarkdown": "AUC Score of 0.5 also suggests that you might have an issue with the creation of your submission file as this value indicates a random distribution of target values.    \n\nIf you still hitting 0.5 after you increase the size of your validation set than look in the submission creation part of your code. ",
      "replies": [
        {
          "id": 1327237,
          "postDate": "2021-05-29T05:06:40.887Z",
          "content": "<p>exactly, <br>\nat validation dataset, output distribution of both type (1 and 0) are perfectly same (even the case with val_auc = 0.99)</p>",
          "rawMarkdown": "exactly, \nat validation dataset, output distribution of both type (1 and 0) are perfectly same (even the case with val_auc = 0.99)"
        }
      ]
    },
    {
      "id": 1325869,
      "postDate": "2021-05-28T04:43:38.427Z",
      "content": "<p>What is your metric ? Also having only 4 percent of data for validation is a bad practice. A standard rule of thumb is having 20 percent of your data for validation</p>",
      "rawMarkdown": "What is your metric ? Also having only 4 percent of data for validation is a bad practice. A standard rule of thumb is having 20 percent of your data for validation",
      "replies": [
        {
          "id": 1325918,
          "postDate": "2021-05-28T05:31:22.397Z",
          "content": "<p>I use keras.metrics.AUC()</p>\n<p>thanks for your advice.</p>\n<p>I start with 10% validation set. <br>\nI thought reduce validation set size didn't matter (if each class count &gt; 100) .</p>\n<p>I'll implement with reference your advice</p>",
          "rawMarkdown": "I use keras.metrics.AUC()\n\nthanks for your advice.\n\nI start with 10% validation set. \nI thought reduce validation set size didn't matter (if each class count > 100) .\n\nI'll implement with reference your advice"
        },
        {
          "id": 1325923,
          "postDate": "2021-05-28T05:35:23.347Z",
          "content": "<p>Your metric has no problem</p>",
          "rawMarkdown": "Your metric has no problem"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1326779,
      "author_name": "ePolaris",
      "author_url": "",
      "post_date": "2021-05-28T17:22:24.137000",
      "content": "<p>0.99 is a very high score!!!… possibly because of overfitting in the training process or because you are using the valid samples for train the network… Are you mixing the valid dataset with training data in the whole process of training? or is the valid dataset keep apart from training from the begining and used only at the end for the score validation?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1327119,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-05-28T23:39:24.863000",
          "content": "<p>yes I use them definitely separated</p>\n<p>I think it was related to keras bug, their AUC is not real AUC. <br>\nI check with sklearn, validation AUC is 0.5 (not .99)</p>\n<p>also, with new corrected dataset, keras validation AUC = 0.97<br>\nbut, sklearn AUC = .855 (and submission score is 0.87)<br>\ndon't believe keras auc.</p>\n<p>I'm puzzled by the overfitting of deep learning on random labelling<br>\nbecause I can't reach 0.1996 training BCE loss with same model</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1326774,
      "author_name": "ePolaris",
      "author_url": "",
      "post_date": "2021-05-28T17:17:14.663000",
      "content": "<p>I had to face a similar problem with my first model … I had AUC about 0.90 in 10 epochs but 0.50 in LB submissions</p>\n<p>My problem was that in the Dataset clas that I had implemented, I performed a 'shuffle' of the data… for all datasets, train, validation and test datasets</p>\n<p>Randomize the dataset for train dataset is a good idea, but innecesary for valid and test data, and if you do it,  you must be very carefoull from where you take the target ID after to build the submission file… My problem: I was tooking the target ID from the original submission sample file (in the original order), while the predictions had been randomly mixed by the 'shuffle' command, so they didn't match</p>\n<p>I hope it helps you </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1326704,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-05-28T16:08:15.683000",
      "content": "<p>AUC Score of 0.5 also suggests that you might have an issue with the creation of your submission file as this value indicates a random distribution of target values.    </p>\n<p>If you still hitting 0.5 after you increase the size of your validation set than look in the submission creation part of your code. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1327237,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-05-29T05:06:40.887000",
          "content": "<p>exactly, <br>\nat validation dataset, output distribution of both type (1 and 0) are perfectly same (even the case with val_auc = 0.99)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1325869,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-05-28T04:43:38.427000",
      "content": "<p>What is your metric ? Also having only 4 percent of data for validation is a bad practice. A standard rule of thumb is having 20 percent of your data for validation</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1325918,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-05-28T05:31:22.397000",
          "content": "<p>I use keras.metrics.AUC()</p>\n<p>thanks for your advice.</p>\n<p>I start with 10% validation set. <br>\nI thought reduce validation set size didn't matter (if each class count &gt; 100) .</p>\n<p>I'll implement with reference your advice</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325923,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-05-28T05:35:23.347000",
          "content": "<p>Your metric has no problem</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1326779": "0.99 is a very high score!!!... possibly because of overfitting in the training process or because you are using the valid samples for train the network... Are you mixing the valid dataset with training data in the whole process of training? or is the valid dataset keep apart from training from the begining and used only at the end for the score validation?",
    "1326774": "I had to face a similar problem with my first model ... I had AUC about 0.90 in 10 epochs but 0.50 in LB submissions\n\nMy problem was that in the Dataset clas that I had implemented, I performed a 'shuffle' of the data... for all datasets, train, validation and test datasets\n\nRandomize the dataset for train dataset is a good idea, but innecesary for valid and test data, and if you do it,  you must be very carefoull from where you take the target ID after to build the submission file... My problem: I was tooking the target ID from the original submission sample file (in the original order), while the predictions had been randomly mixed by the 'shuffle' command, so they didn't match\n\nI hope it helps you ",
    "1325707": "Hello,\nThis is my first competition for AI in my life!\n\nBut I have a trouble with different expected score.\n\ntrain dataset shows well divided distribution. (BCE + label smoothing 0.1) loss = 0.1996 auc = 0.9900\nvalidation dataset output seems gamma distribution. loss = 0.4244. but auc = 0.9900\ntest dataset outputs are make same distribution with validation dataset output (top1% output value = 0.252)\n\ntrain : valid = 96:4 (4% of data is validation)\n\nBut, I can't understand why submission score is 0.5\nAlso I know BCE loss = 0.4244 is very bad. But auc score can 0.99 if model make weak positive almost time.\n\nI want to post with distribution plot, but image cannot upload as base64 format (how other user upload image?)\n\n+EDIT1: with seabon.distplot, 0.5 is reasonable. but I cant understand why 0.9900 in val auc\n+problem cuase: wrong ordering of input and target. but still I cant understand why 0.9900 in val auc\n+I understand keras AUC has bug ",
    "1326704": "AUC Score of 0.5 also suggests that you might have an issue with the creation of your submission file as this value indicates a random distribution of target values.    \n\nIf you still hitting 0.5 after you increase the size of your validation set than look in the submission creation part of your code. ",
    "1325869": "What is your metric ? Also having only 4 percent of data for validation is a bad practice. A standard rule of thumb is having 20 percent of your data for validation"
  }
}