{
  "id": 154624,
  "title": "how many public test samples are positive?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154624",
  "author_name": "",
  "post_date": "2020-05-29T05:34:38.912122500Z",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>without knowing the distribution of the test set, it could be difficult to make a good validation \nset ...</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F39d79554904b5982516c71d8939d4440%2FSelection_034.png?generation=1590730439956811&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "866101",
      "postDate": "05/29/2020 05:34:38",
      "content": "<p>without knowing the distribution of the test set, it could be difficult to make a good validation \nset ...</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F39d79554904b5982516c71d8939d4440%2FSelection_034.png?generation=1590730439956811&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "without knowing the distribution of the test set, it could be difficult to make a good validation \nset ...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F39d79554904b5982516c71d8939d4440%2FSelection_034.png?generation=1590730439956811&amp;alt=media)",
      "votes": null
    },
    {
      "id": "866120",
      "postDate": "05/29/2020 06:01:26",
      "content": "<p>we know that the train class is imbalance, with majority class=0\n1. submit a all prediction = 1, we should get public LB score=0.5\n2. randomly flip   N, prediction = 0, get Lb score s1\n3. randomly flip 2N, prediction = 0, get Lb score s2 \n4. randomly flip 3N, prediction = 0, get Lb score s3\n5. ...</p>\n\n<p>assuming everything is random, from the difference of (0.5-s1) you can worked out mathematically the number of +ve samples in N.  you can further confirm it with 2N, 3N, ... and worked out some mathematical relationship</p>\n\n<p>you may be surprised that this \"random guess\" performance is quite high (due to class imbalance) . so how much has your model performed better than random guess?</p>",
      "rawMarkdown": "we know that the train class is imbalance, with majority class=0\n1. submit a all prediction = 1, we should get public LB score=0.5\n2. randomly flip   N, prediction = 0, get Lb score s1\n3. randomly flip 2N, prediction = 0, get Lb score s2 \n4. randomly flip 3N, prediction = 0, get Lb score s3\n5. ...\n\nassuming everything is random, from the difference of (0.5-s1) you can worked out mathematically the number of +ve samples in N.  you can further confirm it with 2N, 3N, ... and worked out some mathematical relationship\n\nyou may be surprised that this \"random guess\" performance is quite high (due to class imbalance) . so how much has your model performed better than random guess?",
      "votes": null
    },
    {
      "id": "866241",
      "postDate": "05/29/2020 08:15:26",
      "content": "<p>```</p>\n\n<p>pos_rate = 0.019  # assume same as training data\nnum_test = 3294  #(10982*0.30) public Lb is 30% of all test samples\nnum_pos  = int(pos_rate*3295) \nnum_neg  = num_test - num_pos</p>\n\n<h1>simulate public Lb data</h1>\n\n<p>truth = np.zeros(num_test)\npos_index = np.random.choice(num_test,num_pos,replace=False) \ntruth[pos_index] = 1</p>\n\n<h1>assume we randomly set 2000 predictions to one and the rest to zero</h1>\n\n<p>probability = np.ones(num_test) <br>\nprobability[np.random.choice(neg_index,2000-int(pos_rate*2000),replace=False)]=0 \nprobability[np.random.choice(pos_index,int(pos_rate*2000),replace=False)]=0</p>\n\n<p>auc = np_metric_roc_auc(probability, truth)\nprint(auc)\nprint(num_pos/num_test)</p>\n\n<p>```</p>\n\n<p>from the python code above, we get \n0.49707561481954654\n0.018822100789313904</p>\n\n<p>now if i submit random 2000 zeros prediction to LB server, i get LB score = 0.467\nThis hows that public test LB distribution is indeed \"close to train data\", with about 0.001 to 0.019? +ve train samples</p>",
      "rawMarkdown": "```\n\npos_rate = 0.019  # assume same as training data\nnum_test = 3294  #(10982*0.30) public Lb is 30% of all test samples\nnum_pos  = int(pos_rate*3295) \nnum_neg  = num_test - num_pos\n \n# simulate public Lb data\ntruth = np.zeros(num_test)\npos_index = np.random.choice(num_test,num_pos,replace=False) \ntruth[pos_index] = 1\n\n\n# assume we randomly set 2000 predictions to one and the rest to zero\n\nprobability = np.ones(num_test)  \nprobability[np.random.choice(neg_index,2000-int(pos_rate*2000),replace=False)]=0 \nprobability[np.random.choice(pos_index,int(pos_rate*2000),replace=False)]=0\n\nauc = np_metric_roc_auc(probability, truth)\nprint(auc)\nprint(num_pos/num_test)\n\n\n```\n\nfrom the python code above, we get \n0.49707561481954654\n0.018822100789313904\n\nnow if i submit random 2000 zeros prediction to LB server, i get LB score = 0.467\nThis hows that public test LB distribution is indeed \"close to train data\", with about 0.001 to 0.019? +ve train samples",
      "votes": null
    },
    {
      "id": "866439",
      "postDate": "05/29/2020 12:03:32",
      "content": "<p>Hey <a href=\"/hengck23\">@hengck23</a>,\nSorry I don't understand your demonstration (and the sample code does not run because neg_index is not defined by the way).</p>\n\n<p>An AUC score of 0.5 corresponds to a random guess, also AUC is about target inversion and independent from target distribution.</p>\n\n<p>Here is how I like to think about AUC : pick randomly one positive sample and one negative sample, are they correctly ranked by your algorithm probabilities outputs? Repeat this with all possible positive/negative pairs : the average number of correct ranking is your AUC.</p>\n\n<p>With this in mind how does your probing can show anything about the public test set distribution? Having AUC scores around 0.5 only tells you that your algorithm does not perform better than random.</p>",
      "rawMarkdown": "Hey @hengck23,\nSorry I don't understand your demonstration (and the sample code does not run because neg_index is not defined by the way).\n\nAn AUC score of 0.5 corresponds to a random guess, also AUC is about target inversion and independent from target distribution.\n\nHere is how I like to think about AUC : pick randomly one positive sample and one negative sample, are they correctly ranked by your algorithm probabilities outputs? Repeat this with all possible positive/negative pairs : the average number of correct ranking is your AUC.\n\nWith this in mind how does your probing can show anything about the public test set distribution? Having AUC scores around 0.5 only tells you that your algorithm does not perform better than random.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 866120,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/29/2020 06:01:26",
      "content": "<p>we know that the train class is imbalance, with majority class=0\n1. submit a all prediction = 1, we should get public LB score=0.5\n2. randomly flip   N, prediction = 0, get Lb score s1\n3. randomly flip 2N, prediction = 0, get Lb score s2 \n4. randomly flip 3N, prediction = 0, get Lb score s3\n5. ...</p>\n\n<p>assuming everything is random, from the difference of (0.5-s1) you can worked out mathematically the number of +ve samples in N.  you can further confirm it with 2N, 3N, ... and worked out some mathematical relationship</p>\n\n<p>you may be surprised that this \"random guess\" performance is quite high (due to class imbalance) . so how much has your model performed better than random guess?</p>",
      "votes": null,
      "replies": [
        {
          "id": 866241,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/29/2020 08:15:26",
          "content": "<p>```</p>\n\n<p>pos_rate = 0.019  # assume same as training data\nnum_test = 3294  #(10982*0.30) public Lb is 30% of all test samples\nnum_pos  = int(pos_rate*3295) \nnum_neg  = num_test - num_pos</p>\n\n<h1>simulate public Lb data</h1>\n\n<p>truth = np.zeros(num_test)\npos_index = np.random.choice(num_test,num_pos,replace=False) \ntruth[pos_index] = 1</p>\n\n<h1>assume we randomly set 2000 predictions to one and the rest to zero</h1>\n\n<p>probability = np.ones(num_test) <br>\nprobability[np.random.choice(neg_index,2000-int(pos_rate*2000),replace=False)]=0 \nprobability[np.random.choice(pos_index,int(pos_rate*2000),replace=False)]=0</p>\n\n<p>auc = np_metric_roc_auc(probability, truth)\nprint(auc)\nprint(num_pos/num_test)</p>\n\n<p>```</p>\n\n<p>from the python code above, we get \n0.49707561481954654\n0.018822100789313904</p>\n\n<p>now if i submit random 2000 zeros prediction to LB server, i get LB score = 0.467\nThis hows that public test LB distribution is indeed \"close to train data\", with about 0.001 to 0.019? +ve train samples</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 866439,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "05/29/2020 12:03:32",
          "content": "<p>Hey <a href=\"/hengck23\">@hengck23</a>,\nSorry I don't understand your demonstration (and the sample code does not run because neg_index is not defined by the way).</p>\n\n<p>An AUC score of 0.5 corresponds to a random guess, also AUC is about target inversion and independent from target distribution.</p>\n\n<p>Here is how I like to think about AUC : pick randomly one positive sample and one negative sample, are they correctly ranked by your algorithm probabilities outputs? Repeat this with all possible positive/negative pairs : the average number of correct ranking is your AUC.</p>\n\n<p>With this in mind how does your probing can show anything about the public test set distribution? Having AUC scores around 0.5 only tells you that your algorithm does not perform better than random.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "866101": "without knowing the distribution of the test set, it could be difficult to make a good validation \nset ...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F39d79554904b5982516c71d8939d4440%2FSelection_034.png?generation=1590730439956811&amp;alt=media)",
    "866120": "we know that the train class is imbalance, with majority class=0\n1. submit a all prediction = 1, we should get public LB score=0.5\n2. randomly flip   N, prediction = 0, get Lb score s1\n3. randomly flip 2N, prediction = 0, get Lb score s2 \n4. randomly flip 3N, prediction = 0, get Lb score s3\n5. ...\n\nassuming everything is random, from the difference of (0.5-s1) you can worked out mathematically the number of +ve samples in N.  you can further confirm it with 2N, 3N, ... and worked out some mathematical relationship\n\nyou may be surprised that this \"random guess\" performance is quite high (due to class imbalance) . so how much has your model performed better than random guess?",
    "866241": "```\n\npos_rate = 0.019  # assume same as training data\nnum_test = 3294  #(10982*0.30) public Lb is 30% of all test samples\nnum_pos  = int(pos_rate*3295) \nnum_neg  = num_test - num_pos\n \n# simulate public Lb data\ntruth = np.zeros(num_test)\npos_index = np.random.choice(num_test,num_pos,replace=False) \ntruth[pos_index] = 1\n\n\n# assume we randomly set 2000 predictions to one and the rest to zero\n\nprobability = np.ones(num_test)  \nprobability[np.random.choice(neg_index,2000-int(pos_rate*2000),replace=False)]=0 \nprobability[np.random.choice(pos_index,int(pos_rate*2000),replace=False)]=0\n\nauc = np_metric_roc_auc(probability, truth)\nprint(auc)\nprint(num_pos/num_test)\n\n\n```\n\nfrom the python code above, we get \n0.49707561481954654\n0.018822100789313904\n\nnow if i submit random 2000 zeros prediction to LB server, i get LB score = 0.467\nThis hows that public test LB distribution is indeed \"close to train data\", with about 0.001 to 0.019? +ve train samples",
    "866439": "Hey @hengck23,\nSorry I don't understand your demonstration (and the sample code does not run because neg_index is not defined by the way).\n\nAn AUC score of 0.5 corresponds to a random guess, also AUC is about target inversion and independent from target distribution.\n\nHere is how I like to think about AUC : pick randomly one positive sample and one negative sample, are they correctly ranked by your algorithm probabilities outputs? Repeat this with all possible positive/negative pairs : the average number of correct ranking is your AUC.\n\nWith this in mind how does your probing can show anything about the public test set distribution? Having AUC scores around 0.5 only tells you that your algorithm does not perform better than random."
  },
  "source": "meta"
}