{
  "id": 268612,
  "title": "How can I use pseudo label properly？",
  "url": "/competitions/seti-breakthrough-listen/discussion/268612",
  "author_name": "Gainover",
  "post_date": "2021-08-28T03:39:10.149000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Friends, I have two questions I want to ask.</p>\n<ol>\n<li><p>How should I use pseudo tags. My question is: our prediction on the test set in this competition is not 0 or 1, but a decimal. After I trained the model in the stage1 and made predictions on the test set, how should I merge the corresponding prediction results with <code>tain_labels.csv</code>? What I think of is to convert these decimals to 0 or 1 through a threshold, but I don't know how to do it. Because as far as I know, when using the <code>sklearn</code> library to calculate <code>auc</code>, the threshold for distinguishing the sample label as 0 or 1 is automatically determined by the algorithm. How do I know this threshold?</p></li>\n<li><p>After I finish the pseudo-labeling, how can I select samples with high confidence to join the original train set for training in stage2? What kind of sample is considered a high degree of confidence?</p></li>\n</ol>\n<p>In addition, due to my limited strength, if there is a problem with my statement, please correct me. thanks.</p>",
  "messages": [
    {
      "id": 1493578,
      "postDate": "2021-08-28T03:39:10.150Z",
      "content": "<p>Friends, I have two questions I want to ask.</p>\n<ol>\n<li><p>How should I use pseudo tags. My question is: our prediction on the test set in this competition is not 0 or 1, but a decimal. After I trained the model in the stage1 and made predictions on the test set, how should I merge the corresponding prediction results with <code>tain_labels.csv</code>? What I think of is to convert these decimals to 0 or 1 through a threshold, but I don't know how to do it. Because as far as I know, when using the <code>sklearn</code> library to calculate <code>auc</code>, the threshold for distinguishing the sample label as 0 or 1 is automatically determined by the algorithm. How do I know this threshold?</p></li>\n<li><p>After I finish the pseudo-labeling, how can I select samples with high confidence to join the original train set for training in stage2? What kind of sample is considered a high degree of confidence?</p></li>\n</ol>\n<p>In addition, due to my limited strength, if there is a problem with my statement, please correct me. thanks.</p>",
      "rawMarkdown": "Friends, I have two questions I want to ask.\n\n1. How should I use pseudo tags. My question is: our prediction on the test set in this competition is not 0 or 1, but a decimal. After I trained the model in the stage1 and made predictions on the test set, how should I merge the corresponding prediction results with `tain_labels.csv`? What I think of is to convert these decimals to 0 or 1 through a threshold, but I don't know how to do it. Because as far as I know, when using the `sklearn` library to calculate `auc`, the threshold for distinguishing the sample label as 0 or 1 is automatically determined by the algorithm. How do I know this threshold?\n\n2. After I finish the pseudo-labeling, how can I select samples with high confidence to join the original train set for training in stage2? What kind of sample is considered a high degree of confidence?\n\nIn addition, due to my limited strength, if there is a problem with my statement, please correct me. thanks.",
      "votes": 1
    },
    {
      "id": 1494371,
      "postDate": "2021-08-28T16:12:40.020Z",
      "content": "<p>Hi. You can include pseudo labels as either hard (1s and 0s) or soft targets (decimals). And you can include all test predictions or a subset of test predictions (only confidence predictions).</p>\n<p>AUC is the metric and CrossEntropy is the loss. When you add pseudo labeled test data to train, you should not change your validation data which contains only train data. Therefore when you compute AUC on your validation data, there is no problem (because validation data contains no test data).</p>\n<p>We add pseudo labeled test data to the train data which uses CrossEntropy as loss (and not AUC). There is no problem using decimals with CrossEntropy loss. (CrossEntropy works with both decimals and \"1s and 0s\").</p>\n<p>Since CrossEntropy can handle both decimals and \"1s and 0s\", we have a choice to either leave pseudo test as decimals or we can convert with a threshold like <code>test.loc[test.pred&gt;0.5,'pred'] = 1</code> and <code>test.loc[test.pred&lt;=0.5,'pred']=0</code>. Lastly we can add all the pseudo labeled test data to train, or we can add only test data with certain confidence, like <code>add_data = test.loc[(test.pred&gt;0.8)|(test.pred&lt;0.2)]</code>.</p>\n<p>Lastly we should point out that this SETI competition test images are very different than train images. So adding pseudo label will only work if we simultaneously use aggressive mixup augmentation too (alpha&gt;3, and max target). Otherwise, your model will just learn to recognize the difference between train and test, and not target=0 and target=1.</p>",
      "rawMarkdown": "Hi. You can include pseudo labels as either hard (1s and 0s) or soft targets (decimals). And you can include all test predictions or a subset of test predictions (only confidence predictions).\n\nAUC is the metric and CrossEntropy is the loss. When you add pseudo labeled test data to train, you should not change your validation data which contains only train data. Therefore when you compute AUC on your validation data, there is no problem (because validation data contains no test data).\n\nWe add pseudo labeled test data to the train data which uses CrossEntropy as loss (and not AUC). There is no problem using decimals with CrossEntropy loss. (CrossEntropy works with both decimals and \"1s and 0s\").\n\nSince CrossEntropy can handle both decimals and \"1s and 0s\", we have a choice to either leave pseudo test as decimals or we can convert with a threshold like `test.loc[test.pred>0.5,'pred'] = 1` and `test.loc[test.pred<=0.5,'pred']=0`. Lastly we can add all the pseudo labeled test data to train, or we can add only test data with certain confidence, like `add_data = test.loc[(test.pred>0.8)|(test.pred<0.2)]`.\n\nLastly we should point out that this SETI competition test images are very different than train images. So adding pseudo label will only work if we simultaneously use aggressive mixup augmentation too (alpha>3, and max target). Otherwise, your model will just learn to recognize the difference between train and test, and not target=0 and target=1.",
      "votes": 2,
      "replies": [
        {
          "id": 1494732,
          "postDate": "2021-08-29T03:53:29.477Z",
          "content": "<p>Hi, Chris, thank you for your patient and detailed answers.</p>\n<p>I still have a question: as you mentioned \"we can convert with a threshold like <code>test.loc[test.pred&gt;0.5,'pred'] = 1</code> and <code>test.loc[test.pred&lt;=0.5,'pred' ]=0.</code>\" But the threshold that we decide by ourselves, is there a way to know the threshold calculated by AUC (I mean I want to know what is the exact threshold calculated by AUC?).</p>",
          "rawMarkdown": "Hi, Chris, thank you for your patient and detailed answers.\n\nI still have a question: as you mentioned \"we can convert with a threshold like `test.loc[test.pred>0.5,'pred'] = 1` and `test.loc[test.pred<=0.5,'pred' ]=0.`\" But the threshold that we decide by ourselves, is there a way to know the threshold calculated by AUC (I mean I want to know what is the exact threshold calculated by AUC?).",
          "votes": 1
        },
        {
          "id": 1495371,
          "postDate": "2021-08-29T14:04:40.387Z",
          "content": "<p>AUC is the \"area under the <a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic\" target=\"_blank\">ROC curve</a>\". Neither the AUC nor ROC curve finds a threshold. Instead the ROC plots all thresholds at once.</p>\n<p>When we see the curve (i.e a line drawn from coordinates (0,0) to (1,1) ), that curve is really lots of connected dots. Each dot represents one threshold. Therefore the \"curve\" is showing us all thresholds at once.</p>\n<p>For each threshold (dot on the curve), we have an x coordinate and y coordinate. The x coordinate is FPR (false positive rate) and the y coordinate is TPR (ie. true positive rate or recall).</p>\n<p>Therefore there never is a <strong>best threshold</strong>. If our application requires us to find as many positives as possible, we choose a threshold with higher TPR (recall). If our application dictates that we waste too much money with false positives, then we choose a threshold with lower FPR (which gives us lower TPR too).</p>\n<p>So in conclusion, you need to pick your own threshold. And there are different ways to do this, but using <code>0.5</code> is the most naive and common (if your model is outputting probabilities).</p>",
          "rawMarkdown": "AUC is the \"area under the [ROC curve][1]\". Neither the AUC nor ROC curve finds a threshold. Instead the ROC plots all thresholds at once.\n\nWhen we see the curve (i.e a line drawn from coordinates (0,0) to (1,1) ), that curve is really lots of connected dots. Each dot represents one threshold. Therefore the \"curve\" is showing us all thresholds at once.\n\nFor each threshold (dot on the curve), we have an x coordinate and y coordinate. The x coordinate is FPR (false positive rate) and the y coordinate is TPR (ie. true positive rate or recall).\n\nTherefore there never is a **best threshold**. If our application requires us to find as many positives as possible, we choose a threshold with higher TPR (recall). If our application dictates that we waste too much money with false positives, then we choose a threshold with lower FPR (which gives us lower TPR too).\n\nSo in conclusion, you need to pick your own threshold. And there are different ways to do this, but using `0.5` is the most naive and common (if your model is outputting probabilities).\n\n[1]: https://en.wikipedia.org/wiki/Receiver_operating_characteristic ",
          "votes": 2
        },
        {
          "id": 1496063,
          "postDate": "2021-08-30T04:20:22.563Z",
          "content": "<p>Thank you for answering my doubts, I have learned a lot from you.</p>",
          "rawMarkdown": "Thank you for answering my doubts, I have learned a lot from you.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1494371,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-28T16:12:40.020000",
      "content": "<p>Hi. You can include pseudo labels as either hard (1s and 0s) or soft targets (decimals). And you can include all test predictions or a subset of test predictions (only confidence predictions).</p>\n<p>AUC is the metric and CrossEntropy is the loss. When you add pseudo labeled test data to train, you should not change your validation data which contains only train data. Therefore when you compute AUC on your validation data, there is no problem (because validation data contains no test data).</p>\n<p>We add pseudo labeled test data to the train data which uses CrossEntropy as loss (and not AUC). There is no problem using decimals with CrossEntropy loss. (CrossEntropy works with both decimals and \"1s and 0s\").</p>\n<p>Since CrossEntropy can handle both decimals and \"1s and 0s\", we have a choice to either leave pseudo test as decimals or we can convert with a threshold like <code>test.loc[test.pred&gt;0.5,'pred'] = 1</code> and <code>test.loc[test.pred&lt;=0.5,'pred']=0</code>. Lastly we can add all the pseudo labeled test data to train, or we can add only test data with certain confidence, like <code>add_data = test.loc[(test.pred&gt;0.8)|(test.pred&lt;0.2)]</code>.</p>\n<p>Lastly we should point out that this SETI competition test images are very different than train images. So adding pseudo label will only work if we simultaneously use aggressive mixup augmentation too (alpha&gt;3, and max target). Otherwise, your model will just learn to recognize the difference between train and test, and not target=0 and target=1.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1494732,
          "author_name": "Gainover",
          "author_url": "",
          "post_date": "2021-08-29T03:53:29.477000",
          "content": "<p>Hi, Chris, thank you for your patient and detailed answers.</p>\n<p>I still have a question: as you mentioned \"we can convert with a threshold like <code>test.loc[test.pred&gt;0.5,'pred'] = 1</code> and <code>test.loc[test.pred&lt;=0.5,'pred' ]=0.</code>\" But the threshold that we decide by ourselves, is there a way to know the threshold calculated by AUC (I mean I want to know what is the exact threshold calculated by AUC?).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1495371,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-29T14:04:40.387000",
          "content": "<p>AUC is the \"area under the <a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic\" target=\"_blank\">ROC curve</a>\". Neither the AUC nor ROC curve finds a threshold. Instead the ROC plots all thresholds at once.</p>\n<p>When we see the curve (i.e a line drawn from coordinates (0,0) to (1,1) ), that curve is really lots of connected dots. Each dot represents one threshold. Therefore the \"curve\" is showing us all thresholds at once.</p>\n<p>For each threshold (dot on the curve), we have an x coordinate and y coordinate. The x coordinate is FPR (false positive rate) and the y coordinate is TPR (ie. true positive rate or recall).</p>\n<p>Therefore there never is a <strong>best threshold</strong>. If our application requires us to find as many positives as possible, we choose a threshold with higher TPR (recall). If our application dictates that we waste too much money with false positives, then we choose a threshold with lower FPR (which gives us lower TPR too).</p>\n<p>So in conclusion, you need to pick your own threshold. And there are different ways to do this, but using <code>0.5</code> is the most naive and common (if your model is outputting probabilities).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1496063,
          "author_name": "Gainover",
          "author_url": "",
          "post_date": "2021-08-30T04:20:22.563000",
          "content": "<p>Thank you for answering my doubts, I have learned a lot from you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1493578": "Friends, I have two questions I want to ask.\n\n1. How should I use pseudo tags. My question is: our prediction on the test set in this competition is not 0 or 1, but a decimal. After I trained the model in the stage1 and made predictions on the test set, how should I merge the corresponding prediction results with `tain_labels.csv`? What I think of is to convert these decimals to 0 or 1 through a threshold, but I don't know how to do it. Because as far as I know, when using the `sklearn` library to calculate `auc`, the threshold for distinguishing the sample label as 0 or 1 is automatically determined by the algorithm. How do I know this threshold?\n\n2. After I finish the pseudo-labeling, how can I select samples with high confidence to join the original train set for training in stage2? What kind of sample is considered a high degree of confidence?\n\nIn addition, due to my limited strength, if there is a problem with my statement, please correct me. thanks.",
    "1494371": "Hi. You can include pseudo labels as either hard (1s and 0s) or soft targets (decimals). And you can include all test predictions or a subset of test predictions (only confidence predictions).\n\nAUC is the metric and CrossEntropy is the loss. When you add pseudo labeled test data to train, you should not change your validation data which contains only train data. Therefore when you compute AUC on your validation data, there is no problem (because validation data contains no test data).\n\nWe add pseudo labeled test data to the train data which uses CrossEntropy as loss (and not AUC). There is no problem using decimals with CrossEntropy loss. (CrossEntropy works with both decimals and \"1s and 0s\").\n\nSince CrossEntropy can handle both decimals and \"1s and 0s\", we have a choice to either leave pseudo test as decimals or we can convert with a threshold like `test.loc[test.pred>0.5,'pred'] = 1` and `test.loc[test.pred<=0.5,'pred']=0`. Lastly we can add all the pseudo labeled test data to train, or we can add only test data with certain confidence, like `add_data = test.loc[(test.pred>0.8)|(test.pred<0.2)]`.\n\nLastly we should point out that this SETI competition test images are very different than train images. So adding pseudo label will only work if we simultaneously use aggressive mixup augmentation too (alpha>3, and max target). Otherwise, your model will just learn to recognize the difference between train and test, and not target=0 and target=1."
  }
}