{
  "id": 223734,
  "title": "i make a mistake for pseudo label",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/223734",
  "author_name": "",
  "post_date": "2021-03-05T09:59:15.905483800Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>i begin experiments for knowledge distillation ….<br>\nthen i realize i make a mistake.</p>\n<p>the LB metric is AUC and not accuracy.<br>\nhence the pseudo label may not be the label itself …</p>\n<p>AUC is the probability of correct ranking. with so high score of LB 0.96+,  pseudo rank label is most likely to be correct! (though the pseudo label may be slightly worse) </p>",
  "messages": [
    {
      "id": "1227212",
      "postDate": "03/05/2021 09:59:15",
      "content": "<p>i begin experiments for knowledge distillation ….<br>\nthen i realize i make a mistake.</p>\n<p>the LB metric is AUC and not accuracy.<br>\nhence the pseudo label may not be the label itself …</p>\n<p>AUC is the probability of correct ranking. with so high score of LB 0.96+,  pseudo rank label is most likely to be correct! (though the pseudo label may be slightly worse) </p>",
      "rawMarkdown": "i begin experiments for knowledge distillation ....\nthen i realize i make a mistake.\n\nthe LB metric is AUC and not accuracy.\nhence the pseudo label may not be the label itself ...\n\nAUC is the probability of correct ranking. with so high score of LB 0.96+,  pseudo rank label is most likely to be correct! (though the pseudo label may be slightly worse)",
      "votes": null
    },
    {
      "id": "1227216",
      "postDate": "03/05/2021 10:02:23",
      "content": "<p>in fact, i have read paper that says  rank loss benefits from +ve and unlabel data (external or test in this case)<br>\nbut i hvaen't try these before:<br>\ne.g. <a href=\"https://openaccess.thecvf.com/content_cvpr_2016/papers/Kanehira_Multi-Label_Ranking_From_CVPR_2016_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_cvpr_2016/papers/Kanehira_Multi-Label_Ranking_From_CVPR_2016_paper.pdf</a></p>\n<p><a href=\"https://github.com/t-sakai-kure/pywsl\" target=\"_blank\">https://github.com/t-sakai-kure/pywsl</a><br>\n<a href=\"https://github.com/aldro61/pu-learning\" target=\"_blank\">https://github.com/aldro61/pu-learning</a></p>",
      "rawMarkdown": "in fact, i have read paper that says  rank loss benefits from +ve and unlabel data (external or test in this case)\nbut i hvaen't try these before:\ne.g. https://openaccess.thecvf.com/content_cvpr_2016/papers/Kanehira_Multi-Label_Ranking_From_CVPR_2016_paper.pdf\n\nhttps://github.com/t-sakai-kure/pywsl\nhttps://github.com/aldro61/pu-learning",
      "votes": null
    },
    {
      "id": "1228518",
      "postDate": "03/06/2021 13:57:18",
      "content": "<p>just a thought experiment.</p>\n<p>after making pseudo labels, train an adversarial binary classifier for true labels versus pseudo labels.<br>\nThis tells you if the distributions are the same or not</p>",
      "rawMarkdown": "just a thought experiment.\n\nafter making pseudo labels, train an adversarial binary classifier for true labels versus pseudo labels.\nThis tells you if the distributions are the same or not",
      "votes": null
    },
    {
      "id": "1228639",
      "postDate": "03/06/2021 16:02:46",
      "content": "<p>It seems like this should not really be a concern if you just only validate on original labeled data. I have been doing experiments with pseudolabeling and I add the pseudolabeled data into train, but then leave the validation fold as is. Training metrics look a little weird. If I really wanted to I could also show train metrics on only the truly labeled train samples, but seeing accurate training loss and auc numbers arent as important in my opinion. </p>",
      "rawMarkdown": "It seems like this should not really be a concern if you just only validate on original labeled data. I have been doing experiments with pseudolabeling and I add the pseudolabeled data into train, but then leave the validation fold as is. Training metrics look a little weird. If I really wanted to I could also show train metrics on only the truly labeled train samples, but seeing accurate training loss and auc numbers arent as important in my opinion.",
      "votes": null
    },
    {
      "id": "1232079",
      "postDate": "03/09/2021 13:25:40",
      "content": "<p>too tough to boost😳</p>",
      "rawMarkdown": "too tough to boost😳",
      "votes": null
    },
    {
      "id": "1232453",
      "postDate": "03/09/2021 18:55:32",
      "content": "<p>Even if the metric was accuracy, f1 score or whatever that requires class labels, isn't it better to use soft predictions as pseudo labels? The motivation is to distinguish between confident and not confident predictions.</p>",
      "rawMarkdown": "Even if the metric was accuracy, f1 score or whatever that requires class labels, isn't it better to use soft predictions as pseudo labels? The motivation is to distinguish between confident and not confident predictions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1227216,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/05/2021 10:02:23",
      "content": "<p>in fact, i have read paper that says  rank loss benefits from +ve and unlabel data (external or test in this case)<br>\nbut i hvaen't try these before:<br>\ne.g. <a href=\"https://openaccess.thecvf.com/content_cvpr_2016/papers/Kanehira_Multi-Label_Ranking_From_CVPR_2016_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_cvpr_2016/papers/Kanehira_Multi-Label_Ranking_From_CVPR_2016_paper.pdf</a></p>\n<p><a href=\"https://github.com/t-sakai-kure/pywsl\" target=\"_blank\">https://github.com/t-sakai-kure/pywsl</a><br>\n<a href=\"https://github.com/aldro61/pu-learning\" target=\"_blank\">https://github.com/aldro61/pu-learning</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1228518,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/06/2021 13:57:18",
      "content": "<p>just a thought experiment.</p>\n<p>after making pseudo labels, train an adversarial binary classifier for true labels versus pseudo labels.<br>\nThis tells you if the distributions are the same or not</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1228639,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "03/06/2021 16:02:46",
      "content": "<p>It seems like this should not really be a concern if you just only validate on original labeled data. I have been doing experiments with pseudolabeling and I add the pseudolabeled data into train, but then leave the validation fold as is. Training metrics look a little weird. If I really wanted to I could also show train metrics on only the truly labeled train samples, but seeing accurate training loss and auc numbers arent as important in my opinion. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1232079,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "03/09/2021 13:25:40",
      "content": "<p>too tough to boost😳</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1232453,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "03/09/2021 18:55:32",
      "content": "<p>Even if the metric was accuracy, f1 score or whatever that requires class labels, isn't it better to use soft predictions as pseudo labels? The motivation is to distinguish between confident and not confident predictions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1227212": "i begin experiments for knowledge distillation ....\nthen i realize i make a mistake.\n\nthe LB metric is AUC and not accuracy.\nhence the pseudo label may not be the label itself ...\n\nAUC is the probability of correct ranking. with so high score of LB 0.96+,  pseudo rank label is most likely to be correct! (though the pseudo label may be slightly worse)",
    "1227216": "in fact, i have read paper that says  rank loss benefits from +ve and unlabel data (external or test in this case)\nbut i hvaen't try these before:\ne.g. https://openaccess.thecvf.com/content_cvpr_2016/papers/Kanehira_Multi-Label_Ranking_From_CVPR_2016_paper.pdf\n\nhttps://github.com/t-sakai-kure/pywsl\nhttps://github.com/aldro61/pu-learning",
    "1228518": "just a thought experiment.\n\nafter making pseudo labels, train an adversarial binary classifier for true labels versus pseudo labels.\nThis tells you if the distributions are the same or not",
    "1228639": "It seems like this should not really be a concern if you just only validate on original labeled data. I have been doing experiments with pseudolabeling and I add the pseudolabeled data into train, but then leave the validation fold as is. Training metrics look a little weird. If I really wanted to I could also show train metrics on only the truly labeled train samples, but seeing accurate training loss and auc numbers arent as important in my opinion.",
    "1232079": "too tough to boost😳",
    "1232453": "Even if the metric was accuracy, f1 score or whatever that requires class labels, isn't it better to use soft predictions as pseudo labels? The motivation is to distinguish between confident and not confident predictions."
  },
  "source": "meta"
}