{
  "id": 238504,
  "title": "43rd : Positive-Unlabeled Learning based Solution",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238504",
  "author_name": "ryomak",
  "post_date": "2021-05-12T11:29:01.175000",
  "votes": 12,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congratulations to the winners and thanks for all the competitors, great hosts!🎉<br>\nWe are members of Harada Laboratory at Tokyo University.</p>\n<p>This competition was very interesting in that classification should be done with weakly supervised labels.<br>\nAnd I thought even not high-score solution is worth to be made public.</p>\n<p><strong>Digging and saying \"here was nothing\" is also important.</strong></p>\n<p>I'll share our Positive-Unlabeled based solution.</p>\n<hr>\n<h1>Overview</h1>\n<ul>\n<li>cut images into cells, treat labels as Negative-Unlabeled then optimized AUC</li>\n<li>optimization of AUC which uses sigmoid was hard and the score didn't improved enough</li>\n<li>treating images as i.i.d. samples could induce not sufficiently tight loss</li>\n</ul>\n<p>We are members of the same laboratory. We participated this competition to search our master course theme. m0ka took care of image level classifier solution and I searched image-level classifier solution. Considering PublicLeaderBoard scores, we decided to use the image-level classifier as our final model.</p>\n<h1>Pipeline</h1>\n<ul>\n<li>ensemble of 10 ResNet34 models.</li>\n<li>AUC optimization under Negative-Unlabeled label setting</li>\n</ul>\n<p><img src=\"https://user-images.githubusercontent.com/45588624/120889912-7350a580-c63a-11eb-92a1-14f3d193b7ed.png\" alt=\"image\"></p>\n<h1>Purpose</h1>\n<ul>\n<li>train models without label noise</li>\n<li>Optimization minimize PR-AUC</li>\n<li>validation of models can be done under label noise</li>\n</ul>\n<p>We had image-level labels, but we didn't have cell-level labels. If we classify cell images in this situation, we would suffer from falsely added labels. And even after training, we can't validate our models with these noisy labels.</p>\n<p>Even under this situation, we can train our models without such bias if we use statistical machine learning technique. Below I will explain our solution.</p>\n<h1>Assumption and Setting</h1>\n<ul>\n<li>treat labels as Negative-Unlabeled.</li>\n</ul>\n<p>Let's consider cell-images with 0-17 class labels.</p>\n<p>Because these labels were added to original whole image, there are many False-Positively added labels. Then we could assume that</p>\n<ol>\n<li>Added labels can be actually negative.</li>\n<li>Not-Added labels are always negative.</li>\n</ol>\n<p>In this point of view, we can see this competition setting as Negative-Unlabeled Setting. Negative label is always negative and positive labels are always positive.</p>\n<p><strong>[Edit]</strong> After this competition, I found a paper, <a href=\"https://arxiv.org/abs/1905.12226\" target=\"_blank\">[Peng and Zhang, 2019]</a> which treats this setting😅. However, this paper's method requires class prior P(y_{k}) for the unbiased estimation, which can't be accessed. To resolve this issue and optimize AUC rather than Bayes-Risk, we introduce AUC based solution. </p>\n<h1>Optimization of AUC</h1>\n<ul>\n<li>Optimize AUC so that PR-AUC improves.</li>\n<li>we don't have to know class-prior</li>\n</ul>\n<p>In normal setting, mAP is enhanced by optimizing some losses like BCE Loss. However because labels are noisy, this loss can be hard to optimize.</p>\n<p>Rather I decided to optimize AUC which resembles ROC-AUC.</p>\n<p>Notate sample x's class i score output as fx. P is positive samples' probability and N is negative one. AUC of class i is calculated as</p>\n<p><img src=\"https://user-images.githubusercontent.com/45588624/120889964-db06f080-c63a-11eb-8975-b3882c50be9d.png\" alt=\"image\"></p>\n<p>I don't write precise theory here, but in PU Learning setting, it is known that</p>\n<ul>\n<li>using symmetric loss, which satisfies l(x) + l(-x) = const. can be reduce bias caused by Falsely added labels.</li>\n</ul>\n<p>Combining AUC optimization and symmetric loss leads object function, which we don't have to use class prior. we can notate it as</p>\n<p><img src=\"https://user-images.githubusercontent.com/45588624/120889998-0f7aac80-c63b-11eb-8a6a-5ce819864a4b.png\" alt=\"image\"></p>\n<h1>Good points &amp; Bad points</h1>\n<ul>\n<li>Good points<ul>\n<li>we can use whole images!<ul>\n<li>only 1 labeled images are limited.</li></ul></li>\n<li>there was correlation between LB and validation score relatively.<ul>\n<li>this wasn't true when I used BCE Loss and score improved +0.06pt which was large in this competition.</li></ul></li></ul></li>\n<li>Bad points<ul>\n<li>optimization of sigmoid was hard.</li>\n<li>taking bags of image apart could make loss loose.</li></ul></li>\n</ul>\n<p>After training this method, I couldn't improve our model much. I tried hard to solve issue caused by sigmoid which is difficult to opmize, ended up first 3weeks solution.</p>\n<p>8th and 9th also use image-based classifier, but they treated image-level labels as no-noise labels. This can be tight loss compared with this solution.</p>\n<h1>Others</h1>\n<p>If we could, we wanted to give some contribution to HPA community like bestfitting did in the last competition. We couldn't and he did it again, congrats!🎉</p>\n<p>We can't say our solution worked enough, but I hope you enjoyed this solution. Thanks!</p>\n<h2>Appendix</h2>\n<h3>solutions' background</h3>\n<p>positive label is described as y=1, negative is y=0.<br>\nunlike PU setting, we can't assume P(y=1) = P(y=1|s=1). So I changed some assumptions.</p>\n<p><img src=\"https://user-images.githubusercontent.com/45588624/117967161-b41d0d80-b35f-11eb-9eee-21f4bbaee9c7.png\" alt=\"image\"></p>",
  "messages": [
    {
      "id": 1304013,
      "postDate": "2021-05-12T11:29:01.177Z",
      "content": "<p>Congratulations to the winners and thanks for all the competitors, great hosts!🎉<br>\nWe are members of Harada Laboratory at Tokyo University.</p>\n<p>This competition was very interesting in that classification should be done with weakly supervised labels.<br>\nAnd I thought even not high-score solution is worth to be made public.</p>\n<p><strong>Digging and saying \"here was nothing\" is also important.</strong></p>\n<p>I'll share our Positive-Unlabeled based solution.</p>\n<hr>\n<h1>Overview</h1>\n<ul>\n<li>cut images into cells, treat labels as Negative-Unlabeled then optimized AUC</li>\n<li>optimization of AUC which uses sigmoid was hard and the score didn't improved enough</li>\n<li>treating images as i.i.d. samples could induce not sufficiently tight loss</li>\n</ul>\n<p>We are members of the same laboratory. We participated this competition to search our master course theme. m0ka took care of image level classifier solution and I searched image-level classifier solution. Considering PublicLeaderBoard scores, we decided to use the image-level classifier as our final model.</p>\n<h1>Pipeline</h1>\n<ul>\n<li>ensemble of 10 ResNet34 models.</li>\n<li>AUC optimization under Negative-Unlabeled label setting</li>\n</ul>\n<p><img src=\"https://user-images.githubusercontent.com/45588624/120889912-7350a580-c63a-11eb-92a1-14f3d193b7ed.png\" alt=\"image\"></p>\n<h1>Purpose</h1>\n<ul>\n<li>train models without label noise</li>\n<li>Optimization minimize PR-AUC</li>\n<li>validation of models can be done under label noise</li>\n</ul>\n<p>We had image-level labels, but we didn't have cell-level labels. If we classify cell images in this situation, we would suffer from falsely added labels. And even after training, we can't validate our models with these noisy labels.</p>\n<p>Even under this situation, we can train our models without such bias if we use statistical machine learning technique. Below I will explain our solution.</p>\n<h1>Assumption and Setting</h1>\n<ul>\n<li>treat labels as Negative-Unlabeled.</li>\n</ul>\n<p>Let's consider cell-images with 0-17 class labels.</p>\n<p>Because these labels were added to original whole image, there are many False-Positively added labels. Then we could assume that</p>\n<ol>\n<li>Added labels can be actually negative.</li>\n<li>Not-Added labels are always negative.</li>\n</ol>\n<p>In this point of view, we can see this competition setting as Negative-Unlabeled Setting. Negative label is always negative and positive labels are always positive.</p>\n<p><strong>[Edit]</strong> After this competition, I found a paper, <a href=\"https://arxiv.org/abs/1905.12226\" target=\"_blank\">[Peng and Zhang, 2019]</a> which treats this setting😅. However, this paper's method requires class prior P(y_{k}) for the unbiased estimation, which can't be accessed. To resolve this issue and optimize AUC rather than Bayes-Risk, we introduce AUC based solution. </p>\n<h1>Optimization of AUC</h1>\n<ul>\n<li>Optimize AUC so that PR-AUC improves.</li>\n<li>we don't have to know class-prior</li>\n</ul>\n<p>In normal setting, mAP is enhanced by optimizing some losses like BCE Loss. However because labels are noisy, this loss can be hard to optimize.</p>\n<p>Rather I decided to optimize AUC which resembles ROC-AUC.</p>\n<p>Notate sample x's class i score output as fx. P is positive samples' probability and N is negative one. AUC of class i is calculated as</p>\n<p><img src=\"https://user-images.githubusercontent.com/45588624/120889964-db06f080-c63a-11eb-8975-b3882c50be9d.png\" alt=\"image\"></p>\n<p>I don't write precise theory here, but in PU Learning setting, it is known that</p>\n<ul>\n<li>using symmetric loss, which satisfies l(x) + l(-x) = const. can be reduce bias caused by Falsely added labels.</li>\n</ul>\n<p>Combining AUC optimization and symmetric loss leads object function, which we don't have to use class prior. we can notate it as</p>\n<p><img src=\"https://user-images.githubusercontent.com/45588624/120889998-0f7aac80-c63b-11eb-8a6a-5ce819864a4b.png\" alt=\"image\"></p>\n<h1>Good points &amp; Bad points</h1>\n<ul>\n<li>Good points<ul>\n<li>we can use whole images!<ul>\n<li>only 1 labeled images are limited.</li></ul></li>\n<li>there was correlation between LB and validation score relatively.<ul>\n<li>this wasn't true when I used BCE Loss and score improved +0.06pt which was large in this competition.</li></ul></li></ul></li>\n<li>Bad points<ul>\n<li>optimization of sigmoid was hard.</li>\n<li>taking bags of image apart could make loss loose.</li></ul></li>\n</ul>\n<p>After training this method, I couldn't improve our model much. I tried hard to solve issue caused by sigmoid which is difficult to opmize, ended up first 3weeks solution.</p>\n<p>8th and 9th also use image-based classifier, but they treated image-level labels as no-noise labels. This can be tight loss compared with this solution.</p>\n<h1>Others</h1>\n<p>If we could, we wanted to give some contribution to HPA community like bestfitting did in the last competition. We couldn't and he did it again, congrats!🎉</p>\n<p>We can't say our solution worked enough, but I hope you enjoyed this solution. Thanks!</p>\n<h2>Appendix</h2>\n<h3>solutions' background</h3>\n<p>positive label is described as y=1, negative is y=0.<br>\nunlike PU setting, we can't assume P(y=1) = P(y=1|s=1). So I changed some assumptions.</p>\n<p><img src=\"https://user-images.githubusercontent.com/45588624/117967161-b41d0d80-b35f-11eb-9eee-21f4bbaee9c7.png\" alt=\"image\"></p>",
      "rawMarkdown": "Congratulations to the winners and thanks for all the competitors, great hosts!🎉\nWe are members of Harada Laboratory at Tokyo University.\n\nThis competition was very interesting in that classification should be done with weakly supervised labels.\nAnd I thought even not high-score solution is worth to be made public.\n\n**Digging and saying \"here was nothing\" is also important.**\n\nI'll share our Positive-Unlabeled based solution.\n\n---\n\n# Overview\n\n- cut images into cells, treat labels as Negative-Unlabeled then optimized AUC\n- optimization of AUC which uses sigmoid was hard and the score didn't improved enough\n- treating images as i.i.d. samples could induce not sufficiently tight loss\n\nWe are members of the same laboratory. We participated this competition to search our master course theme. m0ka took care of image level classifier solution and I searched image-level classifier solution. Considering PublicLeaderBoard scores, we decided to use the image-level classifier as our final model.\n\n# Pipeline\n\n- ensemble of 10 ResNet34 models.\n- AUC optimization under Negative-Unlabeled label setting\n\n![image](https://user-images.githubusercontent.com/45588624/120889912-7350a580-c63a-11eb-92a1-14f3d193b7ed.png)\n\n# Purpose\n\n- train models without label noise\n- Optimization minimize PR-AUC\n- validation of models can be done under label noise\n\nWe had image-level labels, but we didn't have cell-level labels. If we classify cell images in this situation, we would suffer from falsely added labels. And even after training, we can't validate our models with these noisy labels.\n\nEven under this situation, we can train our models without such bias if we use statistical machine learning technique. Below I will explain our solution.\n\n# Assumption and Setting\n\n- treat labels as Negative-Unlabeled.\n\nLet's consider cell-images with 0-17 class labels.\n\nBecause these labels were added to original whole image, there are many False-Positively added labels. Then we could assume that\n\n1. Added labels can be actually negative.\n2. Not-Added labels are always negative.\n\nIn this point of view, we can see this competition setting as Negative-Unlabeled Setting. Negative label is always negative and positive labels are always positive.\n\n**[Edit]** After this competition, I found a paper, [[Peng and Zhang, 2019]](https://arxiv.org/abs/1905.12226) which treats this setting😅. However, this paper's method requires class prior P(y_{k}) for the unbiased estimation, which can't be accessed. To resolve this issue and optimize AUC rather than Bayes-Risk, we introduce AUC based solution. \n\n# Optimization of AUC\n\n- Optimize AUC so that PR-AUC improves.\n- we don't have to know class-prior\n\nIn normal setting, mAP is enhanced by optimizing some losses like BCE Loss. However because labels are noisy, this loss can be hard to optimize.\n\nRather I decided to optimize AUC which resembles ROC-AUC.\n\nNotate sample x's class i score output as fx. P is positive samples' probability and N is negative one. AUC of class i is calculated as\n\n![image](https://user-images.githubusercontent.com/45588624/120889964-db06f080-c63a-11eb-8975-b3882c50be9d.png)\n\n\nI don't write precise theory here, but in PU Learning setting, it is known that\n\n- using symmetric loss, which satisfies l(x) + l(-x) = const. can be reduce bias caused by Falsely added labels.\n\nCombining AUC optimization and symmetric loss leads object function, which we don't have to use class prior. we can notate it as\n\n![image](https://user-images.githubusercontent.com/45588624/120889998-0f7aac80-c63b-11eb-8a6a-5ce819864a4b.png)\n\n# Good points & Bad points\n\n- Good points\n    - we can use whole images!\n        - only 1 labeled images are limited.\n    - there was correlation between LB and validation score relatively.\n        - this wasn't true when I used BCE Loss and score improved +0.06pt which was large in this competition.\n- Bad points\n    - optimization of sigmoid was hard.\n    - taking bags of image apart could make loss loose.\n\nAfter training this method, I couldn't improve our model much. I tried hard to solve issue caused by sigmoid which is difficult to opmize, ended up first 3weeks solution.\n\n8th and 9th also use image-based classifier, but they treated image-level labels as no-noise labels. This can be tight loss compared with this solution.\n\n# Others\n\nIf we could, we wanted to give some contribution to HPA community like bestfitting did in the last competition. We couldn't and he did it again, congrats!🎉\n\nWe can't say our solution worked enough, but I hope you enjoyed this solution. Thanks!\n\n## Appendix\n\n### solutions' background\n\npositive label is described as y=1, negative is y=0.\nunlike PU setting, we can't assume P(y=1) = P(y=1|s=1). So I changed some assumptions.\n\n![image](https://user-images.githubusercontent.com/45588624/117967161-b41d0d80-b35f-11eb-9eee-21f4bbaee9c7.png)",
      "votes": 12
    },
    {
      "id": 1305028,
      "postDate": "2021-05-13T04:27:02.997Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryomak\" target=\"_blank\">@ryomak</a> Congratulations  and Thanks for sharing the approach</p>",
      "rawMarkdown": "@ryomak Congratulations  and Thanks for sharing the approach",
      "votes": 1
    },
    {
      "id": 1304024,
      "postDate": "2021-05-12T11:35:45.593Z",
      "content": "<p>Thanks for the post.<br>\nHow long did it take to train all of 1 million cells for 1 epoch?<br>\nIn my case, EfficientNetB7 single fold trained with about 70000 cells led to public 0.452.<br>\nActually I was so confused how many cells we should use to train in the first place.</p>",
      "rawMarkdown": "Thanks for the post.\nHow long did it take to train all of 1 million cells for 1 epoch?\nIn my case, EfficientNetB7 single fold trained with about 70000 cells led to public 0.452.\nActually I was so confused how many cells we should use to train in the first place.",
      "replies": [
        {
          "id": 1304031,
          "postDate": "2021-05-12T11:43:20.120Z",
          "content": "<p>thanks for the comment. it took around 1.2 hour.<br>\nmachine : P100.<br>\ntotal was 2 days…</p>",
          "rawMarkdown": "thanks for the comment. it took around 1.2 hour.\nmachine : P100.\ntotal was 2 days...",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1305028,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:27:02.997000",
      "content": "<p><a href=\"https://www.kaggle.com/ryomak\" target=\"_blank\">@ryomak</a> Congratulations  and Thanks for sharing the approach</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1304024,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-05-12T11:35:45.593000",
      "content": "<p>Thanks for the post.<br>\nHow long did it take to train all of 1 million cells for 1 epoch?<br>\nIn my case, EfficientNetB7 single fold trained with about 70000 cells led to public 0.452.<br>\nActually I was so confused how many cells we should use to train in the first place.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1304031,
          "author_name": "ryomak",
          "author_url": "",
          "post_date": "2021-05-12T11:43:20.120000",
          "content": "<p>thanks for the comment. it took around 1.2 hour.<br>\nmachine : P100.<br>\ntotal was 2 days…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1304013": "Congratulations to the winners and thanks for all the competitors, great hosts!🎉\nWe are members of Harada Laboratory at Tokyo University.\n\nThis competition was very interesting in that classification should be done with weakly supervised labels.\nAnd I thought even not high-score solution is worth to be made public.\n\n**Digging and saying \"here was nothing\" is also important.**\n\nI'll share our Positive-Unlabeled based solution.\n\n---\n\n# Overview\n\n- cut images into cells, treat labels as Negative-Unlabeled then optimized AUC\n- optimization of AUC which uses sigmoid was hard and the score didn't improved enough\n- treating images as i.i.d. samples could induce not sufficiently tight loss\n\nWe are members of the same laboratory. We participated this competition to search our master course theme. m0ka took care of image level classifier solution and I searched image-level classifier solution. Considering PublicLeaderBoard scores, we decided to use the image-level classifier as our final model.\n\n# Pipeline\n\n- ensemble of 10 ResNet34 models.\n- AUC optimization under Negative-Unlabeled label setting\n\n![image](https://user-images.githubusercontent.com/45588624/120889912-7350a580-c63a-11eb-92a1-14f3d193b7ed.png)\n\n# Purpose\n\n- train models without label noise\n- Optimization minimize PR-AUC\n- validation of models can be done under label noise\n\nWe had image-level labels, but we didn't have cell-level labels. If we classify cell images in this situation, we would suffer from falsely added labels. And even after training, we can't validate our models with these noisy labels.\n\nEven under this situation, we can train our models without such bias if we use statistical machine learning technique. Below I will explain our solution.\n\n# Assumption and Setting\n\n- treat labels as Negative-Unlabeled.\n\nLet's consider cell-images with 0-17 class labels.\n\nBecause these labels were added to original whole image, there are many False-Positively added labels. Then we could assume that\n\n1. Added labels can be actually negative.\n2. Not-Added labels are always negative.\n\nIn this point of view, we can see this competition setting as Negative-Unlabeled Setting. Negative label is always negative and positive labels are always positive.\n\n**[Edit]** After this competition, I found a paper, [[Peng and Zhang, 2019]](https://arxiv.org/abs/1905.12226) which treats this setting😅. However, this paper's method requires class prior P(y_{k}) for the unbiased estimation, which can't be accessed. To resolve this issue and optimize AUC rather than Bayes-Risk, we introduce AUC based solution. \n\n# Optimization of AUC\n\n- Optimize AUC so that PR-AUC improves.\n- we don't have to know class-prior\n\nIn normal setting, mAP is enhanced by optimizing some losses like BCE Loss. However because labels are noisy, this loss can be hard to optimize.\n\nRather I decided to optimize AUC which resembles ROC-AUC.\n\nNotate sample x's class i score output as fx. P is positive samples' probability and N is negative one. AUC of class i is calculated as\n\n![image](https://user-images.githubusercontent.com/45588624/120889964-db06f080-c63a-11eb-8975-b3882c50be9d.png)\n\n\nI don't write precise theory here, but in PU Learning setting, it is known that\n\n- using symmetric loss, which satisfies l(x) + l(-x) = const. can be reduce bias caused by Falsely added labels.\n\nCombining AUC optimization and symmetric loss leads object function, which we don't have to use class prior. we can notate it as\n\n![image](https://user-images.githubusercontent.com/45588624/120889998-0f7aac80-c63b-11eb-8a6a-5ce819864a4b.png)\n\n# Good points & Bad points\n\n- Good points\n    - we can use whole images!\n        - only 1 labeled images are limited.\n    - there was correlation between LB and validation score relatively.\n        - this wasn't true when I used BCE Loss and score improved +0.06pt which was large in this competition.\n- Bad points\n    - optimization of sigmoid was hard.\n    - taking bags of image apart could make loss loose.\n\nAfter training this method, I couldn't improve our model much. I tried hard to solve issue caused by sigmoid which is difficult to opmize, ended up first 3weeks solution.\n\n8th and 9th also use image-based classifier, but they treated image-level labels as no-noise labels. This can be tight loss compared with this solution.\n\n# Others\n\nIf we could, we wanted to give some contribution to HPA community like bestfitting did in the last competition. We couldn't and he did it again, congrats!🎉\n\nWe can't say our solution worked enough, but I hope you enjoyed this solution. Thanks!\n\n## Appendix\n\n### solutions' background\n\npositive label is described as y=1, negative is y=0.\nunlike PU setting, we can't assume P(y=1) = P(y=1|s=1). So I changed some assumptions.\n\n![image](https://user-images.githubusercontent.com/45588624/117967161-b41d0d80-b35f-11eb-9eee-21f4bbaee9c7.png)",
    "1305028": "@ryomak Congratulations  and Thanks for sharing the approach",
    "1304024": "Thanks for the post.\nHow long did it take to train all of 1 million cells for 1 epoch?\nIn my case, EfficientNetB7 single fold trained with about 70000 cells led to public 0.452.\nActually I was so confused how many cells we should use to train in the first place."
  }
}