{
  "id": 238729,
  "title": "Meta-summary of the Top 30",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238729",
  "author_name": "Andrew Tratz",
  "post_date": "2021-05-13T08:01:09.948000",
  "votes": 36,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I'm fascinated by the incredible diversity of approaches among the top-performing teams, particularly as compared to many competitions where a single dominant approach emerges.</p>\n<p>I've attempted to summarize the main similarities and differences among the top teams' approaches, hoping to distill some of the common themes which emerged.</p>\n<p>I need to apologize in advance for oversimplifying, mischaracterizing, or otherwise \"missing the point\" of some of the teams' ideas. Some of the techniques are new to me as well and the way I've clustered the themes may or not make sense. I'd welcome any feedback, suggestions, or corrections that any teams might be willing to offer.</p>\n<p>I'm also limited by the descriptions provided by the teams. Blank cells mean that the team did not mention this in their approach summary, and I broadly assume this means they did not adopt a particular approach.</p>\n<p>With these caveats, hope you will find this useful:</p>\n<p><img src=\"https://raw.githubusercontent.com/pinky1812/images/main/summary9.png\" alt=\"https://raw.githubusercontent.com/pinky1812/images/main/summary9.png\"></p>",
  "messages": [
    {
      "id": 1305326,
      "postDate": "2021-05-13T08:01:09.947Z",
      "content": "<p>I'm fascinated by the incredible diversity of approaches among the top-performing teams, particularly as compared to many competitions where a single dominant approach emerges.</p>\n<p>I've attempted to summarize the main similarities and differences among the top teams' approaches, hoping to distill some of the common themes which emerged.</p>\n<p>I need to apologize in advance for oversimplifying, mischaracterizing, or otherwise \"missing the point\" of some of the teams' ideas. Some of the techniques are new to me as well and the way I've clustered the themes may or not make sense. I'd welcome any feedback, suggestions, or corrections that any teams might be willing to offer.</p>\n<p>I'm also limited by the descriptions provided by the teams. Blank cells mean that the team did not mention this in their approach summary, and I broadly assume this means they did not adopt a particular approach.</p>\n<p>With these caveats, hope you will find this useful:</p>\n<p><img src=\"https://raw.githubusercontent.com/pinky1812/images/main/summary9.png\" alt=\"https://raw.githubusercontent.com/pinky1812/images/main/summary9.png\"></p>",
      "rawMarkdown": "I'm fascinated by the incredible diversity of approaches among the top-performing teams, particularly as compared to many competitions where a single dominant approach emerges.\n\nI've attempted to summarize the main similarities and differences among the top teams' approaches, hoping to distill some of the common themes which emerged.\n\nI need to apologize in advance for oversimplifying, mischaracterizing, or otherwise \"missing the point\" of some of the teams' ideas. Some of the techniques are new to me as well and the way I've clustered the themes may or not make sense. I'd welcome any feedback, suggestions, or corrections that any teams might be willing to offer.\n\nI'm also limited by the descriptions provided by the teams. Blank cells mean that the team did not mention this in their approach summary, and I broadly assume this means they did not adopt a particular approach.\n\nWith these caveats, hope you will find this useful:\n\n![https://raw.githubusercontent.com/pinky1812/images/main/summary9.png](https://raw.githubusercontent.com/pinky1812/images/main/summary9.png)\n",
      "votes": 36
    },
    {
      "id": 1305374,
      "postDate": "2021-05-13T08:28:29.630Z",
      "content": "<p>Thank you for making this. Very helpful. Please keep updating this.</p>",
      "rawMarkdown": "Thank you for making this. Very helpful. Please keep updating this.",
      "votes": 3
    },
    {
      "id": 1306515,
      "postDate": "2021-05-13T20:39:04.883Z",
      "content": "<p>Great summery! It's a tough job to make a comparison considering all the different approaches. But you did a nice job to look at the solutions and categorizing the flavors! This is useful representation of collective knowledge.</p>",
      "rawMarkdown": "Great summery! It's a tough job to make a comparison considering all the different approaches. But you did a nice job to look at the solutions and categorizing the flavors! This is useful representation of collective knowledge.",
      "votes": 2
    },
    {
      "id": 1322693,
      "postDate": "2021-05-25T15:48:37.853Z",
      "content": "<p>This is super cool! Thank you for summarizing things!</p>",
      "rawMarkdown": "This is super cool! Thank you for summarizing things!"
    },
    {
      "id": 1308056,
      "postDate": "2021-05-14T23:36:44.070Z",
      "content": "<p>That's great summary to view the solutions, thanks!</p>",
      "rawMarkdown": "That's great summary to view the solutions, thanks!"
    },
    {
      "id": 1306746,
      "postDate": "2021-05-14T03:59:33.397Z",
      "content": "<p>Very useful summary! Thank you for making it!</p>",
      "rawMarkdown": "Very useful summary! Thank you for making it!"
    },
    {
      "id": 1305682,
      "postDate": "2021-05-13T12:29:17.563Z",
      "content": "<p>Hi Andrew! Thank you a lot for putting it together, it's very interesting to analyze!</p>\n<p>Just wanted to let you know that I used focal loss, as well as the hard-negatives mining loss and the Lovasz hinge loss, in the same way as <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> did in his solution to the first HPA challenge.</p>\n<p>Also, I'm not sure if <code>metric modeling</code> is the right term for the label noise reduction I used, but I'm no expert, so please feel free to correct me. I didn't model the metric (it's definitely not about the metric learning), I rather \"extracted\" the metric based on the last layer embedding of the classifier. The more important part was de-noising on top of the knn graph constructed using distances between those extracted embeddings.</p>\n<p>In the paper <a href=\"https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y\" target=\"_blank\"><strong>Learning from Weak and Noisy Labels for Semantic Segmentation</strong></a>, which I implemented the algorithm from, the authors described the essence of their approach as follows </p>\n<blockquote>\n  <p> to identify and correct the noisy labels,  a L1-optimisation based sparse learning model is formulated</p>\n</blockquote>\n<p>I believe that something like <code>mathematical optimization for label noise reduction</code> grasps the main theme of the approach better. What do you think? :) </p>",
      "rawMarkdown": "Hi Andrew! Thank you a lot for putting it together, it's very interesting to analyze!\n\nJust wanted to let you know that I used focal loss, as well as the hard-negatives mining loss and the Lovasz hinge loss, in the same way as @bestfitting did in his solution to the first HPA challenge.\n\nAlso, I'm not sure if `metric modeling` is the right term for the label noise reduction I used, but I'm no expert, so please feel free to correct me. I didn't model the metric (it's definitely not about the metric learning), I rather \"extracted\" the metric based on the last layer embedding of the classifier. The more important part was de-noising on top of the knn graph constructed using distances between those extracted embeddings.\n\nIn the paper [**Learning from Weak and Noisy Labels for Semantic Segmentation**](https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&isAllowed=y), which I implemented the algorithm from, the authors described the essence of their approach as follows \n>  <..> to identify and correct the noisy labels, <..> a L1-optimisation based sparse learning model is formulated\n\nI believe that something like `mathematical optimization for label noise reduction` grasps the main theme of the approach better. What do you think? :) ",
      "replies": [
        {
          "id": 1305760,
          "postDate": "2021-05-13T13:13:47.073Z",
          "content": "<p>Thanks - I've updated to try and reflect this more accurately.</p>",
          "rawMarkdown": "Thanks - I've updated to try and reflect this more accurately.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1305538,
      "postDate": "2021-05-13T10:48:50.303Z",
      "content": "<p>I remember there was a post in this competition saying that Kaggle competitions are becoming harder. But, its also true that competitors are getting smarter and better. Congrats to all the winners, and thanks for compiling this.</p>",
      "rawMarkdown": "I remember there was a post in this competition saying that Kaggle competitions are becoming harder. But, its also true that competitors are getting smarter and better. Congrats to all the winners, and thanks for compiling this."
    },
    {
      "id": 1305446,
      "postDate": "2021-05-13T09:28:51.067Z",
      "content": "<p>Thank you so much for the summary. I'm 22nd rabbit 🐇<br>\nI did manual-labeling of class \"11\" if their pseudolabel &gt;= 0.3.<br>\n(There are sometimes \"11\" cells inside the image without \"11\" image level label.)</p>",
      "rawMarkdown": "Thank you so much for the summary. I'm 22nd rabbit 🐇\nI did manual-labeling of class \"11\" if their pseudolabel >= 0.3.\n(There are sometimes \"11\" cells inside the image without \"11\" image level label.)",
      "replies": [
        {
          "id": 1305499,
          "postDate": "2021-05-13T10:15:36.887Z",
          "content": "<p>Updated - thank you.</p>\n<p>I also found many mitotic spindles in cells without the label. However I was worried about reclassifying them since many didn't have strong activation in the green channel. (Actually, the same could be argued for many examples which were labeled by the annotators). I started creating a dataset with these examples in case my model started accidentally classifying \"mitotic negatives\" as positives. I was worried that my model would base its mitotic predictions off of the blue or red channel rather than the green, but this problem didn't seem to be a significant one after all.</p>",
          "rawMarkdown": "Updated - thank you.\n\nI also found many mitotic spindles in cells without the label. However I was worried about reclassifying them since many didn't have strong activation in the green channel. (Actually, the same could be argued for many examples which were labeled by the annotators). I started creating a dataset with these examples in case my model started accidentally classifying \"mitotic negatives\" as positives. I was worried that my model would base its mitotic predictions off of the blue or red channel rather than the green, but this problem didn't seem to be a significant one after all.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1305374,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-05-13T08:28:29.630000",
      "content": "<p>Thank you for making this. Very helpful. Please keep updating this.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1306515,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2021-05-13T20:39:04.883000",
      "content": "<p>Great summery! It's a tough job to make a comparison considering all the different approaches. But you did a nice job to look at the solutions and categorizing the flavors! This is useful representation of collective knowledge.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1322693,
      "author_name": "Casper Winsnes",
      "author_url": "",
      "post_date": "2021-05-25T15:48:37.853000",
      "content": "<p>This is super cool! Thank you for summarizing things!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1308056,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-14T23:36:44.070000",
      "content": "<p>That's great summary to view the solutions, thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1306746,
      "author_name": "KaizaburoChubachi",
      "author_url": "",
      "post_date": "2021-05-14T03:59:33.397000",
      "content": "<p>Very useful summary! Thank you for making it!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1305682,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-13T12:29:17.563000",
      "content": "<p>Hi Andrew! Thank you a lot for putting it together, it's very interesting to analyze!</p>\n<p>Just wanted to let you know that I used focal loss, as well as the hard-negatives mining loss and the Lovasz hinge loss, in the same way as <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> did in his solution to the first HPA challenge.</p>\n<p>Also, I'm not sure if <code>metric modeling</code> is the right term for the label noise reduction I used, but I'm no expert, so please feel free to correct me. I didn't model the metric (it's definitely not about the metric learning), I rather \"extracted\" the metric based on the last layer embedding of the classifier. The more important part was de-noising on top of the knn graph constructed using distances between those extracted embeddings.</p>\n<p>In the paper <a href=\"https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y\" target=\"_blank\"><strong>Learning from Weak and Noisy Labels for Semantic Segmentation</strong></a>, which I implemented the algorithm from, the authors described the essence of their approach as follows </p>\n<blockquote>\n  <p> to identify and correct the noisy labels,  a L1-optimisation based sparse learning model is formulated</p>\n</blockquote>\n<p>I believe that something like <code>mathematical optimization for label noise reduction</code> grasps the main theme of the approach better. What do you think? :) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1305760,
          "author_name": "Andrew Tratz",
          "author_url": "",
          "post_date": "2021-05-13T13:13:47.073000",
          "content": "<p>Thanks - I've updated to try and reflect this more accurately.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1305538,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "2021-05-13T10:48:50.303000",
      "content": "<p>I remember there was a post in this competition saying that Kaggle competitions are becoming harder. But, its also true that competitors are getting smarter and better. Congrats to all the winners, and thanks for compiling this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1305446,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-05-13T09:28:51.067000",
      "content": "<p>Thank you so much for the summary. I'm 22nd rabbit 🐇<br>\nI did manual-labeling of class \"11\" if their pseudolabel &gt;= 0.3.<br>\n(There are sometimes \"11\" cells inside the image without \"11\" image level label.)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1305499,
          "author_name": "Andrew Tratz",
          "author_url": "",
          "post_date": "2021-05-13T10:15:36.887000",
          "content": "<p>Updated - thank you.</p>\n<p>I also found many mitotic spindles in cells without the label. However I was worried about reclassifying them since many didn't have strong activation in the green channel. (Actually, the same could be argued for many examples which were labeled by the annotators). I started creating a dataset with these examples in case my model started accidentally classifying \"mitotic negatives\" as positives. I was worried that my model would base its mitotic predictions off of the blue or red channel rather than the green, but this problem didn't seem to be a significant one after all.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1305326": "I'm fascinated by the incredible diversity of approaches among the top-performing teams, particularly as compared to many competitions where a single dominant approach emerges.\n\nI've attempted to summarize the main similarities and differences among the top teams' approaches, hoping to distill some of the common themes which emerged.\n\nI need to apologize in advance for oversimplifying, mischaracterizing, or otherwise \"missing the point\" of some of the teams' ideas. Some of the techniques are new to me as well and the way I've clustered the themes may or not make sense. I'd welcome any feedback, suggestions, or corrections that any teams might be willing to offer.\n\nI'm also limited by the descriptions provided by the teams. Blank cells mean that the team did not mention this in their approach summary, and I broadly assume this means they did not adopt a particular approach.\n\nWith these caveats, hope you will find this useful:\n\n![https://raw.githubusercontent.com/pinky1812/images/main/summary9.png](https://raw.githubusercontent.com/pinky1812/images/main/summary9.png)\n",
    "1305374": "Thank you for making this. Very helpful. Please keep updating this.",
    "1306515": "Great summery! It's a tough job to make a comparison considering all the different approaches. But you did a nice job to look at the solutions and categorizing the flavors! This is useful representation of collective knowledge.",
    "1322693": "This is super cool! Thank you for summarizing things!",
    "1308056": "That's great summary to view the solutions, thanks!",
    "1306746": "Very useful summary! Thank you for making it!",
    "1305682": "Hi Andrew! Thank you a lot for putting it together, it's very interesting to analyze!\n\nJust wanted to let you know that I used focal loss, as well as the hard-negatives mining loss and the Lovasz hinge loss, in the same way as @bestfitting did in his solution to the first HPA challenge.\n\nAlso, I'm not sure if `metric modeling` is the right term for the label noise reduction I used, but I'm no expert, so please feel free to correct me. I didn't model the metric (it's definitely not about the metric learning), I rather \"extracted\" the metric based on the last layer embedding of the classifier. The more important part was de-noising on top of the knn graph constructed using distances between those extracted embeddings.\n\nIn the paper [**Learning from Weak and Noisy Labels for Semantic Segmentation**](https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&isAllowed=y), which I implemented the algorithm from, the authors described the essence of their approach as follows \n>  <..> to identify and correct the noisy labels, <..> a L1-optimisation based sparse learning model is formulated\n\nI believe that something like `mathematical optimization for label noise reduction` grasps the main theme of the approach better. What do you think? :) ",
    "1305538": "I remember there was a post in this competition saying that Kaggle competitions are becoming harder. But, its also true that competitors are getting smarter and better. Congrats to all the winners, and thanks for compiling this.",
    "1305446": "Thank you so much for the summary. I'm 22nd rabbit 🐇\nI did manual-labeling of class \"11\" if their pseudolabel >= 0.3.\n(There are sometimes \"11\" cells inside the image without \"11\" image level label.)"
  }
}