{
  "id": 263665,
  "title": "why joint classification and detection don't work?",
  "url": "/competitions/siim-covid19-detection/discussion/263665",
  "author_name": "",
  "post_date": "2021-08-10T01:13:31.842555700Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I think many teams will have realized the following down work:</p>\n<ul>\n<li>joint task of 4 study classification classes and 1 opacity detection class  </li>\n<li>joint task of 1 none classification class and 1 opacity detection class  </li>\n</ul>\n<p>I was wondering why it doesn't work?</p>\n<p>but if you use classification + segmentation (as aux loss), it works very well</p>\n<p>Anyone has an explanation or did experiments on this?<br>\nis it because:</p>\n<ul>\n<li>lack of train data</li>\n<li>the label themselves are ambiguous</li>\n<li>metric issues (i.e. if you use accuracy metric, results may improve)</li>\n<li>modeling issue: global pooling effect/back propagation of classifier head? </li>\n</ul>\n<p>by the way, did anyone notice that if you train a \"single none vs non-none class binary\" classifier, you have map for non in the range of 0.80, while non-none in the range of high 0.95</p>",
  "messages": [
    {
      "id": "1462703",
      "postDate": "08/10/2021 01:13:31",
      "content": "<p>I think many teams will have realized the following down work:</p>\n<ul>\n<li>joint task of 4 study classification classes and 1 opacity detection class  </li>\n<li>joint task of 1 none classification class and 1 opacity detection class  </li>\n</ul>\n<p>I was wondering why it doesn't work?</p>\n<p>but if you use classification + segmentation (as aux loss), it works very well</p>\n<p>Anyone has an explanation or did experiments on this?<br>\nis it because:</p>\n<ul>\n<li>lack of train data</li>\n<li>the label themselves are ambiguous</li>\n<li>metric issues (i.e. if you use accuracy metric, results may improve)</li>\n<li>modeling issue: global pooling effect/back propagation of classifier head? </li>\n</ul>\n<p>by the way, did anyone notice that if you train a \"single none vs non-none class binary\" classifier, you have map for non in the range of 0.80, while non-none in the range of high 0.95</p>",
      "rawMarkdown": "I think many teams will have realized the following down work:\n- joint task of 4 study classification classes and 1 opacity detection class  \n- joint task of 1 none classification class and 1 opacity detection class  \n\nI was wondering why it doesn't work?\n\nbut if you use classification + segmentation (as aux loss), it works very well\n\nAnyone has an explanation or did experiments on this?\nis it because:\n- lack of train data\n- the label themselves are ambiguous\n- metric issues (i.e. if you use accuracy metric, results may improve)\n- modeling issue: global pooling effect/back propagation of classifier head? \n\nby the way, did anyone notice that if you train a \"single none vs non-none class binary\" classifier, you have map for non in the range of 0.80, while non-none in the range of high 0.95",
      "votes": null
    },
    {
      "id": "1462844",
      "postDate": "08/10/2021 02:35:44",
      "content": "<p>It is mentioned here:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263683\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/263683</a></p>\n<p>\"To predict all the six targets, the whole image was used as bbox for the four study level labels,\"<br>\n\"My 5-fold efficientdet-D5 achieved public LB 0.622 and private LB 0.618.\"</p>\n<p>so it is possible to work  …</p>\n<p>note that if you use a bounding box for the whole image, you do not average pool as you whole for classifier model in one stage detection network like efficientnet.</p>\n<p>for two-stage network like faster RCNN, the classifier head use pooling.</p>",
      "rawMarkdown": "It is mentioned here:\n\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/263683\n\n\"To predict all the six targets, the whole image was used as bbox for the four study level labels,\"\n\"My 5-fold efficientdet-D5 achieved public LB 0.622 and private LB 0.618.\"\n\nso it is possible to work  ...\n\nnote that if you use a bounding box for the whole image, you do not average pool as you whole for classifier model in one stage detection network like efficientnet.\n\nfor two-stage network like faster RCNN, the classifier head use pooling.",
      "votes": null
    },
    {
      "id": "1463345",
      "postDate": "08/10/2021 07:04:00",
      "content": "<p><a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> Made it work -&gt; <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263683\" target=\"_blank\">link</a><br>\nBut he's kaggle #1 (or #2 idk) so it's different :)</p>",
      "rawMarkdown": "wowfattie Made it work -> [link](https://www.kaggle.com/c/siim-covid19-detection/discussion/263683)\nBut he's kaggle #1 (or #2 idk) so it's different :)",
      "votes": null
    },
    {
      "id": "1463452",
      "postDate": "08/10/2021 07:39:50",
      "content": "<p>I think it works quite well, you just have to have a good pipeline, we only tried it with efficientdet though and there it worked well.</p>",
      "rawMarkdown": "I think it works quite well, you just have to have a good pipeline, we only tried it with efficientdet though and there it worked well.",
      "votes": null
    },
    {
      "id": "1463566",
      "postDate": "08/10/2021 08:34:31",
      "content": "<p>i tried it. In my experiments, I can get some decent results. But the performance is always worse (MAP about 0.05 worse) than using the separate models.</p>\n<p>do you get better results?</p>",
      "rawMarkdown": "i tried it. In my experiments, I can get some decent results. But the performance is always worse (MAP about 0.05 worse) than using the separate models.\n\ndo you get better results?",
      "votes": null
    },
    {
      "id": "1463578",
      "postDate": "08/10/2021 08:39:13",
      "content": "<p>Not better, but similar. Actually for effdet detection it was better to train it combined. But it all depends on your individual tuning, this problem is very sensitive to over/underfit.</p>",
      "rawMarkdown": "Not better, but similar. Actually for effdet detection it was better to train it combined. But it all depends on your individual tuning, this problem is very sensitive to over/underfit.",
      "votes": null
    },
    {
      "id": "1464414",
      "postDate": "08/10/2021 15:05:21",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  I don't understand how did you guys go to <strong>1323</strong> place? You guys had a good score on <strong>public-test</strong> . Perhaps you guys didn't infer on the <strong>private-test</strong> ? I wonder why. You guys had a great team…</p>",
      "rawMarkdown": "hengck23  I don't understand how did you guys go to **1323** place? You guys had a good score on **public-test** . Perhaps you guys didn't infer on the **private-test** ? I wonder why. You guys had a great team...",
      "votes": null
    },
    {
      "id": "1465193",
      "postDate": "08/10/2021 23:29:55",
      "content": "<p>Regarding the last thing you mentioned, mAP tends to be lower for sparse classes. The class none is observed much less frequent than not-none. It is reasonable to have these results then. </p>",
      "rawMarkdown": "Regarding the last thing you mentioned, mAP tends to be lower for sparse classes. The class none is observed much less frequent than not-none. It is reasonable to have these results then.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1462844,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/10/2021 02:35:44",
      "content": "<p>It is mentioned here:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263683\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/263683</a></p>\n<p>\"To predict all the six targets, the whole image was used as bbox for the four study level labels,\"<br>\n\"My 5-fold efficientdet-D5 achieved public LB 0.622 and private LB 0.618.\"</p>\n<p>so it is possible to work  …</p>\n<p>note that if you use a bounding box for the whole image, you do not average pool as you whole for classifier model in one stage detection network like efficientnet.</p>\n<p>for two-stage network like faster RCNN, the classifier head use pooling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1463345,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "08/10/2021 07:04:00",
      "content": "<p><a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> Made it work -&gt; <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263683\" target=\"_blank\">link</a><br>\nBut he's kaggle #1 (or #2 idk) so it's different :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1463452,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "08/10/2021 07:39:50",
      "content": "<p>I think it works quite well, you just have to have a good pipeline, we only tried it with efficientdet though and there it worked well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1463566,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/10/2021 08:34:31",
          "content": "<p>i tried it. In my experiments, I can get some decent results. But the performance is always worse (MAP about 0.05 worse) than using the separate models.</p>\n<p>do you get better results?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1463578,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/10/2021 08:39:13",
          "content": "<p>Not better, but similar. Actually for effdet detection it was better to train it combined. But it all depends on your individual tuning, this problem is very sensitive to over/underfit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1464414,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "08/10/2021 15:05:21",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  I don't understand how did you guys go to <strong>1323</strong> place? You guys had a good score on <strong>public-test</strong> . Perhaps you guys didn't infer on the <strong>private-test</strong> ? I wonder why. You guys had a great team…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1465193,
      "author_name": "mikecho",
      "author_url": "",
      "post_date": "08/10/2021 23:29:55",
      "content": "<p>Regarding the last thing you mentioned, mAP tends to be lower for sparse classes. The class none is observed much less frequent than not-none. It is reasonable to have these results then. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1462703": "I think many teams will have realized the following down work:\n- joint task of 4 study classification classes and 1 opacity detection class  \n- joint task of 1 none classification class and 1 opacity detection class  \n\nI was wondering why it doesn't work?\n\nbut if you use classification + segmentation (as aux loss), it works very well\n\nAnyone has an explanation or did experiments on this?\nis it because:\n- lack of train data\n- the label themselves are ambiguous\n- metric issues (i.e. if you use accuracy metric, results may improve)\n- modeling issue: global pooling effect/back propagation of classifier head? \n\nby the way, did anyone notice that if you train a \"single none vs non-none class binary\" classifier, you have map for non in the range of 0.80, while non-none in the range of high 0.95",
    "1462844": "It is mentioned here:\n\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/263683\n\n\"To predict all the six targets, the whole image was used as bbox for the four study level labels,\"\n\"My 5-fold efficientdet-D5 achieved public LB 0.622 and private LB 0.618.\"\n\nso it is possible to work  ...\n\nnote that if you use a bounding box for the whole image, you do not average pool as you whole for classifier model in one stage detection network like efficientnet.\n\nfor two-stage network like faster RCNN, the classifier head use pooling.",
    "1463345": "wowfattie Made it work -> [link](https://www.kaggle.com/c/siim-covid19-detection/discussion/263683)\nBut he's kaggle #1 (or #2 idk) so it's different :)",
    "1463452": "I think it works quite well, you just have to have a good pipeline, we only tried it with efficientdet though and there it worked well.",
    "1463566": "i tried it. In my experiments, I can get some decent results. But the performance is always worse (MAP about 0.05 worse) than using the separate models.\n\ndo you get better results?",
    "1463578": "Not better, but similar. Actually for effdet detection it was better to train it combined. But it all depends on your individual tuning, this problem is very sensitive to over/underfit.",
    "1464414": "hengck23  I don't understand how did you guys go to **1323** place? You guys had a good score on **public-test** . Perhaps you guys didn't infer on the **private-test** ? I wonder why. You guys had a great team...",
    "1465193": "Regarding the last thing you mentioned, mAP tends to be lower for sparse classes. The class none is observed much less frequent than not-none. It is reasonable to have these results then."
  },
  "source": "meta"
}