{
  "id": 295724,
  "title": "Huge difference among class-wise APs",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/295724",
  "author_name": "",
  "post_date": "2021-12-17T13:14:29.125429200Z",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I submitted my results with each prediction belonging to each class only, and I found there is a huge difference among class-wise APs on the LB.</p>\n<p>cort: 0.25<br>\nshsy5y: 0.038<br>\nastro : 0.034</p>\n<p>I guess that portion of cort images is large, but the scores are still weird to me. Have anyone tried to check class-wise on LB and observed the same score portion to me. </p>",
  "messages": [
    {
      "id": "1621123",
      "postDate": "12/17/2021 13:14:29",
      "content": "<p>I submitted my results with each prediction belonging to each class only, and I found there is a huge difference among class-wise APs on the LB.</p>\n<p>cort: 0.25<br>\nshsy5y: 0.038<br>\nastro : 0.034</p>\n<p>I guess that portion of cort images is large, but the scores are still weird to me. Have anyone tried to check class-wise on LB and observed the same score portion to me. </p>",
      "rawMarkdown": "I submitted my results with each prediction belonging to each class only, and I found there is a huge difference among class-wise APs on the LB.\n\ncort: 0.25\nshsy5y: 0.038\nastro : 0.034\n\nI guess that portion of cort images is large, but the scores are still weird to me. Have anyone tried to check class-wise on LB and observed the same score portion to me.",
      "votes": null
    },
    {
      "id": "1621218",
      "postDate": "12/17/2021 14:39:38",
      "content": "<p>How do you eval dataset? mAP code did you use?</p>",
      "rawMarkdown": "How do you eval dataset? mAP code did you use?",
      "votes": null
    },
    {
      "id": "1621272",
      "postDate": "12/17/2021 15:34:51",
      "content": "<p>I just submit the notebook to LB with predictions belonging to single class only, while predictions belonging to others classes are ignored. Repeat that for 3 classes. </p>",
      "rawMarkdown": "I just submit the notebook to LB with predictions belonging to single class only, while predictions belonging to others classes are ignored. Repeat that for 3 classes.",
      "votes": null
    },
    {
      "id": "1621314",
      "postDate": "12/17/2021 16:14:52",
      "content": "<p>How are you deciding which test file is which type? Maybe your classifier returns cort for most of the test files :)</p>",
      "rawMarkdown": "How are you deciding which test file is which type? Maybe your classifier returns cort for most of the test files :)",
      "votes": null
    },
    {
      "id": "1621629",
      "postDate": "12/17/2021 22:01:51",
      "content": "<p>Even if the numbers look odd at first, they are correct.</p>\n<p>When you calculate the score locally per category, you only average across the images of that category. However, when you submit the masks for a single category in Kaggle, the score is averaged across all images of all categories which gives you a much lower score. As an example, let's say the astro category makes up 25% of all images. If you only submit masks for astro and your model gives a perfect score of 1.0, you would only get a score of 0.25 when submitting to Kaggle since there are additional 75% of images with score 0. In order to get the real category score, you would have to multiply the score from Kaggle by 4 (due to astro making up 25% of the images).</p>\n<p>Also, if you sum up your individual submitted category scores, you should get your overall LB score. It doesn't fully match your current LB score but I think it's close enough ;-) Deviations probably come from some misclassification on individual images.</p>",
      "rawMarkdown": "Even if the numbers look odd at first, they are correct.\n\nWhen you calculate the score locally per category, you only average across the images of that category. However, when you submit the masks for a single category in Kaggle, the score is averaged across all images of all categories which gives you a much lower score. As an example, let's say the astro category makes up 25% of all images. If you only submit masks for astro and your model gives a perfect score of 1.0, you would only get a score of 0.25 when submitting to Kaggle since there are additional 75% of images with score 0. In order to get the real category score, you would have to multiply the score from Kaggle by 4 (due to astro making up 25% of the images).\n\nAlso, if you sum up your individual submitted category scores, you should get your overall LB score. It doesn't fully match your current LB score but I think it's close enough ;-) Deviations probably come from some misclassification on individual images.",
      "votes": null
    },
    {
      "id": "1621712",
      "postDate": "12/18/2021 01:36:11",
      "content": "<p>I dont decide at instance level. Keep only instance belonging to a target class.</p>",
      "rawMarkdown": "I dont decide at instance level. Keep only instance belonging to a target class.",
      "votes": null
    },
    {
      "id": "1621713",
      "postDate": "12/18/2021 01:38:47",
      "content": "<p>I got it. Thank you very much. </p>",
      "rawMarkdown": "I got it. Thank you very much.",
      "votes": null
    },
    {
      "id": "1623033",
      "postDate": "12/19/2021 12:10:58",
      "content": "<p>My mAP for each class is weird, too.</p>\n<p>Detectron2 R50</p>\n<p>Astro got high and cort is lower.</p>\n<p>[0mcopypaste: MaP IoU=0.28160431798918684<br>\n[0mcopypaste: mAP cort=0.21041490919401545<br>\n[0mcopypaste: mAP shsy5y=0.16366705343985546<br>\n[0mcopypaste: mAP astro=0.3572197447479763</p>",
      "rawMarkdown": "My mAP for each class is weird, too.\n\nDetectron2 R50\n\nAstro got high and cort is lower.\n\n[0mcopypaste: MaP IoU=0.28160431798918684\n[0mcopypaste: mAP cort=0.21041490919401545\n[0mcopypaste: mAP shsy5y=0.16366705343985546\n[0mcopypaste: mAP astro=0.3572197447479763",
      "votes": null
    },
    {
      "id": "1623038",
      "postDate": "12/19/2021 12:19:52",
      "content": "<p>Did you double check whether you might be mixing up the category indexes and their names? The raw numbers would look reasonable to me if they were for the categories shsy5y, astro, and cort (in that order).</p>",
      "rawMarkdown": "Did you double check whether you might be mixing up the category indexes and their names? The raw numbers would look reasonable to me if they were for the categories shsy5y, astro, and cort (in that order).",
      "votes": null
    },
    {
      "id": "1623277",
      "postDate": "12/19/2021 17:17:48",
      "content": "<p>Ah, Thanks. I mess up the index. It should be shsy5y, astro, cort.</p>",
      "rawMarkdown": "Ah, Thanks. I mess up the index. It should be shsy5y, astro, cort.",
      "votes": null
    },
    {
      "id": "1623407",
      "postDate": "12/19/2021 20:37:10",
      "content": "<p>Interesting experiment. The numbers you listed divided by your validation class-wise APs would be the class portion in the test dataset. It seems that 80% of our test images are of cell type cort.</p>",
      "rawMarkdown": "Interesting experiment. The numbers you listed divided by your validation class-wise APs would be the class portion in the test dataset. It seems that 80% of our test images are of cell type cort.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1621218,
      "author_name": "nochanged",
      "author_url": "",
      "post_date": "12/17/2021 14:39:38",
      "content": "<p>How do you eval dataset? mAP code did you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1621272,
          "author_name": "thangvubka",
          "author_url": "",
          "post_date": "12/17/2021 15:34:51",
          "content": "<p>I just submit the notebook to LB with predictions belonging to single class only, while predictions belonging to others classes are ignored. Repeat that for 3 classes. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1621314,
      "author_name": "slawekbiel",
      "author_url": "",
      "post_date": "12/17/2021 16:14:52",
      "content": "<p>How are you deciding which test file is which type? Maybe your classifier returns cort for most of the test files :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1621712,
          "author_name": "thangvubka",
          "author_url": "",
          "post_date": "12/18/2021 01:36:11",
          "content": "<p>I dont decide at instance level. Keep only instance belonging to a target class.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1621629,
      "author_name": "omallo",
      "author_url": "",
      "post_date": "12/17/2021 22:01:51",
      "content": "<p>Even if the numbers look odd at first, they are correct.</p>\n<p>When you calculate the score locally per category, you only average across the images of that category. However, when you submit the masks for a single category in Kaggle, the score is averaged across all images of all categories which gives you a much lower score. As an example, let's say the astro category makes up 25% of all images. If you only submit masks for astro and your model gives a perfect score of 1.0, you would only get a score of 0.25 when submitting to Kaggle since there are additional 75% of images with score 0. In order to get the real category score, you would have to multiply the score from Kaggle by 4 (due to astro making up 25% of the images).</p>\n<p>Also, if you sum up your individual submitted category scores, you should get your overall LB score. It doesn't fully match your current LB score but I think it's close enough ;-) Deviations probably come from some misclassification on individual images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1621713,
          "author_name": "thangvubka",
          "author_url": "",
          "post_date": "12/18/2021 01:38:47",
          "content": "<p>I got it. Thank you very much. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1623033,
      "author_name": "ichimarugin",
      "author_url": "",
      "post_date": "12/19/2021 12:10:58",
      "content": "<p>My mAP for each class is weird, too.</p>\n<p>Detectron2 R50</p>\n<p>Astro got high and cort is lower.</p>\n<p>[0mcopypaste: MaP IoU=0.28160431798918684<br>\n[0mcopypaste: mAP cort=0.21041490919401545<br>\n[0mcopypaste: mAP shsy5y=0.16366705343985546<br>\n[0mcopypaste: mAP astro=0.3572197447479763</p>",
      "votes": null,
      "replies": [
        {
          "id": 1623038,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "12/19/2021 12:19:52",
          "content": "<p>Did you double check whether you might be mixing up the category indexes and their names? The raw numbers would look reasonable to me if they were for the categories shsy5y, astro, and cort (in that order).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1623277,
          "author_name": "ichimarugin",
          "author_url": "",
          "post_date": "12/19/2021 17:17:48",
          "content": "<p>Ah, Thanks. I mess up the index. It should be shsy5y, astro, cort.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1623407,
      "author_name": "junxhuang",
      "author_url": "",
      "post_date": "12/19/2021 20:37:10",
      "content": "<p>Interesting experiment. The numbers you listed divided by your validation class-wise APs would be the class portion in the test dataset. It seems that 80% of our test images are of cell type cort.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1621123": "I submitted my results with each prediction belonging to each class only, and I found there is a huge difference among class-wise APs on the LB.\n\ncort: 0.25\nshsy5y: 0.038\nastro : 0.034\n\nI guess that portion of cort images is large, but the scores are still weird to me. Have anyone tried to check class-wise on LB and observed the same score portion to me.",
    "1621218": "How do you eval dataset? mAP code did you use?",
    "1621272": "I just submit the notebook to LB with predictions belonging to single class only, while predictions belonging to others classes are ignored. Repeat that for 3 classes.",
    "1621314": "How are you deciding which test file is which type? Maybe your classifier returns cort for most of the test files :)",
    "1621629": "Even if the numbers look odd at first, they are correct.\n\nWhen you calculate the score locally per category, you only average across the images of that category. However, when you submit the masks for a single category in Kaggle, the score is averaged across all images of all categories which gives you a much lower score. As an example, let's say the astro category makes up 25% of all images. If you only submit masks for astro and your model gives a perfect score of 1.0, you would only get a score of 0.25 when submitting to Kaggle since there are additional 75% of images with score 0. In order to get the real category score, you would have to multiply the score from Kaggle by 4 (due to astro making up 25% of the images).\n\nAlso, if you sum up your individual submitted category scores, you should get your overall LB score. It doesn't fully match your current LB score but I think it's close enough ;-) Deviations probably come from some misclassification on individual images.",
    "1621712": "I dont decide at instance level. Keep only instance belonging to a target class.",
    "1621713": "I got it. Thank you very much.",
    "1623033": "My mAP for each class is weird, too.\n\nDetectron2 R50\n\nAstro got high and cort is lower.\n\n[0mcopypaste: MaP IoU=0.28160431798918684\n[0mcopypaste: mAP cort=0.21041490919401545\n[0mcopypaste: mAP shsy5y=0.16366705343985546\n[0mcopypaste: mAP astro=0.3572197447479763",
    "1623038": "Did you double check whether you might be mixing up the category indexes and their names? The raw numbers would look reasonable to me if they were for the categories shsy5y, astro, and cort (in that order).",
    "1623277": "Ah, Thanks. I mess up the index. It should be shsy5y, astro, cort.",
    "1623407": "Interesting experiment. The numbers you listed divided by your validation class-wise APs would be the class portion in the test dataset. It seems that 80% of our test images are of cell type cort."
  },
  "source": "meta"
}