{
  "id": 244066,
  "title": "LB Probing - A Hypothesis",
  "url": "/competitions/siim-covid19-detection/discussion/244066",
  "author_name": "",
  "post_date": "2021-06-05T06:51:11.141098500Z",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p><strong>Edit:</strong> See the comments for the discussion first. Probing expectations of 1/6 here is wrong. Instead use this <a href=\"https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929405\" target=\"_blank\">notebook</a> for verification with LB probing below.</p>\n<p>I did some LB probing with all my 5 submissions from today here is a summary: </p>\n<p>Here is the <a href=\"https://www.kaggle.com/keremt/understanding-evaluation/output?scriptVersionId=64877409\" target=\"_blank\">notebook</a></p>\n<h3>LB Probes</h3>\n<h4>Probe 1: study (1,1,2,2 invalid boxes) and image (1,1,2,2 invalid boxes):</h4>\n<p>LB:0.0 Expectation:0.0</p>\n<h4>Probe 2: study (1,1,2,2 invalid boxes) and image (none,1,0,0,1,1):</h4>\n<p>LB:0.0 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>LB: 0.0 also supports the existing problem with image level score calculations. Many say it only contributes 0.001 but infact its not contributing at all, not sure when we will have an explanation about this.</p>\n<h4>Probe 3: study (negative,1,1,0,0) and image (1,1,2,2 boxes):</h4>\n<p>LB:0.050 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>This means mAP is weighted by the number of samples of that class in test set. Probe 4 and 5 supports this idea.</p>\n<h4>Probe 4: study (typical,1,1,0,0) and image (1,1,2,2 boxes):</h4>\n<p>LB:0.079 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>This means mAP is weighted.</p>\n<p>LB negative probe / LB typical probe = 0.050 / 0.079 = 0.63</p>\n<p>no. of negative in train / no. of typical in train =  1676 / 2855 = 0.58</p>\n<p>We can see that ratios of LB scores and number of samples are very close. They are not exactly same since class distributions of training and test are not same, we probably have more negatives in test set.</p>\n<h4>Probe 5: study (indeterminate,1,1,0,0) and image (1,1,2,2 boxes):</h4>\n<p>LB:0.031 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>This means mAP is weighted.</p>\n<p>LB negative probe / LB indeterminate probe = 0.050 / 0.031 = 1.61</p>\n<p>no. of negative in train / no. of indeterminate in train =  1676 / 2855 = 1.59</p>\n<h4>Probe 6: study (atypical,1,1,0,0) and image (1,1,2,2 boxes) (TODO):</h4>\n<p>LB:0.031 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>This means mAP is weighted.</p>\n<p>LB negative probe / LB atypical probe = 0.050 / X = Y (Tomorrow: out of submissions)</p>\n<p>no. of negative in train / no. of atypical in train =  1676 / 2855 = 3.53</p>\n<p>X should be close to 0.050 / 3.53 ~= 0.014</p>\n<h3>What are weights for image and study level predictions?</h3>\n<p>Above we see that mAP is not weigthed equally, hence not 1/6. We can see the relative contributions within study level classes, but image level classes might be differently weighted compared to study level classes.</p>\n<p>We kind of now know that study level predictions are weighted proportionally to their number of samples, if we also assume this for all the 6 classes then here is a rough calculation:</p>\n<p>LB (probe 3): 0.050 -  Expectation: no. of negative in train / all train samples = 0.136</p>\n<p>LB (probe 4): 0.079 -  Expectation: no. of negative in train / all train samples = 0.230</p>\n<p>LB (probe 4): 0.031 -  Expectation: no. of negative in train / all train samples = 0.084</p>\n<p>When we assume equal weight between study and image level predictions numbers don't match. Here if we assume a weight of 0.35 for study level predictions:</p>\n<p>LB (probe 3): 0.050 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.047</p>\n<p>LB (probe 4): 0.079 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.080</p>\n<p>LB (probe 4): 0.031 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.029</p>\n<p>Again numbers are slightly different because train and test distribution differences.</p>\n<h3>Evaluation Formula Hypothesis</h3>\n<p>Score: 0.35x(Weighted mAP Study) + 0.65x(Weighted mAP Image) ???</p>\n<p>(We can't do much probing around images since evaluation is currently broken.)</p>\n<p>Everything would be much easier if we could get more answers about metric calculation. It is well known that key to succeeding in any Kaggle competition is to construct a good validation set, and it would only be possible by having the questions around the metric answered.</p>\n<p>Let me know what you think about my hypothesis in the comments, maybe I am missing something. Thanks.</p>",
  "messages": [
    {
      "id": "1336698",
      "postDate": "06/05/2021 06:51:11",
      "content": "<p><strong>Edit:</strong> See the comments for the discussion first. Probing expectations of 1/6 here is wrong. Instead use this <a href=\"https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929405\" target=\"_blank\">notebook</a> for verification with LB probing below.</p>\n<p>I did some LB probing with all my 5 submissions from today here is a summary: </p>\n<p>Here is the <a href=\"https://www.kaggle.com/keremt/understanding-evaluation/output?scriptVersionId=64877409\" target=\"_blank\">notebook</a></p>\n<h3>LB Probes</h3>\n<h4>Probe 1: study (1,1,2,2 invalid boxes) and image (1,1,2,2 invalid boxes):</h4>\n<p>LB:0.0 Expectation:0.0</p>\n<h4>Probe 2: study (1,1,2,2 invalid boxes) and image (none,1,0,0,1,1):</h4>\n<p>LB:0.0 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>LB: 0.0 also supports the existing problem with image level score calculations. Many say it only contributes 0.001 but infact its not contributing at all, not sure when we will have an explanation about this.</p>\n<h4>Probe 3: study (negative,1,1,0,0) and image (1,1,2,2 boxes):</h4>\n<p>LB:0.050 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>This means mAP is weighted by the number of samples of that class in test set. Probe 4 and 5 supports this idea.</p>\n<h4>Probe 4: study (typical,1,1,0,0) and image (1,1,2,2 boxes):</h4>\n<p>LB:0.079 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>This means mAP is weighted.</p>\n<p>LB negative probe / LB typical probe = 0.050 / 0.079 = 0.63</p>\n<p>no. of negative in train / no. of typical in train =  1676 / 2855 = 0.58</p>\n<p>We can see that ratios of LB scores and number of samples are very close. They are not exactly same since class distributions of training and test are not same, we probably have more negatives in test set.</p>\n<h4>Probe 5: study (indeterminate,1,1,0,0) and image (1,1,2,2 boxes):</h4>\n<p>LB:0.031 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>This means mAP is weighted.</p>\n<p>LB negative probe / LB indeterminate probe = 0.050 / 0.031 = 1.61</p>\n<p>no. of negative in train / no. of indeterminate in train =  1676 / 2855 = 1.59</p>\n<h4>Probe 6: study (atypical,1,1,0,0) and image (1,1,2,2 boxes) (TODO):</h4>\n<p>LB:0.031 Expectation:1/6 (Assuming equal class weights mAP)</p>\n<p>This means mAP is weighted.</p>\n<p>LB negative probe / LB atypical probe = 0.050 / X = Y (Tomorrow: out of submissions)</p>\n<p>no. of negative in train / no. of atypical in train =  1676 / 2855 = 3.53</p>\n<p>X should be close to 0.050 / 3.53 ~= 0.014</p>\n<h3>What are weights for image and study level predictions?</h3>\n<p>Above we see that mAP is not weigthed equally, hence not 1/6. We can see the relative contributions within study level classes, but image level classes might be differently weighted compared to study level classes.</p>\n<p>We kind of now know that study level predictions are weighted proportionally to their number of samples, if we also assume this for all the 6 classes then here is a rough calculation:</p>\n<p>LB (probe 3): 0.050 -  Expectation: no. of negative in train / all train samples = 0.136</p>\n<p>LB (probe 4): 0.079 -  Expectation: no. of negative in train / all train samples = 0.230</p>\n<p>LB (probe 4): 0.031 -  Expectation: no. of negative in train / all train samples = 0.084</p>\n<p>When we assume equal weight between study and image level predictions numbers don't match. Here if we assume a weight of 0.35 for study level predictions:</p>\n<p>LB (probe 3): 0.050 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.047</p>\n<p>LB (probe 4): 0.079 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.080</p>\n<p>LB (probe 4): 0.031 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.029</p>\n<p>Again numbers are slightly different because train and test distribution differences.</p>\n<h3>Evaluation Formula Hypothesis</h3>\n<p>Score: 0.35x(Weighted mAP Study) + 0.65x(Weighted mAP Image) ???</p>\n<p>(We can't do much probing around images since evaluation is currently broken.)</p>\n<p>Everything would be much easier if we could get more answers about metric calculation. It is well known that key to succeeding in any Kaggle competition is to construct a good validation set, and it would only be possible by having the questions around the metric answered.</p>\n<p>Let me know what you think about my hypothesis in the comments, maybe I am missing something. Thanks.</p>",
      "rawMarkdown": "**Edit:** See the comments for the discussion first. Probing expectations of 1/6 here is wrong. Instead use this [notebook](https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929405) for verification with LB probing below.\n\nI did some LB probing with all my 5 submissions from today here is a summary: \n\nHere is the [notebook](https://www.kaggle.com/keremt/understanding-evaluation/output?scriptVersionId=64877409)\n\n### LB Probes\n\n#### Probe 1: study (1,1,2,2 invalid boxes) and image (1,1,2,2 invalid boxes):  \n\nLB:0.0 Expectation:0.0\n\n#### Probe 2: study (1,1,2,2 invalid boxes) and image (none,1,0,0,1,1): \n\nLB:0.0 Expectation:1/6 (Assuming equal class weights mAP)\n\nLB: 0.0 also supports the existing problem with image level score calculations. Many say it only contributes 0.001 but infact its not contributing at all, not sure when we will have an explanation about this.\n\n#### Probe 3: study (negative,1,1,0,0) and image (1,1,2,2 boxes):  \n\nLB:0.050 Expectation:1/6 (Assuming equal class weights mAP)\n\nThis means mAP is weighted by the number of samples of that class in test set. Probe 4 and 5 supports this idea.\n\n#### Probe 4: study (typical,1,1,0,0) and image (1,1,2,2 boxes):  \n\nLB:0.079 Expectation:1/6 (Assuming equal class weights mAP)\n\nThis means mAP is weighted.\n\nLB negative probe / LB typical probe = 0.050 / 0.079 = 0.63\n\nno. of negative in train / no. of typical in train =  1676 / 2855 = 0.58\n\nWe can see that ratios of LB scores and number of samples are very close. They are not exactly same since class distributions of training and test are not same, we probably have more negatives in test set.\n\n#### Probe 5: study (indeterminate,1,1,0,0) and image (1,1,2,2 boxes):  \n\nLB:0.031 Expectation:1/6 (Assuming equal class weights mAP)\n\nThis means mAP is weighted.\n\nLB negative probe / LB indeterminate probe = 0.050 / 0.031 = 1.61\n\nno. of negative in train / no. of indeterminate in train =  1676 / 2855 = 1.59\n\n#### Probe 6: study (atypical,1,1,0,0) and image (1,1,2,2 boxes) (TODO):  \n\nLB:0.031 Expectation:1/6 (Assuming equal class weights mAP)\n\nThis means mAP is weighted.\n\nLB negative probe / LB atypical probe = 0.050 / X = Y (Tomorrow: out of submissions)\n\nno. of negative in train / no. of atypical in train =  1676 / 2855 = 3.53\n\nX should be close to 0.050 / 3.53 ~= 0.014\n\n\n### What are weights for image and study level predictions?\n\nAbove we see that mAP is not weigthed equally, hence not 1/6. We can see the relative contributions within study level classes, but image level classes might be differently weighted compared to study level classes.\n\n\nWe kind of now know that study level predictions are weighted proportionally to their number of samples, if we also assume this for all the 6 classes then here is a rough calculation:\n\n\nLB (probe 3): 0.050 -  Expectation: no. of negative in train / all train samples = 0.136\n\nLB (probe 4): 0.079 -  Expectation: no. of negative in train / all train samples = 0.230\n\nLB (probe 4): 0.031 -  Expectation: no. of negative in train / all train samples = 0.084\n\n\nWhen we assume equal weight between study and image level predictions numbers don't match. Here if we assume a weight of 0.35 for study level predictions:\n\nLB (probe 3): 0.050 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.047\n\nLB (probe 4): 0.079 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.080\n\nLB (probe 4): 0.031 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.029\n\nAgain numbers are slightly different because train and test distribution differences.\n\n\n### Evaluation Formula Hypothesis\n\nScore: 0.35x(Weighted mAP Study) + 0.65x(Weighted mAP Image) ???\n\n\n(We can't do much probing around images since evaluation is currently broken.)\n\nEverything would be much easier if we could get more answers about metric calculation. It is well known that key to succeeding in any Kaggle competition is to construct a good validation set, and it would only be possible by having the questions around the metric answered.\n\nLet me know what you think about my hypothesis in the comments, maybe I am missing something. Thanks.",
      "votes": null
    },
    {
      "id": "1337242",
      "postDate": "06/05/2021 13:46:25",
      "content": "<pre><code>Many say it only contributes 0.001 but infact its not contributing at all\n</code></pre>\n<p>The score of the none class is 0, but the score of the opacity class is 0.001.<br>\nAnd there is no difference in weight between the classes, only in the percentage of positive samples.</p>",
      "rawMarkdown": "```\nMany say it only contributes 0.001 but infact its not contributing at all\n```\nThe score of the none class is 0, but the score of the opacity class is 0.001.\nAnd there is no difference in weight between the classes, only in the percentage of positive samples.",
      "votes": null
    },
    {
      "id": "1337499",
      "postDate": "06/05/2021 16:38:58",
      "content": "<p>So, does it mean number of samples is impacting the <strong>mAP</strong> ?</p>",
      "rawMarkdown": "So, does it mean number of samples is impacting the **mAP** ?",
      "votes": null
    },
    {
      "id": "1337509",
      "postDate": "06/05/2021 16:54:18",
      "content": "<p>Thanks for the reply. After your comment I realize that it might not be possible to probe and expect AP=1.00 for (negative,typical,indeterminate or atypical) classes when we predict all studies as same class since precision and recall will depend on the number of positive cases of that class as you explained. Also, ordering will matter during AP calculation and since we predict all confidences as 1 we can't know how predictions will be ordered, and how TPs and FPs will be ordered. Different orders will have different AP and ranking TP higher will result in higher score. How to optimize mAP in a classification setting is little complicated for me at the moment, thanks for your comment which made me think about it :)</p>",
      "rawMarkdown": "Thanks for the reply. After your comment I realize that it might not be possible to probe and expect AP=1.00 for (negative,typical,indeterminate or atypical) classes when we predict all studies as same class since precision and recall will depend on the number of positive cases of that class as you explained. Also, ordering will matter during AP calculation and since we predict all confidences as 1 we can't know how predictions will be ordered, and how TPs and FPs will be ordered. Different orders will have different AP and ranking TP higher will result in higher score. How to optimize mAP in a classification setting is little complicated for me at the moment, thanks for your comment which made me think about it :)",
      "votes": null
    },
    {
      "id": "1337530",
      "postDate": "06/05/2021 17:12:54",
      "content": "<p>I think in case of this probing, although recall is 1 precision is the ratio of positive samples / number of predictions. To my understanding, thats why LB is proportional to number of positive cases. F1 would be more intuitive but with mAP in classification setting I am not very sure whats happening exactly.</p>",
      "rawMarkdown": "I think in case of this probing, although recall is 1 precision is the ratio of positive samples / number of predictions. To my understanding, thats why LB is proportional to number of positive cases. F1 would be more intuitive but with mAP in classification setting I am not very sure whats happening exactly.",
      "votes": null
    },
    {
      "id": "1337703",
      "postDate": "06/05/2021 19:53:14",
      "content": "<p>I am able to replicate my LB probing results using this <a href=\"https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929087\" target=\"_blank\">notebook</a>. Results using training dataset vs LB are very close ~-+0.001 indicating that distribution of training and public test set are very close.</p>",
      "rawMarkdown": "I am able to replicate my LB probing results using this [notebook](https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929087). Results using training dataset vs LB are very close ~-+0.001 indicating that distribution of training and public test set are very close.",
      "votes": null
    },
    {
      "id": "1337774",
      "postDate": "06/05/2021 22:17:45",
      "content": "<p>It certainly seems to be close to the train distribution. Thanks for the info!</p>",
      "rawMarkdown": "It certainly seems to be close to the train distribution. Thanks for the info!",
      "votes": null
    },
    {
      "id": "1338563",
      "postDate": "06/06/2021 14:48:03",
      "content": "<p>Thanks for sharing with the community!</p>",
      "rawMarkdown": "Thanks for sharing with the community!",
      "votes": null
    },
    {
      "id": "1343203",
      "postDate": "06/10/2021 05:33:03",
      "content": "<p>Great job!👍</p>",
      "rawMarkdown": "Great job!👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1337242,
      "author_name": "yujiariyasu",
      "author_url": "",
      "post_date": "06/05/2021 13:46:25",
      "content": "<pre><code>Many say it only contributes 0.001 but infact its not contributing at all\n</code></pre>\n<p>The score of the none class is 0, but the score of the opacity class is 0.001.<br>\nAnd there is no difference in weight between the classes, only in the percentage of positive samples.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337499,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "06/05/2021 16:38:58",
          "content": "<p>So, does it mean number of samples is impacting the <strong>mAP</strong> ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337509,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "06/05/2021 16:54:18",
          "content": "<p>Thanks for the reply. After your comment I realize that it might not be possible to probe and expect AP=1.00 for (negative,typical,indeterminate or atypical) classes when we predict all studies as same class since precision and recall will depend on the number of positive cases of that class as you explained. Also, ordering will matter during AP calculation and since we predict all confidences as 1 we can't know how predictions will be ordered, and how TPs and FPs will be ordered. Different orders will have different AP and ranking TP higher will result in higher score. How to optimize mAP in a classification setting is little complicated for me at the moment, thanks for your comment which made me think about it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337530,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "06/05/2021 17:12:54",
          "content": "<p>I think in case of this probing, although recall is 1 precision is the ratio of positive samples / number of predictions. To my understanding, thats why LB is proportional to number of positive cases. F1 would be more intuitive but with mAP in classification setting I am not very sure whats happening exactly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337703,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "06/05/2021 19:53:14",
          "content": "<p>I am able to replicate my LB probing results using this <a href=\"https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929087\" target=\"_blank\">notebook</a>. Results using training dataset vs LB are very close ~-+0.001 indicating that distribution of training and public test set are very close.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337774,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/05/2021 22:17:45",
          "content": "<p>It certainly seems to be close to the train distribution. Thanks for the info!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1338563,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "06/06/2021 14:48:03",
      "content": "<p>Thanks for sharing with the community!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1343203,
      "author_name": "shangweichen",
      "author_url": "",
      "post_date": "06/10/2021 05:33:03",
      "content": "<p>Great job!👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1336698": "**Edit:** See the comments for the discussion first. Probing expectations of 1/6 here is wrong. Instead use this [notebook](https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929405) for verification with LB probing below.\n\nI did some LB probing with all my 5 submissions from today here is a summary: \n\nHere is the [notebook](https://www.kaggle.com/keremt/understanding-evaluation/output?scriptVersionId=64877409)\n\n### LB Probes\n\n#### Probe 1: study (1,1,2,2 invalid boxes) and image (1,1,2,2 invalid boxes):  \n\nLB:0.0 Expectation:0.0\n\n#### Probe 2: study (1,1,2,2 invalid boxes) and image (none,1,0,0,1,1): \n\nLB:0.0 Expectation:1/6 (Assuming equal class weights mAP)\n\nLB: 0.0 also supports the existing problem with image level score calculations. Many say it only contributes 0.001 but infact its not contributing at all, not sure when we will have an explanation about this.\n\n#### Probe 3: study (negative,1,1,0,0) and image (1,1,2,2 boxes):  \n\nLB:0.050 Expectation:1/6 (Assuming equal class weights mAP)\n\nThis means mAP is weighted by the number of samples of that class in test set. Probe 4 and 5 supports this idea.\n\n#### Probe 4: study (typical,1,1,0,0) and image (1,1,2,2 boxes):  \n\nLB:0.079 Expectation:1/6 (Assuming equal class weights mAP)\n\nThis means mAP is weighted.\n\nLB negative probe / LB typical probe = 0.050 / 0.079 = 0.63\n\nno. of negative in train / no. of typical in train =  1676 / 2855 = 0.58\n\nWe can see that ratios of LB scores and number of samples are very close. They are not exactly same since class distributions of training and test are not same, we probably have more negatives in test set.\n\n#### Probe 5: study (indeterminate,1,1,0,0) and image (1,1,2,2 boxes):  \n\nLB:0.031 Expectation:1/6 (Assuming equal class weights mAP)\n\nThis means mAP is weighted.\n\nLB negative probe / LB indeterminate probe = 0.050 / 0.031 = 1.61\n\nno. of negative in train / no. of indeterminate in train =  1676 / 2855 = 1.59\n\n#### Probe 6: study (atypical,1,1,0,0) and image (1,1,2,2 boxes) (TODO):  \n\nLB:0.031 Expectation:1/6 (Assuming equal class weights mAP)\n\nThis means mAP is weighted.\n\nLB negative probe / LB atypical probe = 0.050 / X = Y (Tomorrow: out of submissions)\n\nno. of negative in train / no. of atypical in train =  1676 / 2855 = 3.53\n\nX should be close to 0.050 / 3.53 ~= 0.014\n\n\n### What are weights for image and study level predictions?\n\nAbove we see that mAP is not weigthed equally, hence not 1/6. We can see the relative contributions within study level classes, but image level classes might be differently weighted compared to study level classes.\n\n\nWe kind of now know that study level predictions are weighted proportionally to their number of samples, if we also assume this for all the 6 classes then here is a rough calculation:\n\n\nLB (probe 3): 0.050 -  Expectation: no. of negative in train / all train samples = 0.136\n\nLB (probe 4): 0.079 -  Expectation: no. of negative in train / all train samples = 0.230\n\nLB (probe 4): 0.031 -  Expectation: no. of negative in train / all train samples = 0.084\n\n\nWhen we assume equal weight between study and image level predictions numbers don't match. Here if we assume a weight of 0.35 for study level predictions:\n\nLB (probe 3): 0.050 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.047\n\nLB (probe 4): 0.079 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.080\n\nLB (probe 4): 0.031 -  Expectation: 0.35*(no. of negative in train / all train samples) = 0.029\n\nAgain numbers are slightly different because train and test distribution differences.\n\n\n### Evaluation Formula Hypothesis\n\nScore: 0.35x(Weighted mAP Study) + 0.65x(Weighted mAP Image) ???\n\n\n(We can't do much probing around images since evaluation is currently broken.)\n\nEverything would be much easier if we could get more answers about metric calculation. It is well known that key to succeeding in any Kaggle competition is to construct a good validation set, and it would only be possible by having the questions around the metric answered.\n\nLet me know what you think about my hypothesis in the comments, maybe I am missing something. Thanks.",
    "1337242": "```\nMany say it only contributes 0.001 but infact its not contributing at all\n```\nThe score of the none class is 0, but the score of the opacity class is 0.001.\nAnd there is no difference in weight between the classes, only in the percentage of positive samples.",
    "1337499": "So, does it mean number of samples is impacting the **mAP** ?",
    "1337509": "Thanks for the reply. After your comment I realize that it might not be possible to probe and expect AP=1.00 for (negative,typical,indeterminate or atypical) classes when we predict all studies as same class since precision and recall will depend on the number of positive cases of that class as you explained. Also, ordering will matter during AP calculation and since we predict all confidences as 1 we can't know how predictions will be ordered, and how TPs and FPs will be ordered. Different orders will have different AP and ranking TP higher will result in higher score. How to optimize mAP in a classification setting is little complicated for me at the moment, thanks for your comment which made me think about it :)",
    "1337530": "I think in case of this probing, although recall is 1 precision is the ratio of positive samples / number of predictions. To my understanding, thats why LB is proportional to number of positive cases. F1 would be more intuitive but with mAP in classification setting I am not very sure whats happening exactly.",
    "1337703": "I am able to replicate my LB probing results using this [notebook](https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929087). Results using training dataset vs LB are very close ~-+0.001 indicating that distribution of training and public test set are very close.",
    "1337774": "It certainly seems to be close to the train distribution. Thanks for the info!",
    "1338563": "Thanks for sharing with the community!",
    "1343203": "Great job!👍"
  },
  "source": "meta"
}