{
  "id": 310103,
  "title": "How to calculate validation score on Nearest Neighbors",
  "url": "/competitions/happy-whale-and-dolphin/discussion/310103",
  "author_name": "",
  "post_date": "2022-02-27T15:14:04.657430400Z",
  "votes": 43,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have received some questions regarding the function that I implemented in my Pytorch Inference kernel for calculating the validation score. (<a href=\"https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling\" target=\"_blank\">https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling</a>)<br>\nI will try to describe the flow of the logic behind it. Indeed is a little hard to follow, probably my code isn't very descriptive also.<br>\nSo, this are the steps:<br>\nVALIDATION PROCESS</p>\n<ol>\n<li>Fit the embeddings of the training data on the Nearest Neighbors algorithm</li>\n<li>Use the fitted Nearest Neighbors algorithm to get the closest points and distances to them for the valid data</li>\n<li>Create a confidence metric that is 1-distance (being a cosine distance, the max is always 1)</li>\n<li>We use the desired thresholds to check if the prediction confidences are below or above it.<br>\nIf the predicted categories are [a,b,c,d,e] and the confidences are [0.9, 0.8, 0.7, 0.6, 0.5] we start comparing the first element with the threshold and then move along the vector (the vector is always sorted descended, so it starts with the most confident predictions, than follows the second most confident prediction and so on).<br>\nSo, let's say the threshold is 0.75. The first element is above the threshold and keeps his value, the second one the same, but the third one being below the threshold gets replaces with \"new_individual\" and everything makes one more to the right. Result will be [a,b,\"new_individual\",c,d]. </li>\n<li>Now, after getting of the prediction through this methodology we need a way to evaluate the score. Where there is not an \"new_individual\" prediction we simply check the metric for valid predictions vs valid ground truth, simple as that. But how do we score elements with \"new_individual\" ? Well, first we check if the ground truth for that valid sample exist in the training labels. If it does, we leave it as that meaning that that prediction is not right(it is not a new_individual), if the ground truth for that sample does not exist in the training data, that means that our new_individual prediction was right and to calculate the score we are replacing the prediction with the ground truth</li>\n<li>Do the first 5 steps for a range of thresholds and choose the best one for the final prediction. Score function is using map_per_image for getting the metric (originally implemented by <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>, more explanations about it you can find here: <a href=\"https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric\" target=\"_blank\">https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric</a>) </li>\n</ol>",
  "messages": [
    {
      "id": "1706553",
      "postDate": "02/27/2022 15:14:04",
      "content": "<p>I have received some questions regarding the function that I implemented in my Pytorch Inference kernel for calculating the validation score. (<a href=\"https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling\" target=\"_blank\">https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling</a>)<br>\nI will try to describe the flow of the logic behind it. Indeed is a little hard to follow, probably my code isn't very descriptive also.<br>\nSo, this are the steps:<br>\nVALIDATION PROCESS</p>\n<ol>\n<li>Fit the embeddings of the training data on the Nearest Neighbors algorithm</li>\n<li>Use the fitted Nearest Neighbors algorithm to get the closest points and distances to them for the valid data</li>\n<li>Create a confidence metric that is 1-distance (being a cosine distance, the max is always 1)</li>\n<li>We use the desired thresholds to check if the prediction confidences are below or above it.<br>\nIf the predicted categories are [a,b,c,d,e] and the confidences are [0.9, 0.8, 0.7, 0.6, 0.5] we start comparing the first element with the threshold and then move along the vector (the vector is always sorted descended, so it starts with the most confident predictions, than follows the second most confident prediction and so on).<br>\nSo, let's say the threshold is 0.75. The first element is above the threshold and keeps his value, the second one the same, but the third one being below the threshold gets replaces with \"new_individual\" and everything makes one more to the right. Result will be [a,b,\"new_individual\",c,d]. </li>\n<li>Now, after getting of the prediction through this methodology we need a way to evaluate the score. Where there is not an \"new_individual\" prediction we simply check the metric for valid predictions vs valid ground truth, simple as that. But how do we score elements with \"new_individual\" ? Well, first we check if the ground truth for that valid sample exist in the training labels. If it does, we leave it as that meaning that that prediction is not right(it is not a new_individual), if the ground truth for that sample does not exist in the training data, that means that our new_individual prediction was right and to calculate the score we are replacing the prediction with the ground truth</li>\n<li>Do the first 5 steps for a range of thresholds and choose the best one for the final prediction. Score function is using map_per_image for getting the metric (originally implemented by <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>, more explanations about it you can find here: <a href=\"https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric\" target=\"_blank\">https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric</a>) </li>\n</ol>",
      "rawMarkdown": "I have received some questions regarding the function that I implemented in my Pytorch Inference kernel for calculating the validation score. (https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling)\nI will try to describe the flow of the logic behind it. Indeed is a little hard to follow, probably my code isn't very descriptive also.\nSo, this are the steps:\nVALIDATION PROCESS\n1. Fit the embeddings of the training data on the Nearest Neighbors algorithm\n2. Use the fitted Nearest Neighbors algorithm to get the closest points and distances to them for the valid data\n3. Create a confidence metric that is 1-distance (being a cosine distance, the max is always 1)\n4. We use the desired thresholds to check if the prediction confidences are below or above it.\nIf the predicted categories are [a,b,c,d,e] and the confidences are [0.9, 0.8, 0.7, 0.6, 0.5] we start comparing the first element with the threshold and then move along the vector (the vector is always sorted descended, so it starts with the most confident predictions, than follows the second most confident prediction and so on).\nSo, let's say the threshold is 0.75. The first element is above the threshold and keeps his value, the second one the same, but the third one being below the threshold gets replaces with \"new_individual\" and everything makes one more to the right. Result will be [a,b,\"new_individual\",c,d]. \n5. Now, after getting of the prediction through this methodology we need a way to evaluate the score. Where there is not an \"new_individual\" prediction we simply check the metric for valid predictions vs valid ground truth, simple as that. But how do we score elements with \"new_individual\" ? Well, first we check if the ground truth for that valid sample exist in the training labels. If it does, we leave it as that meaning that that prediction is not right(it is not a new_individual), if the ground truth for that sample does not exist in the training data, that means that our new_individual prediction was right and to calculate the score we are replacing the prediction with the ground truth\n6. Do the first 5 steps for a range of thresholds and choose the best one for the final prediction. Score function is using map_per_image for getting the metric (originally implemented by @pestipeti, more explanations about it you can find here: https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric)",
      "votes": null
    },
    {
      "id": "1706699",
      "postDate": "02/27/2022 17:52:28",
      "content": "<p><a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a> yaaay! everything is clear now! thanks for this and the kernel</p>",
      "rawMarkdown": "vladvdv yaaay! everything is clear now! thanks for this and the kernel",
      "votes": null
    },
    {
      "id": "1706748",
      "postDate": "02/27/2022 18:45:02",
      "content": "<p>Glad that it helped. Good luck ! </p>",
      "rawMarkdown": "Glad that it helped. Good luck !",
      "votes": null
    },
    {
      "id": "1710097",
      "postDate": "03/02/2022 18:41:23",
      "content": "<p>Very interesting! I haven't seen something like this before--I assumed the standard was just picking the top k from the softmax confidences. </p>\n<p>Is this a commonplace method or something you originally developed? Very cool</p>",
      "rawMarkdown": "Very interesting! I haven't seen something like this before--I assumed the standard was just picking the top k from the softmax confidences. \n\nIs this a commonplace method or something you originally developed? Very cool",
      "votes": null
    },
    {
      "id": "1710226",
      "postDate": "03/02/2022 20:25:16",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/jtan2231\" target=\"_blank\">@jtan2231</a> <br>\nIt is an adapted version (for this competition context) of the classical validation process</p>",
      "rawMarkdown": "Thank you @jtan2231 \nIt is an adapted version (for this competition context) of the classical validation process",
      "votes": null
    },
    {
      "id": "1712092",
      "postDate": "03/04/2022 15:52:55",
      "content": "<p>Very interesting! Nice work.</p>",
      "rawMarkdown": "Very interesting! Nice work.",
      "votes": null
    },
    {
      "id": "1712202",
      "postDate": "03/04/2022 17:02:52",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/artemburenok\" target=\"_blank\">@artemburenok</a> </p>",
      "rawMarkdown": "Thank you @artemburenok",
      "votes": null
    },
    {
      "id": "1731627",
      "postDate": "03/22/2022 14:49:34",
      "content": "<p>Oh <a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a> thank you for this topic!</p>\n<p>I can tell you a funny story:<br>\nMy best score was around 0.818-0.820, but I have a bug in my submission-processing code.<br>\nI add to predictions only <code>individual_ids &gt; threshold</code>, and all places below threshold were new_individuals.<br>\nFor example, for <code>ind1=0.9, ind2=0.8, ind3=0.7, ind4=0.6, ind5=0.5</code> and <code>threshold=0.75</code> my prediction was like <code>image_name 'ind1 ind2 new_individual new_individual new_individual'</code>.<br>\nWhich is obviously <strong>very</strong> bad decision.<br>\nAfter I saw you topic (unfortunately quite late) and have this bug fixed, I have reached 0.841.<br>\nSo don't make my mistakes 😃</p>",
      "rawMarkdown": "Oh @vladvdv thank you for this topic!\n\nI can tell you a funny story:\nMy best score was around 0.818-0.820, but I have a bug in my submission-processing code.\nI add to predictions only `individual_ids > threshold`, and all places below threshold were new_individuals.\nFor example, for `ind1=0.9, ind2=0.8, ind3=0.7, ind4=0.6, ind5=0.5` and `threshold=0.75` my prediction was like `image_name 'ind1 ind2 new_individual new_individual new_individual'`.\nWhich is obviously **very** bad decision.\nAfter I saw you topic (unfortunately quite late) and have this bug fixed, I have reached 0.841.\nSo don't make my mistakes 😃",
      "votes": null
    },
    {
      "id": "1731815",
      "postDate": "03/22/2022 17:57:54",
      "content": "<p><a href=\"https://www.kaggle.com/ilyadobrynin\" target=\"_blank\">@ilyadobrynin</a> I know exactly what you mean, first I did it the same way like you + other extra 3-4 mistakes 😁</p>",
      "rawMarkdown": "ilyadobrynin I know exactly what you mean, first I did it the same way like you + other extra 3-4 mistakes 😁",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1706699,
      "author_name": "sparkyjunior",
      "author_url": "",
      "post_date": "02/27/2022 17:52:28",
      "content": "<p><a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a> yaaay! everything is clear now! thanks for this and the kernel</p>",
      "votes": null,
      "replies": [
        {
          "id": 1706748,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/27/2022 18:45:02",
          "content": "<p>Glad that it helped. Good luck ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1710097,
      "author_name": "jtan2231",
      "author_url": "",
      "post_date": "03/02/2022 18:41:23",
      "content": "<p>Very interesting! I haven't seen something like this before--I assumed the standard was just picking the top k from the softmax confidences. </p>\n<p>Is this a commonplace method or something you originally developed? Very cool</p>",
      "votes": null,
      "replies": [
        {
          "id": 1710226,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "03/02/2022 20:25:16",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/jtan2231\" target=\"_blank\">@jtan2231</a> <br>\nIt is an adapted version (for this competition context) of the classical validation process</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1712092,
      "author_name": "artemburenok",
      "author_url": "",
      "post_date": "03/04/2022 15:52:55",
      "content": "<p>Very interesting! Nice work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1712202,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "03/04/2022 17:02:52",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/artemburenok\" target=\"_blank\">@artemburenok</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1731627,
      "author_name": "ilyadobrynin",
      "author_url": "",
      "post_date": "03/22/2022 14:49:34",
      "content": "<p>Oh <a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a> thank you for this topic!</p>\n<p>I can tell you a funny story:<br>\nMy best score was around 0.818-0.820, but I have a bug in my submission-processing code.<br>\nI add to predictions only <code>individual_ids &gt; threshold</code>, and all places below threshold were new_individuals.<br>\nFor example, for <code>ind1=0.9, ind2=0.8, ind3=0.7, ind4=0.6, ind5=0.5</code> and <code>threshold=0.75</code> my prediction was like <code>image_name 'ind1 ind2 new_individual new_individual new_individual'</code>.<br>\nWhich is obviously <strong>very</strong> bad decision.<br>\nAfter I saw you topic (unfortunately quite late) and have this bug fixed, I have reached 0.841.<br>\nSo don't make my mistakes 😃</p>",
      "votes": null,
      "replies": [
        {
          "id": 1731815,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "03/22/2022 17:57:54",
          "content": "<p><a href=\"https://www.kaggle.com/ilyadobrynin\" target=\"_blank\">@ilyadobrynin</a> I know exactly what you mean, first I did it the same way like you + other extra 3-4 mistakes 😁</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1706553": "I have received some questions regarding the function that I implemented in my Pytorch Inference kernel for calculating the validation score. (https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling)\nI will try to describe the flow of the logic behind it. Indeed is a little hard to follow, probably my code isn't very descriptive also.\nSo, this are the steps:\nVALIDATION PROCESS\n1. Fit the embeddings of the training data on the Nearest Neighbors algorithm\n2. Use the fitted Nearest Neighbors algorithm to get the closest points and distances to them for the valid data\n3. Create a confidence metric that is 1-distance (being a cosine distance, the max is always 1)\n4. We use the desired thresholds to check if the prediction confidences are below or above it.\nIf the predicted categories are [a,b,c,d,e] and the confidences are [0.9, 0.8, 0.7, 0.6, 0.5] we start comparing the first element with the threshold and then move along the vector (the vector is always sorted descended, so it starts with the most confident predictions, than follows the second most confident prediction and so on).\nSo, let's say the threshold is 0.75. The first element is above the threshold and keeps his value, the second one the same, but the third one being below the threshold gets replaces with \"new_individual\" and everything makes one more to the right. Result will be [a,b,\"new_individual\",c,d]. \n5. Now, after getting of the prediction through this methodology we need a way to evaluate the score. Where there is not an \"new_individual\" prediction we simply check the metric for valid predictions vs valid ground truth, simple as that. But how do we score elements with \"new_individual\" ? Well, first we check if the ground truth for that valid sample exist in the training labels. If it does, we leave it as that meaning that that prediction is not right(it is not a new_individual), if the ground truth for that sample does not exist in the training data, that means that our new_individual prediction was right and to calculate the score we are replacing the prediction with the ground truth\n6. Do the first 5 steps for a range of thresholds and choose the best one for the final prediction. Score function is using map_per_image for getting the metric (originally implemented by @pestipeti, more explanations about it you can find here: https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric)",
    "1706699": "vladvdv yaaay! everything is clear now! thanks for this and the kernel",
    "1706748": "Glad that it helped. Good luck !",
    "1710097": "Very interesting! I haven't seen something like this before--I assumed the standard was just picking the top k from the softmax confidences. \n\nIs this a commonplace method or something you originally developed? Very cool",
    "1710226": "Thank you @jtan2231 \nIt is an adapted version (for this competition context) of the classical validation process",
    "1712092": "Very interesting! Nice work.",
    "1712202": "Thank you @artemburenok",
    "1731627": "Oh @vladvdv thank you for this topic!\n\nI can tell you a funny story:\nMy best score was around 0.818-0.820, but I have a bug in my submission-processing code.\nI add to predictions only `individual_ids > threshold`, and all places below threshold were new_individuals.\nFor example, for `ind1=0.9, ind2=0.8, ind3=0.7, ind4=0.6, ind5=0.5` and `threshold=0.75` my prediction was like `image_name 'ind1 ind2 new_individual new_individual new_individual'`.\nWhich is obviously **very** bad decision.\nAfter I saw you topic (unfortunately quite late) and have this bug fixed, I have reached 0.841.\nSo don't make my mistakes 😃",
    "1731815": "ilyadobrynin I know exactly what you mean, first I did it the same way like you + other extra 3-4 mistakes 😁"
  },
  "source": "meta"
}