{
  "id": 305428,
  "title": "LB probing and train/test split",
  "url": "/competitions/happy-whale-and-dolphin/discussion/305428",
  "author_name": "",
  "post_date": "2022-02-05T10:43:31.022905500Z",
  "votes": 92,
  "comment_count": 19,
  "views": 0,
  "content": "<p>There are <code>51033</code> images in train dataset and <code>27956</code> in test dataset; <code>78989</code> in total.</p>\n<p>There are <code>15587</code> different individual ids in train data, <code>9258</code> (~18% of photos) with 1 photo only. If data was randomly split, those 1 photo ind. ids should be somewhat similar to new_individual present in test data distribution-wise: if we just take random validation split and mark all ids that are not present in the train as new_individual, the majority of those new_individuals would be 1 photo only samples.</p>\n<p>Because of how target metric works, we can directly estimate how many new_individual images there are in LB split by just submitting file with all new_individual predictions, which gives <code>0.112</code> LB score.</p>\n<p>That means there are around <code>3131</code> (~11%) new_individual photos, which is very different from 18% present in train.</p>\n<p>Another interesting thing to note is that there are individuals that have lots of images in train, but are not present in public LB whatsoever (those are top 2 highest train images ids):</p>\n<table>\n<thead>\n<tr>\n<th>id</th>\n<th>Imgs in train</th>\n<th>Percentage in train</th>\n<th>Percentage in public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>37c7aba965a5</td>\n<td>400</td>\n<td>0.007838</td>\n<td>0.000</td>\n</tr>\n<tr>\n<td>114207cab555</td>\n<td>168</td>\n<td>0.003292</td>\n<td>0.000</td>\n</tr>\n</tbody>\n</table>\n<p>All of this means that train/test split is not truly random - some specific choice criteria were applied to get it. It would be good for competition hosts to clarify this and perhaps share splitting algorithm details in order to make validation approaches more stable and less luck-dependent. </p>",
  "messages": [
    {
      "id": "1676854",
      "postDate": "02/05/2022 10:43:31",
      "content": "<p>There are <code>51033</code> images in train dataset and <code>27956</code> in test dataset; <code>78989</code> in total.</p>\n<p>There are <code>15587</code> different individual ids in train data, <code>9258</code> (~18% of photos) with 1 photo only. If data was randomly split, those 1 photo ind. ids should be somewhat similar to new_individual present in test data distribution-wise: if we just take random validation split and mark all ids that are not present in the train as new_individual, the majority of those new_individuals would be 1 photo only samples.</p>\n<p>Because of how target metric works, we can directly estimate how many new_individual images there are in LB split by just submitting file with all new_individual predictions, which gives <code>0.112</code> LB score.</p>\n<p>That means there are around <code>3131</code> (~11%) new_individual photos, which is very different from 18% present in train.</p>\n<p>Another interesting thing to note is that there are individuals that have lots of images in train, but are not present in public LB whatsoever (those are top 2 highest train images ids):</p>\n<table>\n<thead>\n<tr>\n<th>id</th>\n<th>Imgs in train</th>\n<th>Percentage in train</th>\n<th>Percentage in public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>37c7aba965a5</td>\n<td>400</td>\n<td>0.007838</td>\n<td>0.000</td>\n</tr>\n<tr>\n<td>114207cab555</td>\n<td>168</td>\n<td>0.003292</td>\n<td>0.000</td>\n</tr>\n</tbody>\n</table>\n<p>All of this means that train/test split is not truly random - some specific choice criteria were applied to get it. It would be good for competition hosts to clarify this and perhaps share splitting algorithm details in order to make validation approaches more stable and less luck-dependent. </p>",
      "rawMarkdown": "There are `51033` images in train dataset and `27956` in test dataset; `78989` in total.\n\nThere are `15587` different individual ids in train data, `9258` (~18% of photos) with 1 photo only. If data was randomly split, those 1 photo ind. ids should be somewhat similar to new_individual present in test data distribution-wise: if we just take random validation split and mark all ids that are not present in the train as new_individual, the majority of those new_individuals would be 1 photo only samples.\n\nBecause of how target metric works, we can directly estimate how many new_individual images there are in LB split by just submitting file with all new_individual predictions, which gives `0.112` LB score.\n\nThat means there are around `3131` (~11%) new_individual photos, which is very different from 18% present in train.\n\nAnother interesting thing to note is that there are individuals that have lots of images in train, but are not present in public LB whatsoever (those are top 2 highest train images ids):\n\n| id | Imgs in train | Percentage in train | Percentage in public LB |\n| --- | --- | --- | --- |\n| 37c7aba965a5 | 400 | 0.007838 | 0.000 |\n| 114207cab555 | 168 | 0.003292 | 0.000 |\n\nAll of this means that train/test split is not truly random - some specific choice criteria were applied to get it. It would be good for competition hosts to clarify this and perhaps share splitting algorithm details in order to make validation approaches more stable and less luck-dependent.",
      "votes": null
    },
    {
      "id": "1677097",
      "postDate": "02/05/2022 14:26:45",
      "content": "<p>I guess it is common practice in identification tasks: we have ID intersection between train and test, also, we have some ids only in train and some ids only in test (11% in public LB)</p>",
      "rawMarkdown": "I guess it is common practice in identification tasks: we have ID intersection between train and test, also, we have some ids only in train and some ids only in test (11% in public LB)",
      "votes": null
    },
    {
      "id": "1677263",
      "postDate": "02/05/2022 15:56:29",
      "content": "<p>Yeah, I think it is not a big deal but decided to share so everybody knows</p>",
      "rawMarkdown": "Yeah, I think it is not a big deal but decided to share so everybody knows",
      "votes": null
    },
    {
      "id": "1677750",
      "postDate": "02/06/2022 00:46:51",
      "content": "<p>Thanks for the effort in this topic.<br>\nHow to approach the assignment of \"new_individual\" in some test images?<br>\nI thought to set a threshold on the probabilities and do a grid search on the CV score to see what threshold optimizes the MAP@5 CV.</p>",
      "rawMarkdown": "Thanks for the effort in this topic.\nHow to approach the assignment of \"new_individual\" in some test images?\nI thought to set a threshold on the probabilities and do a grid search on the CV score to see what threshold optimizes the MAP@5 CV.",
      "votes": null
    },
    {
      "id": "1678005",
      "postDate": "02/06/2022 07:56:41",
      "content": "<p>I think that this is a good basic approach. In previous Happywhale competition some people were using sophisticated post-processing techniques and standalone new_whale/non new_whale classifier models just for that, so it would be interesting to check them out aswell </p>",
      "rawMarkdown": "I think that this is a good basic approach. In previous Happywhale competition some people were using sophisticated post-processing techniques and standalone new_whale/non new_whale classifier models just for that, so it would be interesting to check them out aswell",
      "votes": null
    },
    {
      "id": "1678129",
      "postDate": "02/06/2022 09:18:46",
      "content": "<p>In the last competition, there was a class for \"new whales\", we don't have this time, I think the new_whale/non new_whale classifier maybe not be possible here. But yes some kind of post-processing can be used</p>",
      "rawMarkdown": "In the last competition, there was a class for \"new whales\", we don't have this time, I think the new_whale/non new_whale classifier maybe not be possible here. But yes some kind of post-processing can be used",
      "votes": null
    },
    {
      "id": "1678552",
      "postDate": "02/06/2022 16:03:16",
      "content": "<p>there is a new_whale class here like in the previous competition -- this is essential to make the winning algorithms applicable in the real world where we often are discovering new individuals to add to our known set of whales and dolphins</p>",
      "rawMarkdown": "there is a new_whale class here like in the previous competition -- this is essential to make the winning algorithms applicable in the real world where we often are discovering new individuals to add to our known set of whales and dolphins",
      "votes": null
    },
    {
      "id": "1683498",
      "postDate": "02/09/2022 20:32:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tedcheese\" target=\"_blank\">@tedcheese</a>,</p>\n<p>Is the distribution of new individuals similar in the public and private test sets?</p>",
      "rawMarkdown": "Hi @tedcheese,\n\nIs the distribution of new individuals similar in the public and private test sets?",
      "votes": null
    },
    {
      "id": "1683518",
      "postDate": "02/09/2022 20:51:07",
      "content": "<p>I'll refer this question to <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> :-)</p>",
      "rawMarkdown": "I'll refer this question to @inversion :-)",
      "votes": null
    },
    {
      "id": "1684147",
      "postDate": "02/10/2022 09:57:08",
      "content": "<p>I think it would be great if they did a time split. This is the true nature of the problem anyway; we use data from before now to build models that predict data after now. This could also explain the top occurring ids not being in the test set. Those whales simply could have passed away. If it is not a time split, I think it's a little frustrating that they would purposefully do things like put all the images for the top occurring ids in the train set. Why would you do this? You would be punishing models that learned well from lots of data. But if they did use a time series split, why not just say so? </p>\n<p>Either way, it would be most beneficial if the organizers could be more explicit in how they did the data split. Otherwise, there will be lots of brain power and submissions committed to probing the leaderboard instead of solving the problem as the organizers want. In fact, I can't think of a better use of submissions other than going down the list of the most common ids and submitting just so I can see what proportion they are of the lb test data. Because the most important step is creating a proper validation strategy. </p>\n<p>Please, <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>, shed some light on the data split.</p>",
      "rawMarkdown": "I think it would be great if they did a time split. This is the true nature of the problem anyway; we use data from before now to build models that predict data after now. This could also explain the top occurring ids not being in the test set. Those whales simply could have passed away. If it is not a time split, I think it's a little frustrating that they would purposefully do things like put all the images for the top occurring ids in the train set. Why would you do this? You would be punishing models that learned well from lots of data. But if they did use a time series split, why not just say so? \n\nEither way, it would be most beneficial if the organizers could be more explicit in how they did the data split. Otherwise, there will be lots of brain power and submissions committed to probing the leaderboard instead of solving the problem as the organizers want. In fact, I can't think of a better use of submissions other than going down the list of the most common ids and submitting just so I can see what proportion they are of the lb test data. Because the most important step is creating a proper validation strategy. \n\nPlease, @inversion, shed some light on the data split.",
      "votes": null
    },
    {
      "id": "1684157",
      "postDate": "02/10/2022 10:05:50",
      "content": "<p>I agree with your comments about importance of details of the data split. I think that train/test split details should be clearly stated in every competition, considering that the entire modern machine learning relies on identical train/test distribution assumption.</p>\n<p>But I think that time split is a little bit redundant in this task. Without the clear time split, models should be able to figure more time-independent factors for identifying animals, which could prove more useful in practical applications (i.e. using the models for a long time without retraining with newer data, and overall more reliable perhaps?)</p>",
      "rawMarkdown": "I agree with your comments about importance of details of the data split. I think that train/test split details should be clearly stated in every competition, considering that the entire modern machine learning relies on identical train/test distribution assumption.\n\nBut I think that time split is a little bit redundant in this task. Without the clear time split, models should be able to figure more time-independent factors for identifying animals, which could prove more useful in practical applications (i.e. using the models for a long time without retraining with newer data, and overall more reliable perhaps?)",
      "votes": null
    },
    {
      "id": "1689145",
      "postDate": "02/14/2022 04:12:42",
      "content": "<p>Very helpful, thank you!</p>",
      "rawMarkdown": "Very helpful, thank you!",
      "votes": null
    },
    {
      "id": "1690953",
      "postDate": "02/15/2022 07:06:40",
      "content": "<p><code>there is a new_whale class here like in the previous competition</code> but there are none in the train data this time, is there some reason this time<br>\n<a href=\"https://www.kaggle.com/tedcheese\" target=\"_blank\">@tedcheese</a> </p>",
      "rawMarkdown": "`there is a new_whale class here like in the previous competition` but there are none in the train data this time, is there some reason this time\n@tedcheese",
      "votes": null
    },
    {
      "id": "1691815",
      "postDate": "02/15/2022 16:19:46",
      "content": "<blockquote>\n  <p>If data was randomly split, those 1 photo ind. ids should be somewhat similar to new_individual present in test data distribution-wise: if we just take random validation split and mark all ids that are not present in the train as new_individual, the majority of those new_individuals would be 1 photo only samples.</p>\n</blockquote>\n<p>But if these whales were to appear in the public or private test set, they would not actually have the label \"new_individual\", right?</p>",
      "rawMarkdown": "> If data was randomly split, those 1 photo ind. ids should be somewhat similar to new_individual present in test data distribution-wise: if we just take random validation split and mark all ids that are not present in the train as new_individual, the majority of those new_individuals would be 1 photo only samples.\n\nBut if these whales were to appear in the public or private test set, they would not actually have the label \"new_individual\", right?",
      "votes": null
    },
    {
      "id": "1692570",
      "postDate": "02/16/2022 05:48:50",
      "content": "<p>Good point <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a>. That was a bit of a fail in the previous competition because if a train data whale is a new_whale class, we're a priori stating that there's no known match in the test set. But in the real world we'd rarely know a test whale couldn't be any of the known train whale, only would know it if the test whale was a known age whale younger than some of the train whales, or vice versa a deceased whale (ie a necropsy photo) from a date past that couldn't match to more recent living whales.</p>\n<p>So, this competition is set to be more reflective of realty.</p>",
      "rawMarkdown": "Good point @mrinath. That was a bit of a fail in the previous competition because if a train data whale is a new_whale class, we're a priori stating that there's no known match in the test set. But in the real world we'd rarely know a test whale couldn't be any of the known train whale, only would know it if the test whale was a known age whale younger than some of the train whales, or vice versa a deceased whale (ie a necropsy photo) from a date past that couldn't match to more recent living whales.\n\nSo, this competition is set to be more reflective of realty.",
      "votes": null
    },
    {
      "id": "1692587",
      "postDate": "02/16/2022 06:10:01",
      "content": "<p>I see, thats actually great thinking! we could know if our models could really be useful in the real world.</p>",
      "rawMarkdown": "I see, thats actually great thinking! we could know if our models could really be useful in the real world.",
      "votes": null
    },
    {
      "id": "1692895",
      "postDate": "02/16/2022 10:33:30",
      "content": "<p>There are a lot of individuals appearing only once in the train set and those of them who you put to your local val set are the \"new whale\" class for the local training effectively.<br>\nYou just need to properly calculate the CV metric to account for this.</p>",
      "rawMarkdown": "There are a lot of individuals appearing only once in the train set and those of them who you put to your local val set are the \"new whale\" class for the local training effectively.\nYou just need to properly calculate the CV metric to account for this.",
      "votes": null
    },
    {
      "id": "1712736",
      "postDate": "03/05/2022 09:00:54",
      "content": "<p>Thank you for sharing the insightful information!</p>",
      "rawMarkdown": "Thank you for sharing the insightful information!",
      "votes": null
    },
    {
      "id": "1715473",
      "postDate": "03/08/2022 03:45:13",
      "content": "<p>I wonder if someone did a complete LB probing to find out all individuals that do not exist in the public LB?😂😂</p>",
      "rawMarkdown": "I wonder if someone did a complete LB probing to find out all individuals that do not exist in the public LB?😂😂",
      "votes": null
    },
    {
      "id": "1715930",
      "postDate": "03/08/2022 14:05:20",
      "content": "<p>Well sadly that would be impossible with limited precision of lb :((</p>",
      "rawMarkdown": "Well sadly that would be impossible with limited precision of lb :((",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1677097,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "02/05/2022 14:26:45",
      "content": "<p>I guess it is common practice in identification tasks: we have ID intersection between train and test, also, we have some ids only in train and some ids only in test (11% in public LB)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1677263,
          "author_name": "olegsidorshin",
          "author_url": "",
          "post_date": "02/05/2022 15:56:29",
          "content": "<p>Yeah, I think it is not a big deal but decided to share so everybody knows</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1677750,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "02/06/2022 00:46:51",
      "content": "<p>Thanks for the effort in this topic.<br>\nHow to approach the assignment of \"new_individual\" in some test images?<br>\nI thought to set a threshold on the probabilities and do a grid search on the CV score to see what threshold optimizes the MAP@5 CV.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1678005,
          "author_name": "olegsidorshin",
          "author_url": "",
          "post_date": "02/06/2022 07:56:41",
          "content": "<p>I think that this is a good basic approach. In previous Happywhale competition some people were using sophisticated post-processing techniques and standalone new_whale/non new_whale classifier models just for that, so it would be interesting to check them out aswell </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678129,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/06/2022 09:18:46",
          "content": "<p>In the last competition, there was a class for \"new whales\", we don't have this time, I think the new_whale/non new_whale classifier maybe not be possible here. But yes some kind of post-processing can be used</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678552,
          "author_name": "tedcheese",
          "author_url": "",
          "post_date": "02/06/2022 16:03:16",
          "content": "<p>there is a new_whale class here like in the previous competition -- this is essential to make the winning algorithms applicable in the real world where we often are discovering new individuals to add to our known set of whales and dolphins</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1683498,
          "author_name": "yousof9",
          "author_url": "",
          "post_date": "02/09/2022 20:32:46",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tedcheese\" target=\"_blank\">@tedcheese</a>,</p>\n<p>Is the distribution of new individuals similar in the public and private test sets?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1683518,
          "author_name": "tedcheese",
          "author_url": "",
          "post_date": "02/09/2022 20:51:07",
          "content": "<p>I'll refer this question to <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690953,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/15/2022 07:06:40",
          "content": "<p><code>there is a new_whale class here like in the previous competition</code> but there are none in the train data this time, is there some reason this time<br>\n<a href=\"https://www.kaggle.com/tedcheese\" target=\"_blank\">@tedcheese</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692570,
          "author_name": "tedcheese",
          "author_url": "",
          "post_date": "02/16/2022 05:48:50",
          "content": "<p>Good point <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a>. That was a bit of a fail in the previous competition because if a train data whale is a new_whale class, we're a priori stating that there's no known match in the test set. But in the real world we'd rarely know a test whale couldn't be any of the known train whale, only would know it if the test whale was a known age whale younger than some of the train whales, or vice versa a deceased whale (ie a necropsy photo) from a date past that couldn't match to more recent living whales.</p>\n<p>So, this competition is set to be more reflective of realty.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692587,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/16/2022 06:10:01",
          "content": "<p>I see, thats actually great thinking! we could know if our models could really be useful in the real world.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692895,
          "author_name": "vsedelnik",
          "author_url": "",
          "post_date": "02/16/2022 10:33:30",
          "content": "<p>There are a lot of individuals appearing only once in the train set and those of them who you put to your local val set are the \"new whale\" class for the local training effectively.<br>\nYou just need to properly calculate the CV metric to account for this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1684147,
      "author_name": "chrisrichardmiles",
      "author_url": "",
      "post_date": "02/10/2022 09:57:08",
      "content": "<p>I think it would be great if they did a time split. This is the true nature of the problem anyway; we use data from before now to build models that predict data after now. This could also explain the top occurring ids not being in the test set. Those whales simply could have passed away. If it is not a time split, I think it's a little frustrating that they would purposefully do things like put all the images for the top occurring ids in the train set. Why would you do this? You would be punishing models that learned well from lots of data. But if they did use a time series split, why not just say so? </p>\n<p>Either way, it would be most beneficial if the organizers could be more explicit in how they did the data split. Otherwise, there will be lots of brain power and submissions committed to probing the leaderboard instead of solving the problem as the organizers want. In fact, I can't think of a better use of submissions other than going down the list of the most common ids and submitting just so I can see what proportion they are of the lb test data. Because the most important step is creating a proper validation strategy. </p>\n<p>Please, <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>, shed some light on the data split.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1684157,
          "author_name": "olegsidorshin",
          "author_url": "",
          "post_date": "02/10/2022 10:05:50",
          "content": "<p>I agree with your comments about importance of details of the data split. I think that train/test split details should be clearly stated in every competition, considering that the entire modern machine learning relies on identical train/test distribution assumption.</p>\n<p>But I think that time split is a little bit redundant in this task. Without the clear time split, models should be able to figure more time-independent factors for identifying animals, which could prove more useful in practical applications (i.e. using the models for a long time without retraining with newer data, and overall more reliable perhaps?)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689145,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "02/14/2022 04:12:42",
      "content": "<p>Very helpful, thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1691815,
      "author_name": "jackchungchiehyu",
      "author_url": "",
      "post_date": "02/15/2022 16:19:46",
      "content": "<blockquote>\n  <p>If data was randomly split, those 1 photo ind. ids should be somewhat similar to new_individual present in test data distribution-wise: if we just take random validation split and mark all ids that are not present in the train as new_individual, the majority of those new_individuals would be 1 photo only samples.</p>\n</blockquote>\n<p>But if these whales were to appear in the public or private test set, they would not actually have the label \"new_individual\", right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1712736,
      "author_name": "namtranase",
      "author_url": "",
      "post_date": "03/05/2022 09:00:54",
      "content": "<p>Thank you for sharing the insightful information!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1715473,
      "author_name": "xstargate",
      "author_url": "",
      "post_date": "03/08/2022 03:45:13",
      "content": "<p>I wonder if someone did a complete LB probing to find out all individuals that do not exist in the public LB?😂😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 1715930,
          "author_name": "olegsidorshin",
          "author_url": "",
          "post_date": "03/08/2022 14:05:20",
          "content": "<p>Well sadly that would be impossible with limited precision of lb :((</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1676854": "There are `51033` images in train dataset and `27956` in test dataset; `78989` in total.\n\nThere are `15587` different individual ids in train data, `9258` (~18% of photos) with 1 photo only. If data was randomly split, those 1 photo ind. ids should be somewhat similar to new_individual present in test data distribution-wise: if we just take random validation split and mark all ids that are not present in the train as new_individual, the majority of those new_individuals would be 1 photo only samples.\n\nBecause of how target metric works, we can directly estimate how many new_individual images there are in LB split by just submitting file with all new_individual predictions, which gives `0.112` LB score.\n\nThat means there are around `3131` (~11%) new_individual photos, which is very different from 18% present in train.\n\nAnother interesting thing to note is that there are individuals that have lots of images in train, but are not present in public LB whatsoever (those are top 2 highest train images ids):\n\n| id | Imgs in train | Percentage in train | Percentage in public LB |\n| --- | --- | --- | --- |\n| 37c7aba965a5 | 400 | 0.007838 | 0.000 |\n| 114207cab555 | 168 | 0.003292 | 0.000 |\n\nAll of this means that train/test split is not truly random - some specific choice criteria were applied to get it. It would be good for competition hosts to clarify this and perhaps share splitting algorithm details in order to make validation approaches more stable and less luck-dependent.",
    "1677097": "I guess it is common practice in identification tasks: we have ID intersection between train and test, also, we have some ids only in train and some ids only in test (11% in public LB)",
    "1677263": "Yeah, I think it is not a big deal but decided to share so everybody knows",
    "1677750": "Thanks for the effort in this topic.\nHow to approach the assignment of \"new_individual\" in some test images?\nI thought to set a threshold on the probabilities and do a grid search on the CV score to see what threshold optimizes the MAP@5 CV.",
    "1678005": "I think that this is a good basic approach. In previous Happywhale competition some people were using sophisticated post-processing techniques and standalone new_whale/non new_whale classifier models just for that, so it would be interesting to check them out aswell",
    "1678129": "In the last competition, there was a class for \"new whales\", we don't have this time, I think the new_whale/non new_whale classifier maybe not be possible here. But yes some kind of post-processing can be used",
    "1678552": "there is a new_whale class here like in the previous competition -- this is essential to make the winning algorithms applicable in the real world where we often are discovering new individuals to add to our known set of whales and dolphins",
    "1683498": "Hi @tedcheese,\n\nIs the distribution of new individuals similar in the public and private test sets?",
    "1683518": "I'll refer this question to @inversion :-)",
    "1684147": "I think it would be great if they did a time split. This is the true nature of the problem anyway; we use data from before now to build models that predict data after now. This could also explain the top occurring ids not being in the test set. Those whales simply could have passed away. If it is not a time split, I think it's a little frustrating that they would purposefully do things like put all the images for the top occurring ids in the train set. Why would you do this? You would be punishing models that learned well from lots of data. But if they did use a time series split, why not just say so? \n\nEither way, it would be most beneficial if the organizers could be more explicit in how they did the data split. Otherwise, there will be lots of brain power and submissions committed to probing the leaderboard instead of solving the problem as the organizers want. In fact, I can't think of a better use of submissions other than going down the list of the most common ids and submitting just so I can see what proportion they are of the lb test data. Because the most important step is creating a proper validation strategy. \n\nPlease, @inversion, shed some light on the data split.",
    "1684157": "I agree with your comments about importance of details of the data split. I think that train/test split details should be clearly stated in every competition, considering that the entire modern machine learning relies on identical train/test distribution assumption.\n\nBut I think that time split is a little bit redundant in this task. Without the clear time split, models should be able to figure more time-independent factors for identifying animals, which could prove more useful in practical applications (i.e. using the models for a long time without retraining with newer data, and overall more reliable perhaps?)",
    "1689145": "Very helpful, thank you!",
    "1690953": "`there is a new_whale class here like in the previous competition` but there are none in the train data this time, is there some reason this time\n@tedcheese",
    "1691815": "> If data was randomly split, those 1 photo ind. ids should be somewhat similar to new_individual present in test data distribution-wise: if we just take random validation split and mark all ids that are not present in the train as new_individual, the majority of those new_individuals would be 1 photo only samples.\n\nBut if these whales were to appear in the public or private test set, they would not actually have the label \"new_individual\", right?",
    "1692570": "Good point @mrinath. That was a bit of a fail in the previous competition because if a train data whale is a new_whale class, we're a priori stating that there's no known match in the test set. But in the real world we'd rarely know a test whale couldn't be any of the known train whale, only would know it if the test whale was a known age whale younger than some of the train whales, or vice versa a deceased whale (ie a necropsy photo) from a date past that couldn't match to more recent living whales.\n\nSo, this competition is set to be more reflective of realty.",
    "1692587": "I see, thats actually great thinking! we could know if our models could really be useful in the real world.",
    "1692895": "There are a lot of individuals appearing only once in the train set and those of them who you put to your local val set are the \"new whale\" class for the local training effectively.\nYou just need to properly calculate the CV metric to account for this.",
    "1712736": "Thank you for sharing the insightful information!",
    "1715473": "I wonder if someone did a complete LB probing to find out all individuals that do not exist in the public LB?😂😂",
    "1715930": "Well sadly that would be impossible with limited precision of lb :(("
  },
  "source": "meta"
}