{
  "id": 314165,
  "title": "About test data",
  "url": "/competitions/happy-whale-and-dolphin/discussion/314165",
  "author_name": "",
  "post_date": "2022-03-21T09:58:59.806686400Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Sorry for the noob question here, I am quite new to such competitions and I want to ask one basic question:</p>\n<p>How to utilize preprocessed test dataset? It is clear to me that if we use preprocessed train dataset, we need to apply the same transformation on the test dataset in order to make reliable predictions.  </p>\n<p>But one thing is not clear to me. How the private test set will be processed, coz I see many of you are using  preprocessed test datasets like Detic Crop (and many other datasets) in order to make a submission? It is not clear to me, how should I repeat the steps on the private test set in order to make it in the same format as the public test set which is now available ? Should I use the pretrained detector (speaking of detic crop) for example and to make the crops as the data comes by,  or should I use the original test set which is given and only apply the resizing?</p>\n<p>Thank you in advance !<br>\nMs</p>",
  "messages": [
    {
      "id": "1730474",
      "postDate": "03/21/2022 09:58:59",
      "content": "<p>Sorry for the noob question here, I am quite new to such competitions and I want to ask one basic question:</p>\n<p>How to utilize preprocessed test dataset? It is clear to me that if we use preprocessed train dataset, we need to apply the same transformation on the test dataset in order to make reliable predictions.  </p>\n<p>But one thing is not clear to me. How the private test set will be processed, coz I see many of you are using  preprocessed test datasets like Detic Crop (and many other datasets) in order to make a submission? It is not clear to me, how should I repeat the steps on the private test set in order to make it in the same format as the public test set which is now available ? Should I use the pretrained detector (speaking of detic crop) for example and to make the crops as the data comes by,  or should I use the original test set which is given and only apply the resizing?</p>\n<p>Thank you in advance !<br>\nMs</p>",
      "rawMarkdown": "Sorry for the noob question here, I am quite new to such competitions and I want to ask one basic question:\n\nHow to utilize preprocessed test dataset? It is clear to me that if we use preprocessed train dataset, we need to apply the same transformation on the test dataset in order to make reliable predictions.  \n\nBut one thing is not clear to me. How the private test set will be processed, coz I see many of you are using  preprocessed test datasets like Detic Crop (and many other datasets) in order to make a submission? It is not clear to me, how should I repeat the steps on the private test set in order to make it in the same format as the public test set which is now available ? Should I use the pretrained detector (speaking of detic crop) for example and to make the crops as the data comes by,  or should I use the original test set which is given and only apply the resizing?\n\nThank you in advance !\nMs",
      "votes": null
    },
    {
      "id": "1730477",
      "postDate": "03/21/2022 10:02:47",
      "content": "<p>The <strong>Private test data</strong> you said was already provided in 27956 images.<br>\nIn short, if you submit your solution (27956 rows) in submission.csv, your private score has already been computed</p>",
      "rawMarkdown": "The **Private test data** you said was already provided in 27956 images.\nIn short, if you submit your solution (27956 rows) in submission.csv, your private score has already been computed",
      "votes": null
    },
    {
      "id": "1730486",
      "postDate": "03/21/2022 10:17:09",
      "content": "<p>Oh I noticed that , the score popped up immediately after extracting the embeddings and after making the predictions, which appeared strange to me. I know we are using 24% of the data for now (the 27956 images), but what about the remaining 76% ?  I was talking about creating an inference notebook which will apply the same transformations (as for private Detic crop) on the other 76% of the   dataset (when they will compute the final standings) ? This means that I am using other preprocessed test dataset , and after that when they evaluate the final standings on the other part of the test set, the format of the images will be different and unlike the preprocessed test set which i am using now to make submission. Thanks mate, this final question will clarify the things !</p>",
      "rawMarkdown": "Oh I noticed that , the score popped up immediately after extracting the embeddings and after making the predictions, which appeared strange to me. I know we are using 24% of the data for now (the 27956 images), but what about the remaining 76% ?  I was talking about creating an inference notebook which will apply the same transformations (as for private Detic crop) on the other 76% of the   dataset (when they will compute the final standings) ? This means that I am using other preprocessed test dataset , and after that when they evaluate the final standings on the other part of the test set, the format of the images will be different and unlike the preprocessed test set which i am using now to make submission. Thanks mate, this final question will clarify the things !",
      "votes": null
    },
    {
      "id": "1731151",
      "postDate": "03/22/2022 03:11:25",
      "content": "<p>27956 is the total test set (both private and public).<br>\n24% of public test data is 24% of 27956 ~= 6704 images<br>\n76% of prive test is 76% of 27956 ~= 21246 images</p>",
      "rawMarkdown": "27956 is the total test set (both private and public).\n24% of public test data is 24% of 27956 ~= 6704 images\n76% of prive test is 76% of 27956 ~= 21246 images",
      "votes": null
    },
    {
      "id": "1735622",
      "postDate": "03/26/2022 12:52:18",
      "content": "<p><a href=\"https://www.kaggle.com/ptran1203\" target=\"_blank\">@ptran1203</a> The private dataset in the submission is computed simultaneously with the public dataset. It is just that the kaggle system hides the private leaderboard for all participants until the end of the competition and only displays the public leaderboard 24%. The private data (76%) is actually run at the same time as public data, so there is no need to restart it yourself</p>",
      "rawMarkdown": "ptran1203 The private dataset in the submission is computed simultaneously with the public dataset. It is just that the kaggle system hides the private leaderboard for all participants until the end of the competition and only displays the public leaderboard 24%. The private data (76%) is actually run at the same time as public data, so there is no need to restart it yourself",
      "votes": null
    },
    {
      "id": "1735658",
      "postDate": "03/26/2022 13:40:57",
      "content": "<p>Thanks Ksenia for the clarification :)!</p>",
      "rawMarkdown": "Thanks Ksenia for the clarification :)!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1730477,
      "author_name": "ptran1203",
      "author_url": "",
      "post_date": "03/21/2022 10:02:47",
      "content": "<p>The <strong>Private test data</strong> you said was already provided in 27956 images.<br>\nIn short, if you submit your solution (27956 rows) in submission.csv, your private score has already been computed</p>",
      "votes": null,
      "replies": [
        {
          "id": 1730486,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "03/21/2022 10:17:09",
          "content": "<p>Oh I noticed that , the score popped up immediately after extracting the embeddings and after making the predictions, which appeared strange to me. I know we are using 24% of the data for now (the 27956 images), but what about the remaining 76% ?  I was talking about creating an inference notebook which will apply the same transformations (as for private Detic crop) on the other 76% of the   dataset (when they will compute the final standings) ? This means that I am using other preprocessed test dataset , and after that when they evaluate the final standings on the other part of the test set, the format of the images will be different and unlike the preprocessed test set which i am using now to make submission. Thanks mate, this final question will clarify the things !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1731151,
          "author_name": "ptran1203",
          "author_url": "",
          "post_date": "03/22/2022 03:11:25",
          "content": "<p>27956 is the total test set (both private and public).<br>\n24% of public test data is 24% of 27956 ~= 6704 images<br>\n76% of prive test is 76% of 27956 ~= 21246 images</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1735622,
      "author_name": "ivanovakm",
      "author_url": "",
      "post_date": "03/26/2022 12:52:18",
      "content": "<p><a href=\"https://www.kaggle.com/ptran1203\" target=\"_blank\">@ptran1203</a> The private dataset in the submission is computed simultaneously with the public dataset. It is just that the kaggle system hides the private leaderboard for all participants until the end of the competition and only displays the public leaderboard 24%. The private data (76%) is actually run at the same time as public data, so there is no need to restart it yourself</p>",
      "votes": null,
      "replies": [
        {
          "id": 1735658,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "03/26/2022 13:40:57",
          "content": "<p>Thanks Ksenia for the clarification :)!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1730474": "Sorry for the noob question here, I am quite new to such competitions and I want to ask one basic question:\n\nHow to utilize preprocessed test dataset? It is clear to me that if we use preprocessed train dataset, we need to apply the same transformation on the test dataset in order to make reliable predictions.  \n\nBut one thing is not clear to me. How the private test set will be processed, coz I see many of you are using  preprocessed test datasets like Detic Crop (and many other datasets) in order to make a submission? It is not clear to me, how should I repeat the steps on the private test set in order to make it in the same format as the public test set which is now available ? Should I use the pretrained detector (speaking of detic crop) for example and to make the crops as the data comes by,  or should I use the original test set which is given and only apply the resizing?\n\nThank you in advance !\nMs",
    "1730477": "The **Private test data** you said was already provided in 27956 images.\nIn short, if you submit your solution (27956 rows) in submission.csv, your private score has already been computed",
    "1730486": "Oh I noticed that , the score popped up immediately after extracting the embeddings and after making the predictions, which appeared strange to me. I know we are using 24% of the data for now (the 27956 images), but what about the remaining 76% ?  I was talking about creating an inference notebook which will apply the same transformations (as for private Detic crop) on the other 76% of the   dataset (when they will compute the final standings) ? This means that I am using other preprocessed test dataset , and after that when they evaluate the final standings on the other part of the test set, the format of the images will be different and unlike the preprocessed test set which i am using now to make submission. Thanks mate, this final question will clarify the things !",
    "1731151": "27956 is the total test set (both private and public).\n24% of public test data is 24% of 27956 ~= 6704 images\n76% of prive test is 76% of 27956 ~= 21246 images",
    "1735622": "ptran1203 The private dataset in the submission is computed simultaneously with the public dataset. It is just that the kaggle system hides the private leaderboard for all participants until the end of the competition and only displays the public leaderboard 24%. The private data (76%) is actually run at the same time as public data, so there is no need to restart it yourself",
    "1735658": "Thanks Ksenia for the clarification :)!"
  },
  "source": "meta"
}