{
  "id": 537135,
  "title": " I didn't quite understand the hidden test dataset concept",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/537135",
  "author_name": "",
  "post_date": "2024-10-01T15:05:02.520094200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I usually participate in smaller competitions where I submit a submission.csv file, and those datasets are easy to work with. However, for this competition, I wanted to test my beginner skills, but I couldn't quite grasp the submission format.<br>\nHere's what I thought: for each ID in the test.csv file, I assumed that there would be a corresponding hidden parquet file located at …/series_test.parquet/id=00115b9f/part-0.parquet, where I would extract time-dependent features. I wrote my code based on this assumption. However, when I ran my code, I realized that there are no hidden parquet files in the test.csv dataset as I had imagined. Doesn't the competition provide a parquet file for each ID in the test.csv file?<br>\nCould you please review my code and explain why it failed? I believe that some of you, at a beginner level like me, might have experienced a similar issue and will understand my confusion. I would appreciate any help or understanding you can provide. I hope my confusion will be clear once you review my code. Thank you in advance!<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/iamnotindangerskyler/childmindtr</a></p>",
  "messages": [
    {
      "id": "3004111",
      "postDate": "10/01/2024 15:05:02",
      "content": "<p>I usually participate in smaller competitions where I submit a submission.csv file, and those datasets are easy to work with. However, for this competition, I wanted to test my beginner skills, but I couldn't quite grasp the submission format.<br>\nHere's what I thought: for each ID in the test.csv file, I assumed that there would be a corresponding hidden parquet file located at …/series_test.parquet/id=00115b9f/part-0.parquet, where I would extract time-dependent features. I wrote my code based on this assumption. However, when I ran my code, I realized that there are no hidden parquet files in the test.csv dataset as I had imagined. Doesn't the competition provide a parquet file for each ID in the test.csv file?<br>\nCould you please review my code and explain why it failed? I believe that some of you, at a beginner level like me, might have experienced a similar issue and will understand my confusion. I would appreciate any help or understanding you can provide. I hope my confusion will be clear once you review my code. Thank you in advance!<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/iamnotindangerskyler/childmindtr</a></p>",
      "rawMarkdown": "I usually participate in smaller competitions where I submit a submission.csv file, and those datasets are easy to work with. However, for this competition, I wanted to test my beginner skills, but I couldn't quite grasp the submission format.\n\nHere's what I thought: for each ID in the test.csv file, I assumed that there would be a corresponding hidden parquet file located at .../series_test.parquet/id=00115b9f/part-0.parquet, where I would extract time-dependent features. I wrote my code based on this assumption. However, when I ran my code, I realized that there are no hidden parquet files in the test.csv dataset as I had imagined. Doesn't the competition provide a parquet file for each ID in the test.csv file?\n\nCould you please review my code and explain why it failed? I believe that some of you, at a beginner level like me, might have experienced a similar issue and will understand my confusion. I would appreciate any help or understanding you can provide. I hope my confusion will be clear once you review my code. Thank you in advance!\n\n[https://www.kaggle.com/code/iamnotindangerskyler/childmindtr](url)",
      "votes": null
    },
    {
      "id": "3004392",
      "postDate": "10/01/2024 19:09:44",
      "content": "<p>You may consider this as a proxy for a real life situation where you don't know the actual test set and have to build a model for <strong>unseen data</strong>. The 3 rows provided to you are for syntax check only <a href=\"https://www.kaggle.com/iamnotindangerskyler\" target=\"_blank\">@iamnotindangerskyler</a> </p>",
      "rawMarkdown": "You may consider this as a proxy for a real life situation where you don't know the actual test set and have to build a model for **unseen data**. The 3 rows provided to you are for syntax check only @iamnotindangerskyler",
      "votes": null
    },
    {
      "id": "3004409",
      "postDate": "10/01/2024 19:34:47",
      "content": "<p>We will create and submit a model that makes a prediction for the 'sii' category for each 'id' in the test.csv file. However, in this competition format, there are more test 'ids' checked in the background. I understood everything up to this point, but I thought there would be hidden parquet files for some of the test 'ids.' Isn't that a wrong assumption? There are only 2 test parquet files, no extras, right? This means that the training dataset contains less information than the test dataset, specifically regarding the features that can be extracted from the parquet files. This actually requires us to make double predictions. Did I understand correctly? Are there only 2 parquet files for the test 'ids'?</p>",
      "rawMarkdown": "We will create and submit a model that makes a prediction for the 'sii' category for each 'id' in the test.csv file. However, in this competition format, there are more test 'ids' checked in the background. I understood everything up to this point, but I thought there would be hidden parquet files for some of the test 'ids.' Isn't that a wrong assumption? There are only 2 test parquet files, no extras, right? This means that the training dataset contains less information than the test dataset, specifically regarding the features that can be extracted from the parquet files. This actually requires us to make double predictions. Did I understand correctly? Are there only 2 parquet files for the test 'ids'?",
      "votes": null
    },
    {
      "id": "3004425",
      "postDate": "10/01/2024 20:09:46",
      "content": "<p>The test.csv shown to you is not the test set they feed into your code.<br>\nIf you look carefully, you will find that 11 out of 20 id in the test set actually come from train set.</p>\n<p>In the \"real\" test set that is hidden, I guess some of them have parquet and some don't.<br>\nSo in your code you can't call a specific id to get a parquet.</p>\n<p>Below is an excellent and understandable notebook by <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> that I will recommend you to have a read. His way in extracting the parquet is elegant.<br>\n<a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-model\" target=\"_blank\">https://www.kaggle.com/code/abdmental01/cmi-best-single-model</a></p>\n<p>Good luck.</p>",
      "rawMarkdown": "The test.csv shown to you is not the test set they feed into your code.\nIf you look carefully, you will find that 11 out of 20 id in the test set actually come from train set.\n\nIn the \"real\" test set that is hidden, I guess some of them have parquet and some don't.\nSo in your code you can't call a specific id to get a parquet.\n\nBelow is an excellent and understandable notebook by @abdmental01 that I will recommend you to have a read. His way in extracting the parquet is elegant.\n[https://www.kaggle.com/code/abdmental01/cmi-best-single-model](https://www.kaggle.com/code/abdmental01/cmi-best-single-model)\n\nGood luck.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3004392,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/01/2024 19:09:44",
      "content": "<p>You may consider this as a proxy for a real life situation where you don't know the actual test set and have to build a model for <strong>unseen data</strong>. The 3 rows provided to you are for syntax check only <a href=\"https://www.kaggle.com/iamnotindangerskyler\" target=\"_blank\">@iamnotindangerskyler</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3004409,
          "author_name": "iamnotindangerskyler",
          "author_url": "",
          "post_date": "10/01/2024 19:34:47",
          "content": "<p>We will create and submit a model that makes a prediction for the 'sii' category for each 'id' in the test.csv file. However, in this competition format, there are more test 'ids' checked in the background. I understood everything up to this point, but I thought there would be hidden parquet files for some of the test 'ids.' Isn't that a wrong assumption? There are only 2 test parquet files, no extras, right? This means that the training dataset contains less information than the test dataset, specifically regarding the features that can be extracted from the parquet files. This actually requires us to make double predictions. Did I understand correctly? Are there only 2 parquet files for the test 'ids'?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3004425,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "10/01/2024 20:09:46",
      "content": "<p>The test.csv shown to you is not the test set they feed into your code.<br>\nIf you look carefully, you will find that 11 out of 20 id in the test set actually come from train set.</p>\n<p>In the \"real\" test set that is hidden, I guess some of them have parquet and some don't.<br>\nSo in your code you can't call a specific id to get a parquet.</p>\n<p>Below is an excellent and understandable notebook by <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> that I will recommend you to have a read. His way in extracting the parquet is elegant.<br>\n<a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-model\" target=\"_blank\">https://www.kaggle.com/code/abdmental01/cmi-best-single-model</a></p>\n<p>Good luck.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3004111": "I usually participate in smaller competitions where I submit a submission.csv file, and those datasets are easy to work with. However, for this competition, I wanted to test my beginner skills, but I couldn't quite grasp the submission format.\n\nHere's what I thought: for each ID in the test.csv file, I assumed that there would be a corresponding hidden parquet file located at .../series_test.parquet/id=00115b9f/part-0.parquet, where I would extract time-dependent features. I wrote my code based on this assumption. However, when I ran my code, I realized that there are no hidden parquet files in the test.csv dataset as I had imagined. Doesn't the competition provide a parquet file for each ID in the test.csv file?\n\nCould you please review my code and explain why it failed? I believe that some of you, at a beginner level like me, might have experienced a similar issue and will understand my confusion. I would appreciate any help or understanding you can provide. I hope my confusion will be clear once you review my code. Thank you in advance!\n\n[https://www.kaggle.com/code/iamnotindangerskyler/childmindtr](url)",
    "3004392": "You may consider this as a proxy for a real life situation where you don't know the actual test set and have to build a model for **unseen data**. The 3 rows provided to you are for syntax check only @iamnotindangerskyler",
    "3004409": "We will create and submit a model that makes a prediction for the 'sii' category for each 'id' in the test.csv file. However, in this competition format, there are more test 'ids' checked in the background. I understood everything up to this point, but I thought there would be hidden parquet files for some of the test 'ids.' Isn't that a wrong assumption? There are only 2 test parquet files, no extras, right? This means that the training dataset contains less information than the test dataset, specifically regarding the features that can be extracted from the parquet files. This actually requires us to make double predictions. Did I understand correctly? Are there only 2 parquet files for the test 'ids'?",
    "3004425": "The test.csv shown to you is not the test set they feed into your code.\nIf you look carefully, you will find that 11 out of 20 id in the test set actually come from train set.\n\nIn the \"real\" test set that is hidden, I guess some of them have parquet and some don't.\nSo in your code you can't call a specific id to get a parquet.\n\nBelow is an excellent and understandable notebook by @abdmental01 that I will recommend you to have a read. His way in extracting the parquet is elegant.\n[https://www.kaggle.com/code/abdmental01/cmi-best-single-model](https://www.kaggle.com/code/abdmental01/cmi-best-single-model)\n\nGood luck."
  },
  "source": "meta"
}