{
  "id": 488334,
  "title": "why is the test data only 10 values",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/488334",
  "author_name": "",
  "post_date": "2024-04-02T04:25:55.852211200Z",
  "votes": -2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Can someone please tell me why the test data only contains 10 data points?</p>",
  "messages": [
    {
      "id": "2728027",
      "postDate": "04/02/2024 04:25:55",
      "content": "<p>Can someone please tell me why the test data only contains 10 data points?</p>",
      "rawMarkdown": "Can someone please tell me why the test data only contains 10 data points?",
      "votes": null
    },
    {
      "id": "2728110",
      "postDate": "04/02/2024 05:31:36",
      "content": "<p>Because this test data is only for you to test whether your program runs smoothly, when you submit the Notebook, the test data will be replaced with hidden data.</p>",
      "rawMarkdown": "Because this test data is only for you to test whether your program runs smoothly, when you submit the Notebook, the test data will be replaced with hidden data.",
      "votes": null
    },
    {
      "id": "2728174",
      "postDate": "04/02/2024 06:06:35",
      "content": "<p>how will that happen? I don't understand? Will the test_base.csv and other test datas be different?</p>",
      "rawMarkdown": "how will that happen? I don't understand? Will the test_base.csv and other test datas be different?",
      "votes": null
    },
    {
      "id": "2728995",
      "postDate": "04/02/2024 13:56:38",
      "content": "<p>when  submit the code, the test data will be replaced by the actual test data which is much more than 10. of which a certain portion will be used to evaluate the public score. the internal mechanics of how this exactly happens is up to the kaggle internal coding.</p>\n<p>This is done to secure the test data as hidden and can not be used in training the model, otherwise, the best scores can easily be obtained by training the model directly on test data.</p>\n<p>But 10 test data is provided so that we can check if the code is working smoothly or there is some error while predicting on test data.</p>\n<p>*elaboration to the already posted comment by <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
      "rawMarkdown": "when  submit the code, the test data will be replaced by the actual test data which is much more than 10. of which a certain portion will be used to evaluate the public score. the internal mechanics of how this exactly happens is up to the kaggle internal coding.\n\nThis is done to secure the test data as hidden and can not be used in training the model, otherwise, the best scores can easily be obtained by training the model directly on test data.\n\nBut 10 test data is provided so that we can check if the code is working smoothly or there is some error while predicting on test data.\n\n*elaboration to the already posted comment by @yunsuxiaozi",
      "votes": null
    },
    {
      "id": "2731918",
      "postDate": "04/02/2024 23:20:23",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "2734855",
      "postDate": "04/04/2024 11:38:09",
      "content": "<p>Yeah, Dataset of 10 points is too small</p>",
      "rawMarkdown": "Yeah, Dataset of 10 points is too small",
      "votes": null
    },
    {
      "id": "2752467",
      "postDate": "04/15/2024 01:43:32",
      "content": "<p>They will replace test-related files to origin's at the same path. so all you need to do is just testing only with 10 samples and be sure the results are listed in <code>submission.csv</code> file.</p>",
      "rawMarkdown": "They will replace test-related files to origin's at the same path. so all you need to do is just testing only with 10 samples and be sure the results are listed in `submission.csv` file.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2728110,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "04/02/2024 05:31:36",
      "content": "<p>Because this test data is only for you to test whether your program runs smoothly, when you submit the Notebook, the test data will be replaced with hidden data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2728174,
          "author_name": "farazj27",
          "author_url": "",
          "post_date": "04/02/2024 06:06:35",
          "content": "<p>how will that happen? I don't understand? Will the test_base.csv and other test datas be different?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2752467,
              "author_name": "kjihwan",
              "author_url": "",
              "post_date": "04/15/2024 01:43:32",
              "content": "<p>They will replace test-related files to origin's at the same path. so all you need to do is just testing only with 10 samples and be sure the results are listed in <code>submission.csv</code> file.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2728995,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "04/02/2024 13:56:38",
      "content": "<p>when  submit the code, the test data will be replaced by the actual test data which is much more than 10. of which a certain portion will be used to evaluate the public score. the internal mechanics of how this exactly happens is up to the kaggle internal coding.</p>\n<p>This is done to secure the test data as hidden and can not be used in training the model, otherwise, the best scores can easily be obtained by training the model directly on test data.</p>\n<p>But 10 test data is provided so that we can check if the code is working smoothly or there is some error while predicting on test data.</p>\n<p>*elaboration to the already posted comment by <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2731918,
          "author_name": "farazj27",
          "author_url": "",
          "post_date": "04/02/2024 23:20:23",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2734855,
      "author_name": "sultanovar",
      "author_url": "",
      "post_date": "04/04/2024 11:38:09",
      "content": "<p>Yeah, Dataset of 10 points is too small</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2728027": "Can someone please tell me why the test data only contains 10 data points?",
    "2728110": "Because this test data is only for you to test whether your program runs smoothly, when you submit the Notebook, the test data will be replaced with hidden data.",
    "2728174": "how will that happen? I don't understand? Will the test_base.csv and other test datas be different?",
    "2728995": "when  submit the code, the test data will be replaced by the actual test data which is much more than 10. of which a certain portion will be used to evaluate the public score. the internal mechanics of how this exactly happens is up to the kaggle internal coding.\n\nThis is done to secure the test data as hidden and can not be used in training the model, otherwise, the best scores can easily be obtained by training the model directly on test data.\n\nBut 10 test data is provided so that we can check if the code is working smoothly or there is some error while predicting on test data.\n\n*elaboration to the already posted comment by @yunsuxiaozi",
    "2731918": "Thanks a lot!",
    "2734855": "Yeah, Dataset of 10 points is too small",
    "2752467": "They will replace test-related files to origin's at the same path. so all you need to do is just testing only with 10 samples and be sure the results are listed in `submission.csv` file."
  },
  "source": "meta"
}