{
  "id": 536814,
  "title": "Notebook ran successfully but scoring failed",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/536814",
  "author_name": "",
  "post_date": "2024-09-30T02:32:20.323691400Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Why would scoring fail if the notebook runs successfully  ? <br>\nNotebook ran successfully but scoring failed. Tried using the sample submission instead and changing the sii column, same.<br>\n<a href=\"https://www.kaggle.com/code/corneliuantonescu/successful-run-but-scoring-failed\" target=\"_blank\">https://www.kaggle.com/code/corneliuantonescu/successful-run-but-scoring-failed</a><br>\nWhat is missing ?</p>",
  "messages": [
    {
      "id": "3002410",
      "postDate": "09/30/2024 02:32:20",
      "content": "<p>Why would scoring fail if the notebook runs successfully  ? <br>\nNotebook ran successfully but scoring failed. Tried using the sample submission instead and changing the sii column, same.<br>\n<a href=\"https://www.kaggle.com/code/corneliuantonescu/successful-run-but-scoring-failed\" target=\"_blank\">https://www.kaggle.com/code/corneliuantonescu/successful-run-but-scoring-failed</a><br>\nWhat is missing ?</p>",
      "rawMarkdown": "Why would scoring fail if the notebook runs successfully  ? \nNotebook ran successfully but scoring failed. Tried using the sample submission instead and changing the sii column, same.\nhttps://www.kaggle.com/code/corneliuantonescu/successful-run-but-scoring-failed\nWhat is missing ?",
      "votes": null
    },
    {
      "id": "3003132",
      "postDate": "09/30/2024 17:22:53",
      "content": "<p>The test dataset is actually hidden, which means that the actual number of rows in submission.csv is not 20, unlike the number of rows in sample_submission.csv.<br>\nTherefore, you should dynamically predict the test values by using the paths of the input test dataset.</p>",
      "rawMarkdown": "The test dataset is actually hidden, which means that the actual number of rows in submission.csv is not 20, unlike the number of rows in sample_submission.csv.\nTherefore, you should dynamically predict the test values by using the paths of the input test dataset.",
      "votes": null
    },
    {
      "id": "3003246",
      "postDate": "09/30/2024 19:42:35",
      "content": "<p>Thank you, I thought I am already doing that by using<br>\n<code>test=pd.read_csv('../input/child-mind-institute-problematic-internet-use/test.csv')</code><br>\nand<br>\n<code>randomsii=train_mock_df['sii'].sample(n=test.shape[0],ignore_index=True)\nsubmission=pd.DataFrame({'id': test['id'].values,'sii':randomsii })</code><br>\nCan you clarify please ?</p>",
      "rawMarkdown": "Thank you, I thought I am already doing that by using\n`test=pd.read_csv('../input/child-mind-institute-problematic-internet-use/test.csv')`\nand\n`randomsii=train_mock_df['sii'].sample(n=test.shape[0],ignore_index=True)\nsubmission=pd.DataFrame({'id': test['id'].values,'sii':randomsii })`\nCan you clarify please ?",
      "votes": null
    },
    {
      "id": "3003310",
      "postDate": "09/30/2024 21:25:50",
      "content": "<p><a href=\"https://www.kaggle.com/corneliuantonescu\" target=\"_blank\">@corneliuantonescu</a> The hidden test set has about 3800 rows which is much larger than your <code>train_mock_df</code>. You'd need to call <code>pd.DataFrame.sample</code> with <code>replace=True</code>.</p>",
      "rawMarkdown": "corneliuantonescu The hidden test set has about 3800 rows which is much larger than your `train_mock_df`. You'd need to call `pd.DataFrame.sample` with `replace=True`.",
      "votes": null
    },
    {
      "id": "3003321",
      "postDate": "09/30/2024 21:52:24",
      "content": "<p>You're right, I missed the fact that  the made up train set that I was sampling was smaller than the test set. I increased the size (about 100x) and it worked. Thank you!</p>",
      "rawMarkdown": "You're right, I missed the fact that  the made up train set that I was sampling was smaller than the test set. I increased the size (about 100x) and it worked. Thank you!",
      "votes": null
    },
    {
      "id": "3012476",
      "postDate": "10/09/2024 04:56:25",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/donwlee\" target=\"_blank\">@donwlee</a> ! I was doing the same mistake </p>",
      "rawMarkdown": "Thanks @donwlee ! I was doing the same mistake",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3003132,
      "author_name": "donwlee",
      "author_url": "",
      "post_date": "09/30/2024 17:22:53",
      "content": "<p>The test dataset is actually hidden, which means that the actual number of rows in submission.csv is not 20, unlike the number of rows in sample_submission.csv.<br>\nTherefore, you should dynamically predict the test values by using the paths of the input test dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3003246,
          "author_name": "corneliuantonescu",
          "author_url": "",
          "post_date": "09/30/2024 19:42:35",
          "content": "<p>Thank you, I thought I am already doing that by using<br>\n<code>test=pd.read_csv('../input/child-mind-institute-problematic-internet-use/test.csv')</code><br>\nand<br>\n<code>randomsii=train_mock_df['sii'].sample(n=test.shape[0],ignore_index=True)\nsubmission=pd.DataFrame({'id': test['id'].values,'sii':randomsii })</code><br>\nCan you clarify please ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3003310,
              "author_name": "siukeitin",
              "author_url": "",
              "post_date": "09/30/2024 21:25:50",
              "content": "<p><a href=\"https://www.kaggle.com/corneliuantonescu\" target=\"_blank\">@corneliuantonescu</a> The hidden test set has about 3800 rows which is much larger than your <code>train_mock_df</code>. You'd need to call <code>pd.DataFrame.sample</code> with <code>replace=True</code>.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3003321,
                  "author_name": "corneliuantonescu",
                  "author_url": "",
                  "post_date": "09/30/2024 21:52:24",
                  "content": "<p>You're right, I missed the fact that  the made up train set that I was sampling was smaller than the test set. I increased the size (about 100x) and it worked. Thank you!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3012476,
          "author_name": "polygot13",
          "author_url": "",
          "post_date": "10/09/2024 04:56:25",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/donwlee\" target=\"_blank\">@donwlee</a> ! I was doing the same mistake </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3002410": "Why would scoring fail if the notebook runs successfully  ? \nNotebook ran successfully but scoring failed. Tried using the sample submission instead and changing the sii column, same.\nhttps://www.kaggle.com/code/corneliuantonescu/successful-run-but-scoring-failed\nWhat is missing ?",
    "3003132": "The test dataset is actually hidden, which means that the actual number of rows in submission.csv is not 20, unlike the number of rows in sample_submission.csv.\nTherefore, you should dynamically predict the test values by using the paths of the input test dataset.",
    "3003246": "Thank you, I thought I am already doing that by using\n`test=pd.read_csv('../input/child-mind-institute-problematic-internet-use/test.csv')`\nand\n`randomsii=train_mock_df['sii'].sample(n=test.shape[0],ignore_index=True)\nsubmission=pd.DataFrame({'id': test['id'].values,'sii':randomsii })`\nCan you clarify please ?",
    "3003310": "corneliuantonescu The hidden test set has about 3800 rows which is much larger than your `train_mock_df`. You'd need to call `pd.DataFrame.sample` with `replace=True`.",
    "3003321": "You're right, I missed the fact that  the made up train set that I was sampling was smaller than the test set. I increased the size (about 100x) and it worked. Thank you!",
    "3012476": "Thanks @donwlee ! I was doing the same mistake"
  },
  "source": "meta"
}