{
  "id": 537229,
  "title": "submission failed",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/537229",
  "author_name": "Yoonjae Lee",
  "post_date": "2024-10-02T03:49:31.370000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2F5bef936191eac533e0e6333362d8dccd%2Fsubmission_capture01.jpg?generation=1727840679437152&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2Fd74fabb548310ff360628531e5dab0a8%2Fsubmission_capture02.jpg?generation=1727840759667982&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2Fc5c8c5b61813d1578f3b124ba6f02095%2Fsubmission_capture03.jpg?generation=1727840924976492&amp;alt=media\" alt=\"\"></p>\n<p>I'm a newbie in ML competition. My submission.csv is looking good and has a right format I think. But after submitting my notebook, it took 2 hours for running and then failed. What am I missing??</p>",
  "messages": [
    {
      "id": 3007100,
      "postDate": "2024-10-05T00:40:50.960Z",
      "content": "<p><a href=\"https://www.kaggle.com/yoonjaelee\" target=\"_blank\">@yoonjaelee</a> and <a href=\"https://www.kaggle.com/karlbina\" target=\"_blank\">@karlbina</a> Here's what the statement at the bottom of the submission page says with some extra words that I've added -- I hope this makes the concept of a \"hidden test set\" with its own \"submission.csv\" clearer for you.</p>\n<p>\"In this competition, we will <strong>privately</strong> re-run your selected Notebook Version\" -- Privately means your notebook will be copied to some directory you cannot access and it will be run there.</p>\n<p>\"with a <strong>hidden test set</strong> substituted into the competition dataset.\" -- This means that in the private location, there will be the same \"train.csv\" file (and the same train parquet files), and there will also be a test.csv and test parquet files. <strong>But, this test.csv will be different, and longer,</strong> than the little test.csv that is in your working directory. Although the parquet directory for test is the same path as before, now it will have different parquet files, ones that go with (some of) the new test.csv id values. </p>\n<p>\"We then extract your chosen <strong>Output File</strong> from the re-run and use that to determine your score.\" -- When your notebook is run in private: i) it will read in the same training data, build features, and fit a model; ii) it will read in the new, hidden \"test.csv\" and the related parquet files, build the test features, and apply the trained-model to the test features; iii) the test predictions from the model are then written into a <strong>new submission.csv</strong> which has one row for each of the ids in the new, hidden test.csv. This created-in-private submission.csv is then scored based on the unknow-to-us correct answers for the hidden test.csv.  </p>",
      "rawMarkdown": "@yoonjaelee and @karlbina Here's what the statement at the bottom of the submission page says with some extra words that I've added -- I hope this makes the concept of a \"hidden test set\" with its own \"submission.csv\" clearer for you.\n\n\"In this competition, we will **privately** re-run your selected Notebook Version\" -- Privately means your notebook will be copied to some directory you cannot access and it will be run there.\n\n\"with a **hidden test set** substituted into the competition dataset.\" -- This means that in the private location, there will be the same \"train.csv\" file (and the same train parquet files), and there will also be a test.csv and test parquet files. **But, this test.csv will be different, and longer,** than the little test.csv that is in your working directory. Although the parquet directory for test is the same path as before, now it will have different parquet files, ones that go with (some of) the new test.csv id values. \n\n\"We then extract your chosen **Output File** from the re-run and use that to determine your score.\" -- When your notebook is run in private: i) it will read in the same training data, build features, and fit a model; ii) it will read in the new, hidden \"test.csv\" and the related parquet files, build the test features, and apply the trained-model to the test features; iii) the test predictions from the model are then written into a **new submission.csv** which has one row for each of the ids in the new, hidden test.csv. This created-in-private submission.csv is then scored based on the unknow-to-us correct answers for the hidden test.csv.  \n",
      "votes": 1
    },
    {
      "id": 3004615,
      "postDate": "2024-10-02T03:49:31.370Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2F5bef936191eac533e0e6333362d8dccd%2Fsubmission_capture01.jpg?generation=1727840679437152&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2Fd74fabb548310ff360628531e5dab0a8%2Fsubmission_capture02.jpg?generation=1727840759667982&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2Fc5c8c5b61813d1578f3b124ba6f02095%2Fsubmission_capture03.jpg?generation=1727840924976492&amp;alt=media\" alt=\"\"></p>\n<p>I'm a newbie in ML competition. My submission.csv is looking good and has a right format I think. But after submitting my notebook, it took 2 hours for running and then failed. What am I missing??</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2F5bef936191eac533e0e6333362d8dccd%2Fsubmission_capture01.jpg?generation=1727840679437152&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2Fd74fabb548310ff360628531e5dab0a8%2Fsubmission_capture02.jpg?generation=1727840759667982&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2Fc5c8c5b61813d1578f3b124ba6f02095%2Fsubmission_capture03.jpg?generation=1727840924976492&alt=media)\n\nI'm a newbie in ML competition. My submission.csv is looking good and has a right format I think. But after submitting my notebook, it took 2 hours for running and then failed. What am I missing??",
      "votes": 1
    },
    {
      "id": 3013659,
      "postDate": "2024-10-10T11:48:03.763Z",
      "content": "<p>I had this issue very early. For me, my issue was from using multi-class objective function instead of binary objective function. Really, the public set works with multi-class, but hidden set is binary.</p>\n<p>Other issues could be nan, inf, or negatives that arise. Or different features (i.e., OHE where hidden set has more uniques than train and test set).</p>\n<p>It's probably best to add some filters and conditionals in your program to handle exceptional cases, etc.</p>",
      "rawMarkdown": "I had this issue very early. For me, my issue was from using multi-class objective function instead of binary objective function. Really, the public set works with multi-class, but hidden set is binary.\n\nOther issues could be nan, inf, or negatives that arise. Or different features (i.e., OHE where hidden set has more uniques than train and test set).\n\nIt's probably best to add some filters and conditionals in your program to handle exceptional cases, etc."
    },
    {
      "id": 3007073,
      "postDate": "2024-10-04T22:32:47.983Z",
      "content": "<p>I also have the same problem as you, the submit file is good but it also failed in less than 30 minutes of execution.</p>",
      "rawMarkdown": "I also have the same problem as you, the submit file is good but it also failed in less than 30 minutes of execution."
    },
    {
      "id": 3005039,
      "postDate": "2024-10-02T13:52:12.040Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3007100,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-10-05T00:40:50.960000",
      "content": "<p><a href=\"https://www.kaggle.com/yoonjaelee\" target=\"_blank\">@yoonjaelee</a> and <a href=\"https://www.kaggle.com/karlbina\" target=\"_blank\">@karlbina</a> Here's what the statement at the bottom of the submission page says with some extra words that I've added -- I hope this makes the concept of a \"hidden test set\" with its own \"submission.csv\" clearer for you.</p>\n<p>\"In this competition, we will <strong>privately</strong> re-run your selected Notebook Version\" -- Privately means your notebook will be copied to some directory you cannot access and it will be run there.</p>\n<p>\"with a <strong>hidden test set</strong> substituted into the competition dataset.\" -- This means that in the private location, there will be the same \"train.csv\" file (and the same train parquet files), and there will also be a test.csv and test parquet files. <strong>But, this test.csv will be different, and longer,</strong> than the little test.csv that is in your working directory. Although the parquet directory for test is the same path as before, now it will have different parquet files, ones that go with (some of) the new test.csv id values. </p>\n<p>\"We then extract your chosen <strong>Output File</strong> from the re-run and use that to determine your score.\" -- When your notebook is run in private: i) it will read in the same training data, build features, and fit a model; ii) it will read in the new, hidden \"test.csv\" and the related parquet files, build the test features, and apply the trained-model to the test features; iii) the test predictions from the model are then written into a <strong>new submission.csv</strong> which has one row for each of the ids in the new, hidden test.csv. This created-in-private submission.csv is then scored based on the unknow-to-us correct answers for the hidden test.csv.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3013659,
      "author_name": "Alan C52",
      "author_url": "",
      "post_date": "2024-10-10T11:48:03.763000",
      "content": "<p>I had this issue very early. For me, my issue was from using multi-class objective function instead of binary objective function. Really, the public set works with multi-class, but hidden set is binary.</p>\n<p>Other issues could be nan, inf, or negatives that arise. Or different features (i.e., OHE where hidden set has more uniques than train and test set).</p>\n<p>It's probably best to add some filters and conditionals in your program to handle exceptional cases, etc.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3007073,
      "author_name": "Karl BINA",
      "author_url": "",
      "post_date": "2024-10-04T22:32:47.983000",
      "content": "<p>I also have the same problem as you, the submit file is good but it also failed in less than 30 minutes of execution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3005039,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-02T13:52:12.040000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3007100": "@yoonjaelee and @karlbina Here's what the statement at the bottom of the submission page says with some extra words that I've added -- I hope this makes the concept of a \"hidden test set\" with its own \"submission.csv\" clearer for you.\n\n\"In this competition, we will **privately** re-run your selected Notebook Version\" -- Privately means your notebook will be copied to some directory you cannot access and it will be run there.\n\n\"with a **hidden test set** substituted into the competition dataset.\" -- This means that in the private location, there will be the same \"train.csv\" file (and the same train parquet files), and there will also be a test.csv and test parquet files. **But, this test.csv will be different, and longer,** than the little test.csv that is in your working directory. Although the parquet directory for test is the same path as before, now it will have different parquet files, ones that go with (some of) the new test.csv id values. \n\n\"We then extract your chosen **Output File** from the re-run and use that to determine your score.\" -- When your notebook is run in private: i) it will read in the same training data, build features, and fit a model; ii) it will read in the new, hidden \"test.csv\" and the related parquet files, build the test features, and apply the trained-model to the test features; iii) the test predictions from the model are then written into a **new submission.csv** which has one row for each of the ids in the new, hidden test.csv. This created-in-private submission.csv is then scored based on the unknow-to-us correct answers for the hidden test.csv.  \n",
    "3004615": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2F5bef936191eac533e0e6333362d8dccd%2Fsubmission_capture01.jpg?generation=1727840679437152&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2Fd74fabb548310ff360628531e5dab0a8%2Fsubmission_capture02.jpg?generation=1727840759667982&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F78753%2Fc5c8c5b61813d1578f3b124ba6f02095%2Fsubmission_capture03.jpg?generation=1727840924976492&alt=media)\n\nI'm a newbie in ML competition. My submission.csv is looking good and has a right format I think. But after submitting my notebook, it took 2 hours for running and then failed. What am I missing??",
    "3013659": "I had this issue very early. For me, my issue was from using multi-class objective function instead of binary objective function. Really, the public set works with multi-class, but hidden set is binary.\n\nOther issues could be nan, inf, or negatives that arise. Or different features (i.e., OHE where hidden set has more uniques than train and test set).\n\nIt's probably best to add some filters and conditionals in your program to handle exceptional cases, etc.",
    "3007073": "I also have the same problem as you, the submit file is good but it also failed in less than 30 minutes of execution.",
    "3005039": ""
  }
}