{
  "id": 485836,
  "title": "Submissions are not working?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/485836",
  "author_name": "",
  "post_date": "2024-03-22T10:51:58.188627700Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I try to submit my work, and notebook is working fine, but it inevitably fails with Submission Scoring Error, despite submission file meets all the requirements (10 rows, case_id int, score float between 0 and 1). More than that, I tried to submit sample submission, output by starter notebook, which was scored fine previously, and it also fails. I see people face the same problem</p>\n<p><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635</a><br>\n(first comment here) <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262</a></p>\n<p>but their concerns remain unanswered. <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Can you please give any comments on that issue?</p>",
  "messages": [
    {
      "id": "2710544",
      "postDate": "03/22/2024 10:51:58",
      "content": "<p>I try to submit my work, and notebook is working fine, but it inevitably fails with Submission Scoring Error, despite submission file meets all the requirements (10 rows, case_id int, score float between 0 and 1). More than that, I tried to submit sample submission, output by starter notebook, which was scored fine previously, and it also fails. I see people face the same problem</p>\n<p><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635</a><br>\n(first comment here) <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262</a></p>\n<p>but their concerns remain unanswered. <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Can you please give any comments on that issue?</p>",
      "rawMarkdown": "I try to submit my work, and notebook is working fine, but it inevitably fails with Submission Scoring Error, despite submission file meets all the requirements (10 rows, case_id int, score float between 0 and 1). More than that, I tried to submit sample submission, output by starter notebook, which was scored fine previously, and it also fails. I see people face the same problem\n\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635\n(first comment here) https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262\n\nbut their concerns remain unanswered. @jetakow Can you please give any comments on that issue?",
      "votes": null
    },
    {
      "id": "2711486",
      "postDate": "03/22/2024 21:55:05",
      "content": "<p>The test files and sample_submission.csv that you see are dummy files. When you submit your notebook, these files are replaced with the actual test data, and your notebook is run against them. The actual test files are larger (the host mentioned that the whole test set is about 90% of the size of the training set, and the data used to evaluate the public leaderboard score constitutes about 30% of the entire test set). You don't have access to them because they are intended to imitate future, unseen data. </p>\n<p>This means that even if your code works fine on the dummy test files, it may encounter issues during the actual evaluation. Such issues could stem from format discrepancies, duplicated case IDs, among other reasons. </p>\n<p>To debug these problems, I'd suggest creating a test environment where you train your model on, say, 10% of the train case IDs and run your submission script on the other 90% of train case IDs. This approach will help you identify any errors without exhausting your submission quota.</p>",
      "rawMarkdown": "The test files and sample_submission.csv that you see are dummy files. When you submit your notebook, these files are replaced with the actual test data, and your notebook is run against them. The actual test files are larger (the host mentioned that the whole test set is about 90% of the size of the training set, and the data used to evaluate the public leaderboard score constitutes about 30% of the entire test set). You don't have access to them because they are intended to imitate future, unseen data. \n\nThis means that even if your code works fine on the dummy test files, it may encounter issues during the actual evaluation. Such issues could stem from format discrepancies, duplicated case IDs, among other reasons. \n\nTo debug these problems, I'd suggest creating a test environment where you train your model on, say, 10% of the train case IDs and run your submission script on the other 90% of train case IDs. This approach will help you identify any errors without exhausting your submission quota.",
      "votes": null
    },
    {
      "id": "2712094",
      "postDate": "03/23/2024 09:46:08",
      "content": "<p>Same problem here, no solution yet. </p>\n<ul>\n<li>My submission code runs green. </li>\n<li>Then the  Scoring Fails.</li>\n<li>It already worked in the past</li>\n</ul>\n<p>I already tried to:</p>\n<ul>\n<li>substitute null in score with 0</li>\n<li>dedup on case_id</li>\n<li>rounded score to 4 positions</li>\n<li>use outer joins in case a case_id is not included in base file, but checked during eval</li>\n</ul>\n<p>The sample of the submission file i created looks good.<br>\nIts called submission.csv and has these top 3 lines</p>\n<p>case_id,score<br>\n57633,0.0312<br>\n57634,0.1542<br>\n57630,0.054</p>",
      "rawMarkdown": "Same problem here, no solution yet. \n\n- My submission code runs green. \n- Then the  Scoring Fails.\n- It already worked in the past\n\nI already tried to:\n- substitute null in score with 0\n- dedup on case_id\n- rounded score to 4 positions\n- use outer joins in case a case_id is not included in base file, but checked during eval\n\nThe sample of the submission file i created looks good.\nIts called submission.csv and has these top 3 lines\n\ncase_id,score\n57633,0.0312\n57634,0.1542\n57630,0.054",
      "votes": null
    },
    {
      "id": "2712100",
      "postDate": "03/23/2024 09:50:27",
      "content": "<p>The problem the author describes is after all the code ran.<br>\nIts a problem exclusively related to the submission process.</p>\n<ul>\n<li>Either the submission file has an error (format, name, location, values, etc.)</li>\n<li>Or the submission code executed in the blackbox has a problem</li>\n</ul>\n<p>I already tried some format cleaning for potential errors in the submission file, see my other comment.</p>",
      "rawMarkdown": "The problem the author describes is after all the code ran.\nIts a problem exclusively related to the submission process.\n- Either the submission file has an error (format, name, location, values, etc.)\n- Or the submission code executed in the blackbox has a problem\n\nI already tried some format cleaning for potential errors in the submission file, see my other comment.",
      "votes": null
    },
    {
      "id": "2712488",
      "postDate": "03/23/2024 15:48:26",
      "content": "<p>I am also facing similar issue. I tried belows, but the issue was not resolved.</p>\n<ul>\n<li>Even after reverting the version to the successfully submitted notebook and running it, the test data becomes 10 rows.</li>\n<li>The same thing happens when you fork and run other public notebooks.</li>\n</ul>\n<p>I'm guessing something is wrong with the system.</p>",
      "rawMarkdown": "I am also facing similar issue. I tried belows, but the issue was not resolved.\n- Even after reverting the version to the successfully submitted notebook and running it, the test data becomes 10 rows.\n- The same thing happens when you fork and run other public notebooks.\n\nI'm guessing something is wrong with the system.",
      "votes": null
    },
    {
      "id": "2712524",
      "postDate": "03/23/2024 16:13:24",
      "content": "<p>Actually, that information resolved the issue for me. I was using prepared datasets, so evaluation of hidden test set was broken since it basically didn't retrieve the new data. Thanks!</p>",
      "rawMarkdown": "Actually, that information resolved the issue for me. I was using prepared datasets, so evaluation of hidden test set was broken since it basically didn't retrieve the new data. Thanks!",
      "votes": null
    },
    {
      "id": "2712539",
      "postDate": "03/23/2024 16:27:39",
      "content": "<p>Ok, but the original problem is still there isn't it?<br>\nJust want to make sure that people who have the same problem, do not get confused by this post.</p>\n<p>It's great information and valuable advice, but will quite likely not help you solve the problem of this post.</p>",
      "rawMarkdown": "Ok, but the original problem is still there isn't it?\nJust want to make sure that people who have the same problem, do not get confused by this post.\n\nIt's great information and valuable advice, but will quite likely not help you solve the problem of this post.",
      "votes": null
    },
    {
      "id": "2726368",
      "postDate": "04/01/2024 06:38:41",
      "content": "<p>Finally got it. Thanks <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> <br>\nIt’s my first competition and I did not get that the notebook is automatically executed against the dummies first during submission </p>",
      "rawMarkdown": "Finally got it. Thanks @eivolkova \nIt’s my first competition and I did not get that the notebook is automatically executed against the dummies first during submission",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2711486,
      "author_name": "eivolkova",
      "author_url": "",
      "post_date": "03/22/2024 21:55:05",
      "content": "<p>The test files and sample_submission.csv that you see are dummy files. When you submit your notebook, these files are replaced with the actual test data, and your notebook is run against them. The actual test files are larger (the host mentioned that the whole test set is about 90% of the size of the training set, and the data used to evaluate the public leaderboard score constitutes about 30% of the entire test set). You don't have access to them because they are intended to imitate future, unseen data. </p>\n<p>This means that even if your code works fine on the dummy test files, it may encounter issues during the actual evaluation. Such issues could stem from format discrepancies, duplicated case IDs, among other reasons. </p>\n<p>To debug these problems, I'd suggest creating a test environment where you train your model on, say, 10% of the train case IDs and run your submission script on the other 90% of train case IDs. This approach will help you identify any errors without exhausting your submission quota.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2712100,
          "author_name": "aahhammer",
          "author_url": "",
          "post_date": "03/23/2024 09:50:27",
          "content": "<p>The problem the author describes is after all the code ran.<br>\nIts a problem exclusively related to the submission process.</p>\n<ul>\n<li>Either the submission file has an error (format, name, location, values, etc.)</li>\n<li>Or the submission code executed in the blackbox has a problem</li>\n</ul>\n<p>I already tried some format cleaning for potential errors in the submission file, see my other comment.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2712524,
          "author_name": "nerkan",
          "author_url": "",
          "post_date": "03/23/2024 16:13:24",
          "content": "<p>Actually, that information resolved the issue for me. I was using prepared datasets, so evaluation of hidden test set was broken since it basically didn't retrieve the new data. Thanks!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2712539,
              "author_name": "aahhammer",
              "author_url": "",
              "post_date": "03/23/2024 16:27:39",
              "content": "<p>Ok, but the original problem is still there isn't it?<br>\nJust want to make sure that people who have the same problem, do not get confused by this post.</p>\n<p>It's great information and valuable advice, but will quite likely not help you solve the problem of this post.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2726368,
          "author_name": "aahhammer",
          "author_url": "",
          "post_date": "04/01/2024 06:38:41",
          "content": "<p>Finally got it. Thanks <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> <br>\nIt’s my first competition and I did not get that the notebook is automatically executed against the dummies first during submission </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2712094,
      "author_name": "aahhammer",
      "author_url": "",
      "post_date": "03/23/2024 09:46:08",
      "content": "<p>Same problem here, no solution yet. </p>\n<ul>\n<li>My submission code runs green. </li>\n<li>Then the  Scoring Fails.</li>\n<li>It already worked in the past</li>\n</ul>\n<p>I already tried to:</p>\n<ul>\n<li>substitute null in score with 0</li>\n<li>dedup on case_id</li>\n<li>rounded score to 4 positions</li>\n<li>use outer joins in case a case_id is not included in base file, but checked during eval</li>\n</ul>\n<p>The sample of the submission file i created looks good.<br>\nIts called submission.csv and has these top 3 lines</p>\n<p>case_id,score<br>\n57633,0.0312<br>\n57634,0.1542<br>\n57630,0.054</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2712488,
      "author_name": "shisa07",
      "author_url": "",
      "post_date": "03/23/2024 15:48:26",
      "content": "<p>I am also facing similar issue. I tried belows, but the issue was not resolved.</p>\n<ul>\n<li>Even after reverting the version to the successfully submitted notebook and running it, the test data becomes 10 rows.</li>\n<li>The same thing happens when you fork and run other public notebooks.</li>\n</ul>\n<p>I'm guessing something is wrong with the system.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2710544": "I try to submit my work, and notebook is working fine, but it inevitably fails with Submission Scoring Error, despite submission file meets all the requirements (10 rows, case_id int, score float between 0 and 1). More than that, I tried to submit sample submission, output by starter notebook, which was scored fine previously, and it also fails. I see people face the same problem\n\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483635\n(first comment here) https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/483262\n\nbut their concerns remain unanswered. @jetakow Can you please give any comments on that issue?",
    "2711486": "The test files and sample_submission.csv that you see are dummy files. When you submit your notebook, these files are replaced with the actual test data, and your notebook is run against them. The actual test files are larger (the host mentioned that the whole test set is about 90% of the size of the training set, and the data used to evaluate the public leaderboard score constitutes about 30% of the entire test set). You don't have access to them because they are intended to imitate future, unseen data. \n\nThis means that even if your code works fine on the dummy test files, it may encounter issues during the actual evaluation. Such issues could stem from format discrepancies, duplicated case IDs, among other reasons. \n\nTo debug these problems, I'd suggest creating a test environment where you train your model on, say, 10% of the train case IDs and run your submission script on the other 90% of train case IDs. This approach will help you identify any errors without exhausting your submission quota.",
    "2712094": "Same problem here, no solution yet. \n\n- My submission code runs green. \n- Then the  Scoring Fails.\n- It already worked in the past\n\nI already tried to:\n- substitute null in score with 0\n- dedup on case_id\n- rounded score to 4 positions\n- use outer joins in case a case_id is not included in base file, but checked during eval\n\nThe sample of the submission file i created looks good.\nIts called submission.csv and has these top 3 lines\n\ncase_id,score\n57633,0.0312\n57634,0.1542\n57630,0.054",
    "2712100": "The problem the author describes is after all the code ran.\nIts a problem exclusively related to the submission process.\n- Either the submission file has an error (format, name, location, values, etc.)\n- Or the submission code executed in the blackbox has a problem\n\nI already tried some format cleaning for potential errors in the submission file, see my other comment.",
    "2712488": "I am also facing similar issue. I tried belows, but the issue was not resolved.\n- Even after reverting the version to the successfully submitted notebook and running it, the test data becomes 10 rows.\n- The same thing happens when you fork and run other public notebooks.\n\nI'm guessing something is wrong with the system.",
    "2712524": "Actually, that information resolved the issue for me. I was using prepared datasets, so evaluation of hidden test set was broken since it basically didn't retrieve the new data. Thanks!",
    "2712539": "Ok, but the original problem is still there isn't it?\nJust want to make sure that people who have the same problem, do not get confused by this post.\n\nIt's great information and valuable advice, but will quite likely not help you solve the problem of this post.",
    "2726368": "Finally got it. Thanks @eivolkova \nIt’s my first competition and I did not get that the notebook is automatically executed against the dummies first during submission"
  },
  "source": "meta"
}