{
  "id": 198317,
  "title": "Is it possible to make a submission from local CSV?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/198317",
  "author_name": "",
  "post_date": "2020-11-20T17:46:13.001168200Z",
  "votes": 13,
  "comment_count": 28,
  "views": 0,
  "content": "<p>So, if I create a CSV locally for the test set file and upload it as a notebook dataset, and then submit it what problems I'm going to face? I have a much better inference speed locally and no memory bottleneck. I assume this will not work for final submission, but can I use it as a quick LB check?<br>\nThanks!</p>",
  "messages": [
    {
      "id": "1085142",
      "postDate": "11/20/2020 17:46:13",
      "content": "<p>So, if I create a CSV locally for the test set file and upload it as a notebook dataset, and then submit it what problems I'm going to face? I have a much better inference speed locally and no memory bottleneck. I assume this will not work for final submission, but can I use it as a quick LB check?<br>\nThanks!</p>",
      "rawMarkdown": "So, if I create a CSV locally for the test set file and upload it as a notebook dataset, and then submit it what problems I'm going to face? I have a much better inference speed locally and no memory bottleneck. I assume this will not work for final submission, but can I use it as a quick LB check?\nThanks!",
      "votes": null
    },
    {
      "id": "1085194",
      "postDate": "11/20/2020 18:25:51",
      "content": "<p>Yes. This is a quick way to check your public LB score. But as you note, this will not work for the final leaderboard.</p>\n<p>To do it, you can't just copy the submission.csv though. You'll get an error unless your submission.csv contains the private test ids too. Therefor you should copy the predictions from your submission.csv into the sample_submission.csv with something like this:</p>\n<pre><code>submission = pd.read_csv('submission.csv', index_col='id')\nsample_sub = pd.read_csv('sample_submission.csv', index_col='id')\npub_ids = submission.index.values\npredictions = submission.values\nsample_sub.loc[pub_ids] = predictions\nsample_sub.to_csv('submission.csv')\n</code></pre>",
      "rawMarkdown": "Yes. This is a quick way to check your public LB score. But as you note, this will not work for the final leaderboard.\n\nTo do it, you can't just copy the submission.csv though. You'll get an error unless your submission.csv contains the private test ids too. Therefor you should copy the predictions from your submission.csv into the sample_submission.csv with something like this:\n```\nsubmission = pd.read_csv('submission.csv', index_col='id')\nsample_sub = pd.read_csv('sample_submission.csv', index_col='id')\npub_ids = submission.index.values\npredictions = submission.values\nsample_sub.loc[pub_ids] = predictions\nsample_sub.to_csv('submission.csv')\n```",
      "votes": null
    },
    {
      "id": "1085228",
      "postDate": "11/20/2020 19:04:25",
      "content": "<p>Nice! Thanks!</p>",
      "rawMarkdown": "Nice! Thanks!",
      "votes": null
    },
    {
      "id": "1088454",
      "postDate": "11/23/2020 16:14:41",
      "content": "<p>Have anyone tried it? Submission works just fine but the score is extremely low. Visually the masks are more or less ok. Not perfect, but not garbage either. So I wonder if kernel test files are different from test files given for downloading.</p>",
      "rawMarkdown": "Have anyone tried it? Submission works just fine but the score is extremely low. Visually the masks are more or less ok. Not perfect, but not garbage either. So I wonder if kernel test files are different from test files given for downloading.",
      "votes": null
    },
    {
      "id": "1088523",
      "postDate": "11/23/2020 17:32:43",
      "content": "<p>I have tried to check my public LB score by this way and got 0,248. In my case, the masks contain lots of noise, so the result is expected. Hence, it doesn't seem kernel test files are different from test files given for downloading</p>",
      "rawMarkdown": "I have tried to check my public LB score by this way and got 0,248. In my case, the masks contain lots of noise, so the result is expected. Hence, it doesn't seem kernel test files are different from test files given for downloading",
      "votes": null
    },
    {
      "id": "1088580",
      "postDate": "11/23/2020 18:57:09",
      "content": "<p>The files should be the same. I haven't submitted yet though.</p>\n<p>I'm guessing there's a bug somewhere between you visualizing your masks and RLE encoding them. Perhaps the orientation of your labels are incorrect (i.e. vertically flipped).</p>",
      "rawMarkdown": "The files should be the same. I haven't submitted yet though.\n\nI'm guessing there's a bug somewhere between you visualizing your masks and RLE encoding them. Perhaps the orientation of your labels are incorrect (i.e. vertically flipped).",
      "votes": null
    },
    {
      "id": "1088647",
      "postDate": "11/23/2020 20:27:06",
      "content": "<p>God damn this rle encoding. Why it needs to be transposed! I've spent two days trying to debug this issue!</p>",
      "rawMarkdown": "God damn this rle encoding. Why it needs to be transposed! I've spent two days trying to debug this issue!",
      "votes": null
    },
    {
      "id": "1094865",
      "postDate": "11/29/2020 03:11:28",
      "content": "<p>Hi, I just encountered the same problem. The score is extremely low. The generated mask looks ok, and it works when I did leave-one-out validation on training set. <br>\nBut I've already make a transpose on the generated mask before RLE encoding. <br>\nCould you please let me know how to deal with this issue? Many thanks!</p>",
      "rawMarkdown": "Hi, I just encountered the same problem. The score is extremely low. The generated mask looks ok, and it works when I did leave-one-out validation on training set. \nBut I've already make a transpose on the generated mask before RLE encoding. \nCould you please let me know how to deal with this issue? Many thanks!",
      "votes": null
    },
    {
      "id": "1095179",
      "postDate": "11/29/2020 10:29:37",
      "content": "<p>Hi, I updated my submission.csv from my local machine and tried to submit it. But I was told \"Submission Scoring Error\". I've checked the format of my .csv file and it seems OK. Any idea to resolve this problem? I am a newbie on Kaggle paltform. Thanks in advance.</p>",
      "rawMarkdown": "Hi, I updated my submission.csv from my local machine and tried to submit it. But I was told \"Submission Scoring Error\". I've checked the format of my .csv file and it seems OK. Any idea to resolve this problem? I am a newbie on Kaggle paltform. Thanks in advance.",
      "votes": null
    },
    {
      "id": "1095196",
      "postDate": "11/29/2020 10:50:39",
      "content": "<p>I've encountered the same issue. In my case, the error is due to improper assignment of index column. Pandas will add an index column automatically.</p>",
      "rawMarkdown": "I've encountered the same issue. In my case, the error is due to improper assignment of index column. Pandas will add an index column automatically.",
      "votes": null
    },
    {
      "id": "1095374",
      "postDate": "11/29/2020 14:37:16",
      "content": "<p>This could be due to RLE encoding of an image that differs slightly in dimensions from the original image. If you encode an image that is 999x999 and decode it as 1000x1000 there will be a shift of one pixel per line. But visually you will not see this. RLE doesn't save image dimensions so this can be hard to debuig.</p>",
      "rawMarkdown": "This could be due to RLE encoding of an image that differs slightly in dimensions from the original image. If you encode an image that is 999x999 and decode it as 1000x1000 there will be a shift of one pixel per line. But visually you will not see this. RLE doesn't save image dimensions so this can be hard to debuig.",
      "votes": null
    },
    {
      "id": "1095888",
      "postDate": "11/30/2020 03:20:05",
      "content": "<p>Thank you Dennis! I've figured out this issue last night, which is due to the format of submission file. Thanks for your answer!</p>",
      "rawMarkdown": "Thank you Dennis! I've figured out this issue last night, which is due to the format of submission file. Thanks for your answer!",
      "votes": null
    },
    {
      "id": "1095913",
      "postDate": "11/30/2020 04:04:26",
      "content": "<p>When you submit, your notebook is rerun with the full test dataset (public + private). Therefor, if you simply copy your submission.csv, it won't have the expected number of rows and will fail. You should do something like <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198317#1085194\" target=\"_blank\">what I posted below</a> to load your predictions into the sample_submission csv.</p>",
      "rawMarkdown": "When you submit, your notebook is rerun with the full test dataset (public + private). Therefor, if you simply copy your submission.csv, it won't have the expected number of rows and will fail. You should do something like [what I posted below](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198317#1085194) to load your predictions into the sample_submission csv.",
      "votes": null
    },
    {
      "id": "1095930",
      "postDate": "11/30/2020 04:34:22",
      "content": "<p>Does the scoring for LB use public+ private test dataset? If yes, even we convert our \"submission.csv\" to \"sample_submission.csv\", there are still 5 rows.<br>\nSo how is the private test set included in the evaluation as we don't know the image_id for private test dataset.<br>\nThank you in advance!</p>",
      "rawMarkdown": "Does the scoring for LB use public+ private test dataset? If yes, even we convert our \"submission.csv\" to \"sample_submission.csv\", there are still 5 rows.\nSo how is the private test set included in the evaluation as we don't know the image_id for private test dataset.\nThank you in advance!",
      "votes": null
    },
    {
      "id": "1095942",
      "postDate": "11/30/2020 04:47:45",
      "content": "<p>When you commit your notebook, you'll only ever see the public test set examples. The private test set files are only visible to your notebook during submission. When you submit your notebook, the dataset is replaced to include the private test set and the notebook is rerun.</p>\n<p>It will score your predictions for both the private and public test sets, but only shows your public score until the end of the competition. For now, it's okay to leave the predictions for the private set blank. But for your final submission, you're gonna want to perform your prediction in the notebook.</p>\n<p>Hope all that makes sense, I know it is confusing :)</p>",
      "rawMarkdown": "When you commit your notebook, you'll only ever see the public test set examples. The private test set files are only visible to your notebook during submission. When you submit your notebook, the dataset is replaced to include the private test set and the notebook is rerun.\n\nIt will score your predictions for both the private and public test sets, but only shows your public score until the end of the competition. For now, it's okay to leave the predictions for the private set blank. But for your final submission, you're gonna want to perform your prediction in the notebook.\n\nHope all that makes sense, I know it is confusing :)",
      "votes": null
    },
    {
      "id": "1096113",
      "postDate": "11/30/2020 08:17:16",
      "content": "<p>You have already made it clearer! The only question left for me is -- How can I build a notebook that fits both conditions? Currently when doing test, I read the \"sample_submission.csv\" provided in the original dataset, use the \"id\" column to get which samples I need for test. But obviously this method cannot fit the situation with private test dataset. </p>\n<p>BTW is private test dataset used for current LB score? Or the score is just calculated by public dataset.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "You have already made it clearer! The only question left for me is -- How can I build a notebook that fits both conditions? Currently when doing test, I read the \"sample_submission.csv\" provided in the original dataset, use the \"id\" column to get which samples I need for test. But obviously this method cannot fit the situation with private test dataset. \n\nBTW is private test dataset used for current LB score? Or the score is just calculated by public dataset.\n\nThanks!",
      "votes": null
    },
    {
      "id": "1096459",
      "postDate": "11/30/2020 14:12:11",
      "content": "<p><a href=\"https://www.kaggle.com/lijiaqi96\" target=\"_blank\">@lijiaqi96</a> When you submit your notebook the \"sample_submission.csv\" gets replaced with a file that contains all image <strong>id</strong>'s of the hidden test set, so you don't need to make special changes for the private test run.</p>\n<p>The public LB score you see now is using the public test dataset. The private test score is hidden till the end of the competition, which is why a high public LB score doesn't necessarily mean your final position will be the same.</p>",
      "rawMarkdown": "lijiaqi96 When you submit your notebook the \"sample_submission.csv\" gets replaced with a file that contains all image **id**'s of the hidden test set, so you don't need to make special changes for the private test run.\n\nThe public LB score you see now is using the public test dataset. The private test score is hidden till the end of the competition, which is why a high public LB score doesn't necessarily mean your final position will be the same.",
      "votes": null
    },
    {
      "id": "1096482",
      "postDate": "11/30/2020 14:25:49",
      "content": "<p>Hello, have you solved this problem? Is it possible to make a submission from local CSV?</p>",
      "rawMarkdown": "Hello, have you solved this problem? Is it possible to make a submission from local CSV?",
      "votes": null
    },
    {
      "id": "1096491",
      "postDate": "11/30/2020 14:32:01",
      "content": "<p><a href=\"https://www.kaggle.com/tikboa\" target=\"_blank\">@tikboa</a> This is a code competition which means you are not allowed to submit you local CSV through a notebook to get a score. Even if you get a score this way for the public LB, your private test score will get an error since you have not made predictions for the hidden private test dataset.</p>\n<p>You need to use your locally trained models to make inference in real time inside a notebook which you submit to the competition.</p>",
      "rawMarkdown": "tikboa This is a code competition which means you are not allowed to submit you local CSV through a notebook to get a score. Even if you get a score this way for the public LB, your private test score will get an error since you have not made predictions for the hidden private test dataset.\n\nYou need to use your locally trained models to make inference in real time inside a notebook which you submit to the competition.",
      "votes": null
    },
    {
      "id": "1096565",
      "postDate": "11/30/2020 15:54:16",
      "content": "<p>So I just keep getting test image ids from \"sample_submission.csv\" from \"/kaggle/input\". The same code will work for the final submission, am I right?</p>",
      "rawMarkdown": "So I just keep getting test image ids from \"sample_submission.csv\" from \"/kaggle/input\". The same code will work for the final submission, am I right?",
      "votes": null
    },
    {
      "id": "1096571",
      "postDate": "11/30/2020 15:56:54",
      "content": "<p>Yes, that is correct.</p>",
      "rawMarkdown": "Yes, that is correct.",
      "votes": null
    },
    {
      "id": "1096943",
      "postDate": "11/30/2020 22:26:33",
      "content": "<p>I might be missing something here. Is it correct that the models may be trained on personal workstations (without limitations on the training time), but the inference time for the final model when run on the kaggle notebook must be within the time limits specified (assuming the trained model can be uploaded directly from local)? </p>",
      "rawMarkdown": "I might be missing something here. Is it correct that the models may be trained on personal workstations (without limitations on the training time), but the inference time for the final model when run on the kaggle notebook must be within the time limits specified (assuming the trained model can be uploaded directly from local)?",
      "votes": null
    },
    {
      "id": "1097318",
      "postDate": "12/01/2020 02:43:17",
      "content": "<p>You are correct. No limitations on training time or resources.</p>\n<p>Inference must be run within committed model with time limit of 9 hours. No TPU. No Internet. GPU allowed.</p>",
      "rawMarkdown": "You are correct. No limitations on training time or resources.\n\nInference must be run within committed model with time limit of 9 hours. No TPU. No Internet. GPU allowed.",
      "votes": null
    },
    {
      "id": "1098516",
      "postDate": "12/01/2020 17:22:30",
      "content": "<p>Thanks for the quick clarification! </p>",
      "rawMarkdown": "Thanks for the quick clarification!",
      "votes": null
    },
    {
      "id": "1249266",
      "postDate": "03/23/2021 08:00:56",
      "content": "<p>I had Submission Error, but after copying of submission.csv as proposed here, I got public score. <br>\nThank you, <a href=\"https://www.kaggle.com/matthewmasters\" target=\"_blank\">@matthewmasters</a>!</p>\n<p>The problem with my submission was that I've created .tfrec files in kaggle/working/ directory instead of kaggle/temp/.</p>",
      "rawMarkdown": "I had Submission Error, but after copying of submission.csv as proposed here, I got public score. \nThank you, @matthewmasters!\n\nThe problem with my submission was that I've created .tfrec files in kaggle/working/ directory instead of kaggle/temp/.",
      "votes": null
    },
    {
      "id": "1281964",
      "postDate": "04/23/2021 13:22:20",
      "content": "<p>I'm facing the same issue. Can you tell me if we need to save all output including submission.csv to temp folder?</p>",
      "rawMarkdown": "I'm facing the same issue. Can you tell me if we need to save all output including submission.csv to temp folder?",
      "votes": null
    },
    {
      "id": "1283669",
      "postDate": "04/25/2021 07:02:01",
      "content": "<p>I save submission.csv to kaggle/working and other output staff (.tfrec) to kaggle/temp/ and it works.</p>",
      "rawMarkdown": "I save submission.csv to kaggle/working and other output staff (.tfrec) to kaggle/temp/ and it works.",
      "votes": null
    },
    {
      "id": "1284314",
      "postDate": "04/25/2021 19:31:26",
      "content": "<p>I tried it but still not successful. Can I know if you used kaggle resources itself or did you used may be GCP somehow when submitting?</p>",
      "rawMarkdown": "I tried it but still not successful. Can I know if you used kaggle resources itself or did you used may be GCP somehow when submitting?",
      "votes": null
    },
    {
      "id": "1284335",
      "postDate": "04/25/2021 19:47:12",
      "content": "<p>I use only Kaggle resources for submission:</p>\n<pre>subm = {}\n\ndf_sample = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv')\n\nwith open(os.path.join(temp_path, 'test/test_shape.yaml'), 'r') as file:\n    shapes = yaml.safe_load(file)\n\nfor idx, row in tqdm(df_sample.iterrows(), total=len(df_sample)):\n    filename_stem = row['id']\n    filename = os.path.join(test_dir_name, filename_stem + '.tfrec')\n    print('{} Predicting {}'.format(idx+1, filename_stem))\n    shape = shapes[filename_stem]\n    preds = np.zeros(shape, dtype=np.uint8)\n\n    for imgs, X1, Y1 in load_dataset(filenames=filename, batch_size=16):\n        pred_avg_segm = None\n\n        # Predict mask itself\n        for model_segm in fold_models_segm:\n            pred_segm = model_segm.predict(imgs) / len(fold_models_segm)\n\n            if pred_avg_segm is None:\n                pred_avg_segm = pred_segm\n            else:\n                pred_avg_segm += pred_segm\n\n        pred_avg_segm = tf.cast((tf.image.resize(pred_avg_segm, (WINDOW,WINDOW)) &gt; THRESHOLD_SEGM), tf.bool).numpy().squeeze()\n\n        for i in range(X1.shape[0]):\n            preds[X1[i]:(X1[i]+WINDOW),Y1[i]:(Y1[i]+WINDOW)] += pred_avg_segm[i]\n\n    preds = (preds &gt; THRESHOLD_SEGM).astype(np.uint8)\n    print('Shape before: {}, after: {}'.format(shape, preds.shape))\n\n    subm[idx] = {'id':filename_stem, 'predicted': rle_encode_less_memory(preds)}\n\n    del preds\n    gc.collect()\n\ndf_submit = pd.DataFrame.from_dict(subm, orient='index')\ndf_submit = df_submit.set_index('id')\ndf_submit.to_csv('submission.csv')\n</pre>",
      "rawMarkdown": "I use only Kaggle resources for submission:\n\n<pre>\nsubm = {}\n\ndf_sample = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv')\n\nwith open(os.path.join(temp_path, 'test/test_shape.yaml'), 'r') as file:\n    shapes = yaml.safe_load(file)\n\nfor idx, row in tqdm(df_sample.iterrows(), total=len(df_sample)):\n    filename_stem = row['id']\n    filename = os.path.join(test_dir_name, filename_stem + '.tfrec')\n    print('{} Predicting {}'.format(idx+1, filename_stem))\n    shape = shapes[filename_stem]\n    preds = np.zeros(shape, dtype=np.uint8)\n\n    for imgs, X1, Y1 in load_dataset(filenames=filename, batch_size=16):\n        pred_avg_segm = None\n        \n        # Predict mask itself\n        for model_segm in fold_models_segm:\n            pred_segm = model_segm.predict(imgs) / len(fold_models_segm)\n\n            if pred_avg_segm is None:\n                pred_avg_segm = pred_segm\n            else:\n                pred_avg_segm += pred_segm\n\n        pred_avg_segm = tf.cast((tf.image.resize(pred_avg_segm, (WINDOW,WINDOW)) > THRESHOLD_SEGM), tf.bool).numpy().squeeze()\n        \n        for i in range(X1.shape[0]):\n            preds[X1[i]:(X1[i]+WINDOW),Y1[i]:(Y1[i]+WINDOW)] += pred_avg_segm[i]\n\n    preds = (preds > THRESHOLD_SEGM).astype(np.uint8)\n    print('Shape before: {}, after: {}'.format(shape, preds.shape))\n    \n    subm[idx] = {'id':filename_stem, 'predicted': rle_encode_less_memory(preds)}\n\n    del preds\n    gc.collect()\n\ndf_submit = pd.DataFrame.from_dict(subm, orient='index')\ndf_submit = df_submit.set_index('id')\ndf_submit.to_csv('submission.csv')\n</pre>",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1085194,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "11/20/2020 18:25:51",
      "content": "<p>Yes. This is a quick way to check your public LB score. But as you note, this will not work for the final leaderboard.</p>\n<p>To do it, you can't just copy the submission.csv though. You'll get an error unless your submission.csv contains the private test ids too. Therefor you should copy the predictions from your submission.csv into the sample_submission.csv with something like this:</p>\n<pre><code>submission = pd.read_csv('submission.csv', index_col='id')\nsample_sub = pd.read_csv('sample_submission.csv', index_col='id')\npub_ids = submission.index.values\npredictions = submission.values\nsample_sub.loc[pub_ids] = predictions\nsample_sub.to_csv('submission.csv')\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1085228,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "11/20/2020 19:04:25",
          "content": "<p>Nice! Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088454,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "11/23/2020 16:14:41",
          "content": "<p>Have anyone tried it? Submission works just fine but the score is extremely low. Visually the masks are more or less ok. Not perfect, but not garbage either. So I wonder if kernel test files are different from test files given for downloading.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088523,
          "author_name": "oleksandrk",
          "author_url": "",
          "post_date": "11/23/2020 17:32:43",
          "content": "<p>I have tried to check my public LB score by this way and got 0,248. In my case, the masks contain lots of noise, so the result is expected. Hence, it doesn't seem kernel test files are different from test files given for downloading</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088580,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "11/23/2020 18:57:09",
          "content": "<p>The files should be the same. I haven't submitted yet though.</p>\n<p>I'm guessing there's a bug somewhere between you visualizing your masks and RLE encoding them. Perhaps the orientation of your labels are incorrect (i.e. vertically flipped).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088647,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "11/23/2020 20:27:06",
          "content": "<p>God damn this rle encoding. Why it needs to be transposed! I've spent two days trying to debug this issue!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096943,
          "author_name": "bvineeth007",
          "author_url": "",
          "post_date": "11/30/2020 22:26:33",
          "content": "<p>I might be missing something here. Is it correct that the models may be trained on personal workstations (without limitations on the training time), but the inference time for the final model when run on the kaggle notebook must be within the time limits specified (assuming the trained model can be uploaded directly from local)? </p>",
          "votes": null,
          "replies": [
            {
              "id": 1097318,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "12/01/2020 02:43:17",
              "content": "<p>You are correct. No limitations on training time or resources.</p>\n<p>Inference must be run within committed model with time limit of 9 hours. No TPU. No Internet. GPU allowed.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1098516,
              "author_name": "bvineeth007",
              "author_url": "",
              "post_date": "12/01/2020 17:22:30",
              "content": "<p>Thanks for the quick clarification! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1094865,
      "author_name": "lijiaqi96",
      "author_url": "",
      "post_date": "11/29/2020 03:11:28",
      "content": "<p>Hi, I just encountered the same problem. The score is extremely low. The generated mask looks ok, and it works when I did leave-one-out validation on training set. <br>\nBut I've already make a transpose on the generated mask before RLE encoding. <br>\nCould you please let me know how to deal with this issue? Many thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1095374,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "11/29/2020 14:37:16",
          "content": "<p>This could be due to RLE encoding of an image that differs slightly in dimensions from the original image. If you encode an image that is 999x999 and decode it as 1000x1000 there will be a shift of one pixel per line. But visually you will not see this. RLE doesn't save image dimensions so this can be hard to debuig.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095888,
          "author_name": "lijiaqi96",
          "author_url": "",
          "post_date": "11/30/2020 03:20:05",
          "content": "<p>Thank you Dennis! I've figured out this issue last night, which is due to the format of submission file. Thanks for your answer!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1095179,
      "author_name": "liuzhian",
      "author_url": "",
      "post_date": "11/29/2020 10:29:37",
      "content": "<p>Hi, I updated my submission.csv from my local machine and tried to submit it. But I was told \"Submission Scoring Error\". I've checked the format of my .csv file and it seems OK. Any idea to resolve this problem? I am a newbie on Kaggle paltform. Thanks in advance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1095196,
          "author_name": "lijiaqi96",
          "author_url": "",
          "post_date": "11/29/2020 10:50:39",
          "content": "<p>I've encountered the same issue. In my case, the error is due to improper assignment of index column. Pandas will add an index column automatically.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095913,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "11/30/2020 04:04:26",
          "content": "<p>When you submit, your notebook is rerun with the full test dataset (public + private). Therefor, if you simply copy your submission.csv, it won't have the expected number of rows and will fail. You should do something like <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198317#1085194\" target=\"_blank\">what I posted below</a> to load your predictions into the sample_submission csv.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095930,
          "author_name": "lijiaqi96",
          "author_url": "",
          "post_date": "11/30/2020 04:34:22",
          "content": "<p>Does the scoring for LB use public+ private test dataset? If yes, even we convert our \"submission.csv\" to \"sample_submission.csv\", there are still 5 rows.<br>\nSo how is the private test set included in the evaluation as we don't know the image_id for private test dataset.<br>\nThank you in advance!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095942,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "11/30/2020 04:47:45",
          "content": "<p>When you commit your notebook, you'll only ever see the public test set examples. The private test set files are only visible to your notebook during submission. When you submit your notebook, the dataset is replaced to include the private test set and the notebook is rerun.</p>\n<p>It will score your predictions for both the private and public test sets, but only shows your public score until the end of the competition. For now, it's okay to leave the predictions for the private set blank. But for your final submission, you're gonna want to perform your prediction in the notebook.</p>\n<p>Hope all that makes sense, I know it is confusing :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096113,
          "author_name": "lijiaqi96",
          "author_url": "",
          "post_date": "11/30/2020 08:17:16",
          "content": "<p>You have already made it clearer! The only question left for me is -- How can I build a notebook that fits both conditions? Currently when doing test, I read the \"sample_submission.csv\" provided in the original dataset, use the \"id\" column to get which samples I need for test. But obviously this method cannot fit the situation with private test dataset. </p>\n<p>BTW is private test dataset used for current LB score? Or the score is just calculated by public dataset.</p>\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096459,
          "author_name": "yovinyahathugoda",
          "author_url": "",
          "post_date": "11/30/2020 14:12:11",
          "content": "<p><a href=\"https://www.kaggle.com/lijiaqi96\" target=\"_blank\">@lijiaqi96</a> When you submit your notebook the \"sample_submission.csv\" gets replaced with a file that contains all image <strong>id</strong>'s of the hidden test set, so you don't need to make special changes for the private test run.</p>\n<p>The public LB score you see now is using the public test dataset. The private test score is hidden till the end of the competition, which is why a high public LB score doesn't necessarily mean your final position will be the same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096565,
          "author_name": "lijiaqi96",
          "author_url": "",
          "post_date": "11/30/2020 15:54:16",
          "content": "<p>So I just keep getting test image ids from \"sample_submission.csv\" from \"/kaggle/input\". The same code will work for the final submission, am I right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096571,
          "author_name": "yovinyahathugoda",
          "author_url": "",
          "post_date": "11/30/2020 15:56:54",
          "content": "<p>Yes, that is correct.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1096482,
      "author_name": "tikboa",
      "author_url": "",
      "post_date": "11/30/2020 14:25:49",
      "content": "<p>Hello, have you solved this problem? Is it possible to make a submission from local CSV?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1096491,
          "author_name": "yovinyahathugoda",
          "author_url": "",
          "post_date": "11/30/2020 14:32:01",
          "content": "<p><a href=\"https://www.kaggle.com/tikboa\" target=\"_blank\">@tikboa</a> This is a code competition which means you are not allowed to submit you local CSV through a notebook to get a score. Even if you get a score this way for the public LB, your private test score will get an error since you have not made predictions for the hidden private test dataset.</p>\n<p>You need to use your locally trained models to make inference in real time inside a notebook which you submit to the competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1249266,
      "author_name": "vsharkeev",
      "author_url": "",
      "post_date": "03/23/2021 08:00:56",
      "content": "<p>I had Submission Error, but after copying of submission.csv as proposed here, I got public score. <br>\nThank you, <a href=\"https://www.kaggle.com/matthewmasters\" target=\"_blank\">@matthewmasters</a>!</p>\n<p>The problem with my submission was that I've created .tfrec files in kaggle/working/ directory instead of kaggle/temp/.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1281964,
          "author_name": "gssdatamanager",
          "author_url": "",
          "post_date": "04/23/2021 13:22:20",
          "content": "<p>I'm facing the same issue. Can you tell me if we need to save all output including submission.csv to temp folder?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283669,
          "author_name": "vsharkeev",
          "author_url": "",
          "post_date": "04/25/2021 07:02:01",
          "content": "<p>I save submission.csv to kaggle/working and other output staff (.tfrec) to kaggle/temp/ and it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1284314,
          "author_name": "gssdatamanager",
          "author_url": "",
          "post_date": "04/25/2021 19:31:26",
          "content": "<p>I tried it but still not successful. Can I know if you used kaggle resources itself or did you used may be GCP somehow when submitting?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1284335,
          "author_name": "vsharkeev",
          "author_url": "",
          "post_date": "04/25/2021 19:47:12",
          "content": "<p>I use only Kaggle resources for submission:</p>\n<pre>subm = {}\n\ndf_sample = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv')\n\nwith open(os.path.join(temp_path, 'test/test_shape.yaml'), 'r') as file:\n    shapes = yaml.safe_load(file)\n\nfor idx, row in tqdm(df_sample.iterrows(), total=len(df_sample)):\n    filename_stem = row['id']\n    filename = os.path.join(test_dir_name, filename_stem + '.tfrec')\n    print('{} Predicting {}'.format(idx+1, filename_stem))\n    shape = shapes[filename_stem]\n    preds = np.zeros(shape, dtype=np.uint8)\n\n    for imgs, X1, Y1 in load_dataset(filenames=filename, batch_size=16):\n        pred_avg_segm = None\n\n        # Predict mask itself\n        for model_segm in fold_models_segm:\n            pred_segm = model_segm.predict(imgs) / len(fold_models_segm)\n\n            if pred_avg_segm is None:\n                pred_avg_segm = pred_segm\n            else:\n                pred_avg_segm += pred_segm\n\n        pred_avg_segm = tf.cast((tf.image.resize(pred_avg_segm, (WINDOW,WINDOW)) &gt; THRESHOLD_SEGM), tf.bool).numpy().squeeze()\n\n        for i in range(X1.shape[0]):\n            preds[X1[i]:(X1[i]+WINDOW),Y1[i]:(Y1[i]+WINDOW)] += pred_avg_segm[i]\n\n    preds = (preds &gt; THRESHOLD_SEGM).astype(np.uint8)\n    print('Shape before: {}, after: {}'.format(shape, preds.shape))\n\n    subm[idx] = {'id':filename_stem, 'predicted': rle_encode_less_memory(preds)}\n\n    del preds\n    gc.collect()\n\ndf_submit = pd.DataFrame.from_dict(subm, orient='index')\ndf_submit = df_submit.set_index('id')\ndf_submit.to_csv('submission.csv')\n</pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1085142": "So, if I create a CSV locally for the test set file and upload it as a notebook dataset, and then submit it what problems I'm going to face? I have a much better inference speed locally and no memory bottleneck. I assume this will not work for final submission, but can I use it as a quick LB check?\nThanks!",
    "1085194": "Yes. This is a quick way to check your public LB score. But as you note, this will not work for the final leaderboard.\n\nTo do it, you can't just copy the submission.csv though. You'll get an error unless your submission.csv contains the private test ids too. Therefor you should copy the predictions from your submission.csv into the sample_submission.csv with something like this:\n```\nsubmission = pd.read_csv('submission.csv', index_col='id')\nsample_sub = pd.read_csv('sample_submission.csv', index_col='id')\npub_ids = submission.index.values\npredictions = submission.values\nsample_sub.loc[pub_ids] = predictions\nsample_sub.to_csv('submission.csv')\n```",
    "1085228": "Nice! Thanks!",
    "1088454": "Have anyone tried it? Submission works just fine but the score is extremely low. Visually the masks are more or less ok. Not perfect, but not garbage either. So I wonder if kernel test files are different from test files given for downloading.",
    "1088523": "I have tried to check my public LB score by this way and got 0,248. In my case, the masks contain lots of noise, so the result is expected. Hence, it doesn't seem kernel test files are different from test files given for downloading",
    "1088580": "The files should be the same. I haven't submitted yet though.\n\nI'm guessing there's a bug somewhere between you visualizing your masks and RLE encoding them. Perhaps the orientation of your labels are incorrect (i.e. vertically flipped).",
    "1088647": "God damn this rle encoding. Why it needs to be transposed! I've spent two days trying to debug this issue!",
    "1094865": "Hi, I just encountered the same problem. The score is extremely low. The generated mask looks ok, and it works when I did leave-one-out validation on training set. \nBut I've already make a transpose on the generated mask before RLE encoding. \nCould you please let me know how to deal with this issue? Many thanks!",
    "1095179": "Hi, I updated my submission.csv from my local machine and tried to submit it. But I was told \"Submission Scoring Error\". I've checked the format of my .csv file and it seems OK. Any idea to resolve this problem? I am a newbie on Kaggle paltform. Thanks in advance.",
    "1095196": "I've encountered the same issue. In my case, the error is due to improper assignment of index column. Pandas will add an index column automatically.",
    "1095374": "This could be due to RLE encoding of an image that differs slightly in dimensions from the original image. If you encode an image that is 999x999 and decode it as 1000x1000 there will be a shift of one pixel per line. But visually you will not see this. RLE doesn't save image dimensions so this can be hard to debuig.",
    "1095888": "Thank you Dennis! I've figured out this issue last night, which is due to the format of submission file. Thanks for your answer!",
    "1095913": "When you submit, your notebook is rerun with the full test dataset (public + private). Therefor, if you simply copy your submission.csv, it won't have the expected number of rows and will fail. You should do something like [what I posted below](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198317#1085194) to load your predictions into the sample_submission csv.",
    "1095930": "Does the scoring for LB use public+ private test dataset? If yes, even we convert our \"submission.csv\" to \"sample_submission.csv\", there are still 5 rows.\nSo how is the private test set included in the evaluation as we don't know the image_id for private test dataset.\nThank you in advance!",
    "1095942": "When you commit your notebook, you'll only ever see the public test set examples. The private test set files are only visible to your notebook during submission. When you submit your notebook, the dataset is replaced to include the private test set and the notebook is rerun.\n\nIt will score your predictions for both the private and public test sets, but only shows your public score until the end of the competition. For now, it's okay to leave the predictions for the private set blank. But for your final submission, you're gonna want to perform your prediction in the notebook.\n\nHope all that makes sense, I know it is confusing :)",
    "1096113": "You have already made it clearer! The only question left for me is -- How can I build a notebook that fits both conditions? Currently when doing test, I read the \"sample_submission.csv\" provided in the original dataset, use the \"id\" column to get which samples I need for test. But obviously this method cannot fit the situation with private test dataset. \n\nBTW is private test dataset used for current LB score? Or the score is just calculated by public dataset.\n\nThanks!",
    "1096459": "lijiaqi96 When you submit your notebook the \"sample_submission.csv\" gets replaced with a file that contains all image **id**'s of the hidden test set, so you don't need to make special changes for the private test run.\n\nThe public LB score you see now is using the public test dataset. The private test score is hidden till the end of the competition, which is why a high public LB score doesn't necessarily mean your final position will be the same.",
    "1096482": "Hello, have you solved this problem? Is it possible to make a submission from local CSV?",
    "1096491": "tikboa This is a code competition which means you are not allowed to submit you local CSV through a notebook to get a score. Even if you get a score this way for the public LB, your private test score will get an error since you have not made predictions for the hidden private test dataset.\n\nYou need to use your locally trained models to make inference in real time inside a notebook which you submit to the competition.",
    "1096565": "So I just keep getting test image ids from \"sample_submission.csv\" from \"/kaggle/input\". The same code will work for the final submission, am I right?",
    "1096571": "Yes, that is correct.",
    "1096943": "I might be missing something here. Is it correct that the models may be trained on personal workstations (without limitations on the training time), but the inference time for the final model when run on the kaggle notebook must be within the time limits specified (assuming the trained model can be uploaded directly from local)?",
    "1097318": "You are correct. No limitations on training time or resources.\n\nInference must be run within committed model with time limit of 9 hours. No TPU. No Internet. GPU allowed.",
    "1098516": "Thanks for the quick clarification!",
    "1249266": "I had Submission Error, but after copying of submission.csv as proposed here, I got public score. \nThank you, @matthewmasters!\n\nThe problem with my submission was that I've created .tfrec files in kaggle/working/ directory instead of kaggle/temp/.",
    "1281964": "I'm facing the same issue. Can you tell me if we need to save all output including submission.csv to temp folder?",
    "1283669": "I save submission.csv to kaggle/working and other output staff (.tfrec) to kaggle/temp/ and it works.",
    "1284314": "I tried it but still not successful. Can I know if you used kaggle resources itself or did you used may be GCP somehow when submitting?",
    "1284335": "I use only Kaggle resources for submission:\n\n<pre>\nsubm = {}\n\ndf_sample = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv')\n\nwith open(os.path.join(temp_path, 'test/test_shape.yaml'), 'r') as file:\n    shapes = yaml.safe_load(file)\n\nfor idx, row in tqdm(df_sample.iterrows(), total=len(df_sample)):\n    filename_stem = row['id']\n    filename = os.path.join(test_dir_name, filename_stem + '.tfrec')\n    print('{} Predicting {}'.format(idx+1, filename_stem))\n    shape = shapes[filename_stem]\n    preds = np.zeros(shape, dtype=np.uint8)\n\n    for imgs, X1, Y1 in load_dataset(filenames=filename, batch_size=16):\n        pred_avg_segm = None\n        \n        # Predict mask itself\n        for model_segm in fold_models_segm:\n            pred_segm = model_segm.predict(imgs) / len(fold_models_segm)\n\n            if pred_avg_segm is None:\n                pred_avg_segm = pred_segm\n            else:\n                pred_avg_segm += pred_segm\n\n        pred_avg_segm = tf.cast((tf.image.resize(pred_avg_segm, (WINDOW,WINDOW)) > THRESHOLD_SEGM), tf.bool).numpy().squeeze()\n        \n        for i in range(X1.shape[0]):\n            preds[X1[i]:(X1[i]+WINDOW),Y1[i]:(Y1[i]+WINDOW)] += pred_avg_segm[i]\n\n    preds = (preds > THRESHOLD_SEGM).astype(np.uint8)\n    print('Shape before: {}, after: {}'.format(shape, preds.shape))\n    \n    subm[idx] = {'id':filename_stem, 'predicted': rle_encode_less_memory(preds)}\n\n    del preds\n    gc.collect()\n\ndf_submit = pd.DataFrame.from_dict(subm, orient='index')\ndf_submit = df_submit.set_index('id')\ndf_submit.to_csv('submission.csv')\n</pre>"
  },
  "source": "meta"
}