{
  "id": 251021,
  "title": "Fast Sub Code",
  "url": "/competitions/siim-covid19-detection/discussion/251021",
  "author_name": "",
  "post_date": "2021-07-05T14:30:12.070005200Z",
  "votes": 13,
  "comment_count": 5,
  "views": 0,
  "content": "<p>You can submit your predictions only on public data<br>\nthis can save your time (since you will be waiting for inference for the whole test set)</p>\n<pre><code>import pandas as pd\n\nfast_sub = pd.read_csv('../input/mmdetection-submission/submission.csv')\nsample_submission = pd.read_csv('../input/siim-covid19-detection/sample_submission.csv')\nstudy_sample = sample_submission[sample_submission.id.str.contains('study')]\\\n                .drop(['PredictionString'],axis = 1).reset_index(drop = True)\nimage_sample = sample_submission[sample_submission.id.str.contains('image')]\\\n                .drop(['PredictionString'],axis = 1).reset_index(drop = True)\n\nstudy_submission = fast_sub[fast_sub.id.str.contains('study')].reset_index(drop = True)\nimage_submission = fast_sub[fast_sub.id.str.contains('image')].reset_index(drop = True)\n\npd.concat([study_sample.merge(study_submission, on = 'id', how = 'outer').fillna('negative 1 0 0 1 1'),\\\n           image_sample.merge(image_submission, on = 'id', how = 'outer').fillna('none 1 0 0 1 1')])\\\n            .reset_index(drop = True).to_csv('submission.csv', index = False)\n</code></pre>",
  "messages": [
    {
      "id": "1377035",
      "postDate": "07/05/2021 14:30:12",
      "content": "<p>You can submit your predictions only on public data<br>\nthis can save your time (since you will be waiting for inference for the whole test set)</p>\n<pre><code>import pandas as pd\n\nfast_sub = pd.read_csv('../input/mmdetection-submission/submission.csv')\nsample_submission = pd.read_csv('../input/siim-covid19-detection/sample_submission.csv')\nstudy_sample = sample_submission[sample_submission.id.str.contains('study')]\\\n                .drop(['PredictionString'],axis = 1).reset_index(drop = True)\nimage_sample = sample_submission[sample_submission.id.str.contains('image')]\\\n                .drop(['PredictionString'],axis = 1).reset_index(drop = True)\n\nstudy_submission = fast_sub[fast_sub.id.str.contains('study')].reset_index(drop = True)\nimage_submission = fast_sub[fast_sub.id.str.contains('image')].reset_index(drop = True)\n\npd.concat([study_sample.merge(study_submission, on = 'id', how = 'outer').fillna('negative 1 0 0 1 1'),\\\n           image_sample.merge(image_submission, on = 'id', how = 'outer').fillna('none 1 0 0 1 1')])\\\n            .reset_index(drop = True).to_csv('submission.csv', index = False)\n</code></pre>",
      "rawMarkdown": "You can submit your predictions only on public data\nthis can save your time (since you will be waiting for inference for the whole test set)\n\n```\nimport pandas as pd\n\nfast_sub = pd.read_csv('../input/mmdetection-submission/submission.csv')\nsample_submission = pd.read_csv('../input/siim-covid19-detection/sample_submission.csv')\nstudy_sample = sample_submission[sample_submission.id.str.contains('study')]\\\n                .drop(['PredictionString'],axis = 1).reset_index(drop = True)\nimage_sample = sample_submission[sample_submission.id.str.contains('image')]\\\n                .drop(['PredictionString'],axis = 1).reset_index(drop = True)\n\nstudy_submission = fast_sub[fast_sub.id.str.contains('study')].reset_index(drop = True)\nimage_submission = fast_sub[fast_sub.id.str.contains('image')].reset_index(drop = True)\n\npd.concat([study_sample.merge(study_submission, on = 'id', how = 'outer').fillna('negative 1 0 0 1 1'),\\\n           image_sample.merge(image_submission, on = 'id', how = 'outer').fillna('none 1 0 0 1 1')])\\\n            .reset_index(drop = True).to_csv('submission.csv', index = False)\n```",
      "votes": null
    },
    {
      "id": "1378293",
      "postDate": "07/06/2021 12:33:19",
      "content": "<p>I don't understand the concept here <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> .. I thought the sample submission file represented the entire public test set .. for which, we must predict at least one row. Could you explain this a little more? </p>",
      "rawMarkdown": "I don't understand the concept here @morizin .. I thought the sample submission file represented the entire public test set .. for which, we must predict at least one row. Could you explain this a little more?",
      "votes": null
    },
    {
      "id": "1378646",
      "postDate": "07/06/2021 17:22:18",
      "content": "<p>The idea is to get your LB result faster rather than waiting for hrs to see your public LB score.<br>\nif you have a notebook that predicts the public set and generates a submission.csv file out of it<br>\nin this case you have to submit the notebook in order to run the private set also(this part is where it takes time). to get the instant results we can just submit the only pub only by inserting your submission file here. This will only submit public and give you the score of public and private score will remain 0.0 <br>\n<code>fast_sub = pd.read_csv('your file here')</code> in above code</p>",
      "rawMarkdown": "The idea is to get your LB result faster rather than waiting for hrs to see your public LB score.\nif you have a notebook that predicts the public set and generates a submission.csv file out of it\nin this case you have to submit the notebook in order to run the private set also(this part is where it takes time). to get the instant results we can just submit the only pub only by inserting your submission file here. This will only submit public and give you the score of public and private score will remain 0.0 \n```fast_sub = pd.read_csv('your file here')``` in above code",
      "votes": null
    },
    {
      "id": "1378679",
      "postDate": "07/06/2021 17:44:39",
      "content": "<p>Thanks for your reply!</p>\n<p>I get it now. But, if that's the case, what's the splitting and merging for? Couldn't you just load the inference submission into a new notebook and export it directly since it should match the sample_sub already?</p>",
      "rawMarkdown": "Thanks for your reply!\n\nI get it now. But, if that's the case, what's the splitting and merging for? Couldn't you just load the inference submission into a new notebook and export it directly since it should match the sample_sub already?",
      "votes": null
    },
    {
      "id": "1378724",
      "postDate": "07/06/2021 18:38:44",
      "content": "<p>the reason I am splitting merging is because<br>\nActually when we run this in a notebook. and still we are submitting the notebook itself . but the change from previous case is:</p>\n<ol>\n<li>when we submit the notebook the <code>sample_submission.csv</code> will contain both public and private. but since we dont want to private we need to keep them as default values as <code>none 1 0 0 1 1</code> for study level and <code>negative 1 0 0 1 1</code> for study<br>\nSo we need to split the csv file into study and image level so that we can give same default values and then we merge the prediction part of public set with whole submission</li>\n</ol>\n<p>since these processes cannot be seen by eyes. i dont know how to explain it to you more. I hope its clear to you.</p>",
      "rawMarkdown": "the reason I am splitting merging is because\nActually when we run this in a notebook. and still we are submitting the notebook itself . but the change from previous case is:\n1. when we submit the notebook the `sample_submission.csv` will contain both public and private. but since we dont want to private we need to keep them as default values as `none 1 0 0 1 1` for study level and `negative 1 0 0 1 1` for study\nSo we need to split the csv file into study and image level so that we can give same default values and then we merge the prediction part of public set with whole submission\n\nsince these processes cannot be seen by eyes. i dont know how to explain it to you more. I hope its clear to you.",
      "votes": null
    },
    {
      "id": "1378738",
      "postDate": "07/06/2021 18:57:04",
      "content": "<p>Ah, I see. I thought the sample_submission file only contained public data and was replaced with private data upon submission. Thanks for clearing it up.</p>",
      "rawMarkdown": "Ah, I see. I thought the sample_submission file only contained public data and was replaced with private data upon submission. Thanks for clearing it up.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1378293,
      "author_name": "davidbroberts",
      "author_url": "",
      "post_date": "07/06/2021 12:33:19",
      "content": "<p>I don't understand the concept here <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> .. I thought the sample submission file represented the entire public test set .. for which, we must predict at least one row. Could you explain this a little more? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1378646,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "07/06/2021 17:22:18",
          "content": "<p>The idea is to get your LB result faster rather than waiting for hrs to see your public LB score.<br>\nif you have a notebook that predicts the public set and generates a submission.csv file out of it<br>\nin this case you have to submit the notebook in order to run the private set also(this part is where it takes time). to get the instant results we can just submit the only pub only by inserting your submission file here. This will only submit public and give you the score of public and private score will remain 0.0 <br>\n<code>fast_sub = pd.read_csv('your file here')</code> in above code</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1378679,
          "author_name": "davidbroberts",
          "author_url": "",
          "post_date": "07/06/2021 17:44:39",
          "content": "<p>Thanks for your reply!</p>\n<p>I get it now. But, if that's the case, what's the splitting and merging for? Couldn't you just load the inference submission into a new notebook and export it directly since it should match the sample_sub already?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1378724,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "07/06/2021 18:38:44",
          "content": "<p>the reason I am splitting merging is because<br>\nActually when we run this in a notebook. and still we are submitting the notebook itself . but the change from previous case is:</p>\n<ol>\n<li>when we submit the notebook the <code>sample_submission.csv</code> will contain both public and private. but since we dont want to private we need to keep them as default values as <code>none 1 0 0 1 1</code> for study level and <code>negative 1 0 0 1 1</code> for study<br>\nSo we need to split the csv file into study and image level so that we can give same default values and then we merge the prediction part of public set with whole submission</li>\n</ol>\n<p>since these processes cannot be seen by eyes. i dont know how to explain it to you more. I hope its clear to you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1378738,
          "author_name": "davidbroberts",
          "author_url": "",
          "post_date": "07/06/2021 18:57:04",
          "content": "<p>Ah, I see. I thought the sample_submission file only contained public data and was replaced with private data upon submission. Thanks for clearing it up.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1377035": "You can submit your predictions only on public data\nthis can save your time (since you will be waiting for inference for the whole test set)\n\n```\nimport pandas as pd\n\nfast_sub = pd.read_csv('../input/mmdetection-submission/submission.csv')\nsample_submission = pd.read_csv('../input/siim-covid19-detection/sample_submission.csv')\nstudy_sample = sample_submission[sample_submission.id.str.contains('study')]\\\n                .drop(['PredictionString'],axis = 1).reset_index(drop = True)\nimage_sample = sample_submission[sample_submission.id.str.contains('image')]\\\n                .drop(['PredictionString'],axis = 1).reset_index(drop = True)\n\nstudy_submission = fast_sub[fast_sub.id.str.contains('study')].reset_index(drop = True)\nimage_submission = fast_sub[fast_sub.id.str.contains('image')].reset_index(drop = True)\n\npd.concat([study_sample.merge(study_submission, on = 'id', how = 'outer').fillna('negative 1 0 0 1 1'),\\\n           image_sample.merge(image_submission, on = 'id', how = 'outer').fillna('none 1 0 0 1 1')])\\\n            .reset_index(drop = True).to_csv('submission.csv', index = False)\n```",
    "1378293": "I don't understand the concept here @morizin .. I thought the sample submission file represented the entire public test set .. for which, we must predict at least one row. Could you explain this a little more?",
    "1378646": "The idea is to get your LB result faster rather than waiting for hrs to see your public LB score.\nif you have a notebook that predicts the public set and generates a submission.csv file out of it\nin this case you have to submit the notebook in order to run the private set also(this part is where it takes time). to get the instant results we can just submit the only pub only by inserting your submission file here. This will only submit public and give you the score of public and private score will remain 0.0 \n```fast_sub = pd.read_csv('your file here')``` in above code",
    "1378679": "Thanks for your reply!\n\nI get it now. But, if that's the case, what's the splitting and merging for? Couldn't you just load the inference submission into a new notebook and export it directly since it should match the sample_sub already?",
    "1378724": "the reason I am splitting merging is because\nActually when we run this in a notebook. and still we are submitting the notebook itself . but the change from previous case is:\n1. when we submit the notebook the `sample_submission.csv` will contain both public and private. but since we dont want to private we need to keep them as default values as `none 1 0 0 1 1` for study level and `negative 1 0 0 1 1` for study\nSo we need to split the csv file into study and image level so that we can give same default values and then we merge the prediction part of public set with whole submission\n\nsince these processes cannot be seen by eyes. i dont know how to explain it to you more. I hope its clear to you.",
    "1378738": "Ah, I see. I thought the sample_submission file only contained public data and was replaced with private data upon submission. Thanks for clearing it up."
  },
  "source": "meta"
}