{
  "id": 69537,
  "title": "A quick fix to get submission Ids in order",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/69537",
  "author_name": "Tilii",
  "post_date": "2018-10-24T17:34:38.438000",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>It seems that lots of leaderboard scores were lower than what they should have been due to wrong order of submission Ids, as described <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366\"><strong>here</strong></a>. Below is a short code that will fix a submission file by aligning its Ids with <code>sample_submission.csv</code>. It assumes that your submission file is otherwise in correct format.</p>\n\n<pre><code>import pandas as pd\n\nf1 = pd.read_csv('sample_submission.csv')\nf1.drop('Predicted', axis=1, inplace=True)\nf2 = pd.read_csv('my_awesome_submission.csv')\nf1 = f1.merge(f2, left_on='Id', right_on='Id', how='outer')\nf1.to_csv('my_new_submission.csv', index=False)\n</code></pre>",
  "messages": [
    {
      "id": 409680,
      "postDate": "2018-10-24T17:34:38.440Z",
      "content": "<p>It seems that lots of leaderboard scores were lower than what they should have been due to wrong order of submission Ids, as described <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366\"><strong>here</strong></a>. Below is a short code that will fix a submission file by aligning its Ids with <code>sample_submission.csv</code>. It assumes that your submission file is otherwise in correct format.</p>\n\n<pre><code>import pandas as pd\n\nf1 = pd.read_csv('sample_submission.csv')\nf1.drop('Predicted', axis=1, inplace=True)\nf2 = pd.read_csv('my_awesome_submission.csv')\nf1 = f1.merge(f2, left_on='Id', right_on='Id', how='outer')\nf1.to_csv('my_new_submission.csv', index=False)\n</code></pre>",
      "rawMarkdown": "It seems that lots of leaderboard scores were lower than what they should have been due to wrong order of submission Ids, as described [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366). Below is a short code that will fix a submission file by aligning its Ids with `sample_submission.csv`. It assumes that your submission file is otherwise in correct format.\n\n    import pandas as pd\n\n    f1 = pd.read_csv('sample_submission.csv')\n    f1.drop('Predicted', axis=1, inplace=True)\n    f2 = pd.read_csv('my_awesome_submission.csv')\n    f1 = f1.merge(f2, left_on='Id', right_on='Id', how='outer')\n    f1.to_csv('my_new_submission.csv', index=False)\n\n",
      "votes": 15
    },
    {
      "id": 409683,
      "postDate": "2018-10-24T17:37:32.607Z",
      "content": "<p>Thank you! Seems kind of silly to impose an order. I wonder if this is to improve scoring efficiency?</p>",
      "rawMarkdown": "Thank you! Seems kind of silly to impose an order. I wonder if this is to improve scoring efficiency?",
      "votes": 3,
      "replies": [
        {
          "id": 409688,
          "postDate": "2018-10-24T17:44:48.540Z",
          "content": "<p><a href=\"/nnnnick\">@nnnnick</a> I have no idea. It is easy enough for Kaggle to do the same few steps as outlined above before scoring submissions. Anecdotally, we could get away with an arbitrary ID order in previous competitions as long as all samples were there. Not sure why it is not so in this competition.</p>",
          "rawMarkdown": "@nnnnick I have no idea. It is easy enough for Kaggle to do the same few steps as outlined above before scoring submissions. Anecdotally, we could get away with an arbitrary ID order in previous competitions as long as all samples were there. Not sure why it is not so in this competition.\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 409721,
      "postDate": "2018-10-24T18:42:36.077Z",
      "content": "<p>Cool, thanks! Is it in general a good idea to use the sample submission as a template for a new submission or would it take unnecessary computing power?</p>",
      "rawMarkdown": "Cool, thanks! Is it in general a good idea to use the sample submission as a template for a new submission or would it take unnecessary computing power?",
      "votes": 1,
      "replies": [
        {
          "id": 409726,
          "postDate": "2018-10-24T18:52:33.480Z",
          "content": "<blockquote>\n  <p>Cool, thanks! Is it in general a good idea to use the sample submission as a template for a new submission or would it take unnecessary computing power?</p>\n</blockquote>\n\n<p><a href=\"/carlolepelaars\">@carlolepelaars</a> I always submit with Ids in the same order as in sample file. Usually I do that as part of model generation so there is no extra computing power needed. If somehow you end up with a submission file where Ids are out of order compared to sample, running those few lines shown above takes about a second.</p>",
          "rawMarkdown": "&gt; Cool, thanks! Is it in general a good idea to use the sample submission as a template for a new submission or would it take unnecessary computing power?\n\n@carlolepelaars I always submit with Ids in the same order as in sample file. Usually I do that as part of model generation so there is no extra computing power needed. If somehow you end up with a submission file where Ids are out of order compared to sample, running those few lines shown above takes about a second.",
          "votes": 1
        },
        {
          "id": 409829,
          "postDate": "2018-10-24T22:53:23.343Z",
          "content": "<p>Awesome! Thank you for the tip!</p>",
          "rawMarkdown": "Awesome! Thank you for the tip!"
        }
      ]
    },
    {
      "id": 409919,
      "postDate": "2018-10-25T04:16:06.977Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 409683,
      "author_name": "Nick",
      "author_url": "",
      "post_date": "2018-10-24T17:37:32.607000",
      "content": "<p>Thank you! Seems kind of silly to impose an order. I wonder if this is to improve scoring efficiency?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 409688,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-10-24T17:44:48.540000",
          "content": "<p><a href=\"/nnnnick\">@nnnnick</a> I have no idea. It is easy enough for Kaggle to do the same few steps as outlined above before scoring submissions. Anecdotally, we could get away with an arbitrary ID order in previous competitions as long as all samples were there. Not sure why it is not so in this competition.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 409721,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2018-10-24T18:42:36.077000",
      "content": "<p>Cool, thanks! Is it in general a good idea to use the sample submission as a template for a new submission or would it take unnecessary computing power?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 409726,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-10-24T18:52:33.480000",
          "content": "<blockquote>\n  <p>Cool, thanks! Is it in general a good idea to use the sample submission as a template for a new submission or would it take unnecessary computing power?</p>\n</blockquote>\n\n<p><a href=\"/carlolepelaars\">@carlolepelaars</a> I always submit with Ids in the same order as in sample file. Usually I do that as part of model generation so there is no extra computing power needed. If somehow you end up with a submission file where Ids are out of order compared to sample, running those few lines shown above takes about a second.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 409829,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2018-10-24T22:53:23.343000",
          "content": "<p>Awesome! Thank you for the tip!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 409919,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-25T04:16:06.977000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "409680": "It seems that lots of leaderboard scores were lower than what they should have been due to wrong order of submission Ids, as described [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366). Below is a short code that will fix a submission file by aligning its Ids with `sample_submission.csv`. It assumes that your submission file is otherwise in correct format.\n\n    import pandas as pd\n\n    f1 = pd.read_csv('sample_submission.csv')\n    f1.drop('Predicted', axis=1, inplace=True)\n    f2 = pd.read_csv('my_awesome_submission.csv')\n    f1 = f1.merge(f2, left_on='Id', right_on='Id', how='outer')\n    f1.to_csv('my_new_submission.csv', index=False)\n\n",
    "409683": "Thank you! Seems kind of silly to impose an order. I wonder if this is to improve scoring efficiency?",
    "409721": "Cool, thanks! Is it in general a good idea to use the sample submission as a template for a new submission or would it take unnecessary computing power?",
    "409919": ""
  }
}