{
  "id": 118129,
  "title": "Save time committing & submitting",
  "url": "/competitions/tensorflow2-question-answering/discussion/118129",
  "author_name": "",
  "post_date": "2019-11-19T17:10:57.482958Z",
  "votes": 25,
  "comment_count": 13,
  "views": 0,
  "content": "<p><code>\nif len(test) &gt; 346:\n    do_stuff()\nelse:\n    sample_submission.to_csv('submission.csv', index=False)\n</code>\nThus your Commit will be almost instant and you'll go right away to Submitting, i.e. running your <code>do_stuff</code> against the private test part (which is longer than 346). </p>",
  "messages": [
    {
      "id": "676976",
      "postDate": "11/19/2019 17:10:57",
      "content": "<p><code>\nif len(test) &gt; 346:\n    do_stuff()\nelse:\n    sample_submission.to_csv('submission.csv', index=False)\n</code>\nThus your Commit will be almost instant and you'll go right away to Submitting, i.e. running your <code>do_stuff</code> against the private test part (which is longer than 346). </p>",
      "rawMarkdown": "```\nif len(test) &gt; 346:\n    do_stuff()\nelse:\n    sample_submission.to_csv('submission.csv', index=False)\n```\nThus your Commit will be almost instant and you'll go right away to Submitting, i.e. running your `do_stuff` against the private test part (which is longer than 346).",
      "votes": null
    },
    {
      "id": "677001",
      "postDate": "11/19/2019 17:39:24",
      "content": "<p>Thanks, <a href=\"/kashnitsky\">@kashnitsky</a> this helps so much, for no more long waiting</p>",
      "rawMarkdown": "Thanks, @kashnitsky this helps so much, for no more long waiting",
      "votes": null
    },
    {
      "id": "677151",
      "postDate": "11/19/2019 21:25:44",
      "content": "<p>Are we sure that the current test dataset used when we submit to have the public leaderboard will be the same one used  for the final private leaderboard? Very likely no, and we have to be careful (although I think it will be larger than 346 ... so kind safe)</p>",
      "rawMarkdown": "Are we sure that the current test dataset used when we submit to have the public leaderboard will be the same one used  for the final private leaderboard? Very likely no, and we have to be careful (although I think it will be larger than 346 ... so kind safe)",
      "votes": null
    },
    {
      "id": "677174",
      "postDate": "11/19/2019 22:01:05",
      "content": "<p>No one says that it’s a final version of your Kernel :)</p>",
      "rawMarkdown": "No one says that it’s a final version of your Kernel :)",
      "votes": null
    },
    {
      "id": "683359",
      "postDate": "11/28/2019 10:50:19",
      "content": "<p>This needs to be applied only with the 100% working code. See caveats <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280#683358\">here</a>. </p>",
      "rawMarkdown": "This needs to be applied only with the 100% working code. See caveats [here](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280#683358).",
      "votes": null
    },
    {
      "id": "695432",
      "postDate": "12/15/2019 06:22:03",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> thanks for sharing this!\nI'm trying to mimic what you showed here. But it ends up 0.0 public score.  Does anyone have any ideas why?</p>\n\n<p>I published my failing kernel below.\n<a href=\"https://www.kaggle.com/higepon/how-to-get-public-score-quickly\">https://www.kaggle.com/higepon/how-to-get-public-score-quickly</a></p>",
      "rawMarkdown": "kashnitsky thanks for sharing this!\nI'm trying to mimic what you showed here. But it ends up 0.0 public score.  Does anyone have any ideas why?\n\nI published my failing kernel below.\nhttps://www.kaggle.com/higepon/how-to-get-public-score-quickly",
      "votes": null
    },
    {
      "id": "695762",
      "postDate": "12/15/2019 14:32:20",
      "content": "<p>No longer working fo me either. But 0.0 can mean that your code fails on the private test part. Check with a 100% working model like \"first long paragraph\".</p>",
      "rawMarkdown": "No longer working fo me either. But 0.0 can mean that your code fails on the private test part. Check with a 100% working model like \"first long paragraph\".",
      "votes": null
    },
    {
      "id": "695960",
      "postDate": "12/15/2019 23:46:49",
      "content": "<p>Thank you! Good to know it's not working for you either.\nI'll try the other way.</p>",
      "rawMarkdown": "Thank you! Good to know it's not working for you either.\nI'll try the other way.",
      "votes": null
    },
    {
      "id": "697747",
      "postDate": "12/18/2019 10:53:11",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> Unfortunately, at least for me, Submitting takes much more time as compared to Commit. However, something is better than nothing.</p>",
      "rawMarkdown": "kashnitsky Unfortunately, at least for me, Submitting takes much more time as compared to Commit. However, something is better than nothing.",
      "votes": null
    },
    {
      "id": "697844",
      "postDate": "12/18/2019 13:38:24",
      "content": "<p>It’s the same for everybody, because private test set is some 10x larger than the public test set (this 346-long). </p>",
      "rawMarkdown": "It’s the same for everybody, because private test set is some 10x larger than the public test set (this 346-long).",
      "votes": null
    },
    {
      "id": "717932",
      "postDate": "01/13/2020 19:57:36",
      "content": "<p>I am not following how it saves time. <a href=\"/kashnitsky\">@kashnitsky</a> Could you give more details?</p>\n\n<p>If it's committing, I don't see any difference in the execution time with single <code>sample_submission.to_csv('submission.csv', index=False)</code>\nIf it's submitting, it'll spend extra on <code>do_stuff</code>.</p>\n\n<p>Here's the difference for me between public and private commit\\submission - the private test set is larger, hence more time spent on models training, which I naturaly want.</p>",
      "rawMarkdown": "I am not following how it saves time. @kashnitsky Could you give more details?\n\nIf it's committing, I don't see any difference in the execution time with single `sample_submission.to_csv('submission.csv', index=False)`\nIf it's submitting, it'll spend extra on `do_stuff`.\n\nHere's the difference for me between public and private commit\\submission - the private test set is larger, hence more time spent on models training, which I naturaly want.",
      "votes": null
    },
    {
      "id": "718026",
      "postDate": "01/13/2020 22:42:53",
      "content": "<p>It skips public part and performs interesting computations with the hidden part. </p>\n\n<p>But a better approach is already shared. </p>",
      "rawMarkdown": "It skips public part and performs interesting computations with the hidden part. \n\nBut a better approach is already shared.",
      "votes": null
    },
    {
      "id": "718657",
      "postDate": "01/14/2020 16:08:04",
      "content": "<p>Got it - like skeeping EDA and everything else you don't need for submittion.</p>",
      "rawMarkdown": "Got it - like skeeping EDA and everything else you don't need for submittion.",
      "votes": null
    },
    {
      "id": "728639",
      "postDate": "01/25/2020 02:03:31",
      "content": "<p>this is so good. Thanks <a href=\"/kashnitsky\">@kashnitsky</a> </p>",
      "rawMarkdown": "this is so good. Thanks @kashnitsky",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 677001,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "11/19/2019 17:39:24",
      "content": "<p>Thanks, <a href=\"/kashnitsky\">@kashnitsky</a> this helps so much, for no more long waiting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 677151,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "11/19/2019 21:25:44",
      "content": "<p>Are we sure that the current test dataset used when we submit to have the public leaderboard will be the same one used  for the final private leaderboard? Very likely no, and we have to be careful (although I think it will be larger than 346 ... so kind safe)</p>",
      "votes": null,
      "replies": [
        {
          "id": 677174,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "11/19/2019 22:01:05",
          "content": "<p>No one says that it’s a final version of your Kernel :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 683359,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "11/28/2019 10:50:19",
      "content": "<p>This needs to be applied only with the 100% working code. See caveats <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280#683358\">here</a>. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 695432,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "12/15/2019 06:22:03",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> thanks for sharing this!\nI'm trying to mimic what you showed here. But it ends up 0.0 public score.  Does anyone have any ideas why?</p>\n\n<p>I published my failing kernel below.\n<a href=\"https://www.kaggle.com/higepon/how-to-get-public-score-quickly\">https://www.kaggle.com/higepon/how-to-get-public-score-quickly</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 695762,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/15/2019 14:32:20",
          "content": "<p>No longer working fo me either. But 0.0 can mean that your code fails on the private test part. Check with a 100% working model like \"first long paragraph\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 695960,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/15/2019 23:46:49",
          "content": "<p>Thank you! Good to know it's not working for you either.\nI'll try the other way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 697747,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "12/18/2019 10:53:11",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> Unfortunately, at least for me, Submitting takes much more time as compared to Commit. However, something is better than nothing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 697844,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/18/2019 13:38:24",
          "content": "<p>It’s the same for everybody, because private test set is some 10x larger than the public test set (this 346-long). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 717932,
      "author_name": "manyregression",
      "author_url": "",
      "post_date": "01/13/2020 19:57:36",
      "content": "<p>I am not following how it saves time. <a href=\"/kashnitsky\">@kashnitsky</a> Could you give more details?</p>\n\n<p>If it's committing, I don't see any difference in the execution time with single <code>sample_submission.to_csv('submission.csv', index=False)</code>\nIf it's submitting, it'll spend extra on <code>do_stuff</code>.</p>\n\n<p>Here's the difference for me between public and private commit\\submission - the private test set is larger, hence more time spent on models training, which I naturaly want.</p>",
      "votes": null,
      "replies": [
        {
          "id": 718026,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/13/2020 22:42:53",
          "content": "<p>It skips public part and performs interesting computations with the hidden part. </p>\n\n<p>But a better approach is already shared. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 718657,
          "author_name": "manyregression",
          "author_url": "",
          "post_date": "01/14/2020 16:08:04",
          "content": "<p>Got it - like skeeping EDA and everything else you don't need for submittion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 728639,
      "author_name": "adityachhabra",
      "author_url": "",
      "post_date": "01/25/2020 02:03:31",
      "content": "<p>this is so good. Thanks <a href=\"/kashnitsky\">@kashnitsky</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676976": "```\nif len(test) &gt; 346:\n    do_stuff()\nelse:\n    sample_submission.to_csv('submission.csv', index=False)\n```\nThus your Commit will be almost instant and you'll go right away to Submitting, i.e. running your `do_stuff` against the private test part (which is longer than 346).",
    "677001": "Thanks, @kashnitsky this helps so much, for no more long waiting",
    "677151": "Are we sure that the current test dataset used when we submit to have the public leaderboard will be the same one used  for the final private leaderboard? Very likely no, and we have to be careful (although I think it will be larger than 346 ... so kind safe)",
    "677174": "No one says that it’s a final version of your Kernel :)",
    "683359": "This needs to be applied only with the 100% working code. See caveats [here](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280#683358).",
    "695432": "kashnitsky thanks for sharing this!\nI'm trying to mimic what you showed here. But it ends up 0.0 public score.  Does anyone have any ideas why?\n\nI published my failing kernel below.\nhttps://www.kaggle.com/higepon/how-to-get-public-score-quickly",
    "695762": "No longer working fo me either. But 0.0 can mean that your code fails on the private test part. Check with a 100% working model like \"first long paragraph\".",
    "695960": "Thank you! Good to know it's not working for you either.\nI'll try the other way.",
    "697747": "kashnitsky Unfortunately, at least for me, Submitting takes much more time as compared to Commit. However, something is better than nothing.",
    "697844": "It’s the same for everybody, because private test set is some 10x larger than the public test set (this 346-long).",
    "717932": "I am not following how it saves time. @kashnitsky Could you give more details?\n\nIf it's committing, I don't see any difference in the execution time with single `sample_submission.to_csv('submission.csv', index=False)`\nIf it's submitting, it'll spend extra on `do_stuff`.\n\nHere's the difference for me between public and private commit\\submission - the private test set is larger, hence more time spent on models training, which I naturaly want.",
    "718026": "It skips public part and performs interesting computations with the hidden part. \n\nBut a better approach is already shared.",
    "718657": "Got it - like skeeping EDA and everything else you don't need for submittion.",
    "728639": "this is so good. Thanks @kashnitsky"
  },
  "source": "meta"
}