{
  "id": 79593,
  "title": "URGENT. How much runtime buffer are you leaving for stage 2 ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79593",
  "author_name": "",
  "post_date": "2019-02-05T20:46:03.768038Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello !\nHow much runtime buffer are you guys leaving for stage 2 evaluation.</p>",
  "messages": [
    {
      "id": "466698",
      "postDate": "02/05/2019 20:46:03",
      "content": "<p>Hello !\nHow much runtime buffer are you guys leaving for stage 2 evaluation.</p>",
      "rawMarkdown": "Hello !\nHow much runtime buffer are you guys leaving for stage 2 evaluation.",
      "votes": null
    },
    {
      "id": "466717",
      "postDate": "02/05/2019 21:18:22",
      "content": "<p>Why are people downvoting this ? Is there anything wrong in this question ?</p>",
      "rawMarkdown": "Why are people downvoting this ? Is there anything wrong in this question ?",
      "votes": null
    },
    {
      "id": "466721",
      "postDate": "02/05/2019 21:30:42",
      "content": "<p>Highly depends on your preprocessing and number of times the predictions are made of the test data. \nLet's say now you are taking 2 mins for preprocessing test data and making the predictions 5 time (30 secs each)</p>\n\n<p>So you should atleast save (2*60 + 5*30)*7 secs, since the 2nd Stage would be 7 times bigger. A save bet would be to have (2*60 + 5*30)*9 secs buffer to incorporate some other stuff, like if you are using tokenizer on test data as well as train data.</p>\n\n<p>Hope this helps. Although it's a little late now. </p>",
      "rawMarkdown": "Highly depends on your preprocessing and number of times the predictions are made of the test data. \nLet's say now you are taking 2 mins for preprocessing test data and making the predictions 5 time (30 secs each)\n\nSo you should atleast save (2*60 + 5*30)*7 secs, since the 2nd Stage would be 7 times bigger. A save bet would be to have (2*60 + 5*30)*9 secs buffer to incorporate some other stuff, like if you are using tokenizer on test data as well as train data.\n\nHope this helps. Although it's a little late now.",
      "votes": null
    },
    {
      "id": "466794",
      "postDate": "02/05/2019 23:55:19",
      "content": "<p>I mark this question sincere</p>",
      "rawMarkdown": "I mark this question sincere",
      "votes": null
    },
    {
      "id": "466797",
      "postDate": "02/05/2019 23:59:36",
      "content": "<p>To answer your question we simulated the increased test size by repeating the test file 6 more times. That is probably the easiest way to get an estimate. We are still allowing for some small buffer though just in case. </p>",
      "rawMarkdown": "To answer your question we simulated the increased test size by repeating the test file 6 more times. That is probably the easiest way to get an estimate. We are still allowing for some small buffer though just in case.",
      "votes": null
    },
    {
      "id": "468143",
      "postDate": "02/08/2019 11:06:32",
      "content": "<p>I tested my kernel on the new test data, and it runs in time (with 200s to spare). I'm so relieved. Phew !</p>",
      "rawMarkdown": "I tested my kernel on the new test data, and it runs in time (with 200s to spare). I'm so relieved. Phew !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 466717,
      "author_name": "tarunpaparaju",
      "author_url": "",
      "post_date": "02/05/2019 21:18:22",
      "content": "<p>Why are people downvoting this ? Is there anything wrong in this question ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 466794,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/05/2019 23:55:19",
          "content": "<p>I mark this question sincere</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 466721,
      "author_name": "abdurrafae",
      "author_url": "",
      "post_date": "02/05/2019 21:30:42",
      "content": "<p>Highly depends on your preprocessing and number of times the predictions are made of the test data. \nLet's say now you are taking 2 mins for preprocessing test data and making the predictions 5 time (30 secs each)</p>\n\n<p>So you should atleast save (2*60 + 5*30)*7 secs, since the 2nd Stage would be 7 times bigger. A save bet would be to have (2*60 + 5*30)*9 secs buffer to incorporate some other stuff, like if you are using tokenizer on test data as well as train data.</p>\n\n<p>Hope this helps. Although it's a little late now. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 466797,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/05/2019 23:59:36",
      "content": "<p>To answer your question we simulated the increased test size by repeating the test file 6 more times. That is probably the easiest way to get an estimate. We are still allowing for some small buffer though just in case. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 468143,
      "author_name": "tarunpaparaju",
      "author_url": "",
      "post_date": "02/08/2019 11:06:32",
      "content": "<p>I tested my kernel on the new test data, and it runs in time (with 200s to spare). I'm so relieved. Phew !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "466698": "Hello !\nHow much runtime buffer are you guys leaving for stage 2 evaluation.",
    "466717": "Why are people downvoting this ? Is there anything wrong in this question ?",
    "466721": "Highly depends on your preprocessing and number of times the predictions are made of the test data. \nLet's say now you are taking 2 mins for preprocessing test data and making the predictions 5 time (30 secs each)\n\nSo you should atleast save (2*60 + 5*30)*7 secs, since the 2nd Stage would be 7 times bigger. A save bet would be to have (2*60 + 5*30)*9 secs buffer to incorporate some other stuff, like if you are using tokenizer on test data as well as train data.\n\nHope this helps. Although it's a little late now.",
    "466794": "I mark this question sincere",
    "466797": "To answer your question we simulated the increased test size by repeating the test file 6 more times. That is probably the easiest way to get an estimate. We are still allowing for some small buffer though just in case.",
    "468143": "I tested my kernel on the new test data, and it runs in time (with 200s to spare). I'm so relieved. Phew !"
  },
  "source": "meta"
}