{
  "id": 93781,
  "title": "Be care of 5 times test set in second-stage!(generator share)",
  "url": "/competitions/imet-2019-fgvc6/discussion/93781",
  "author_name": "",
  "post_date": "2019-05-30T01:49:43.602367Z",
  "votes": 19,
  "comment_count": 9,
  "views": 0,
  "content": "In Data Description it said that <strong>The second-stage test set is approximately five times the size of the first</strong> , which means more strict time limit, disk limit and memory limit we have to face. Your final kernel for prediction may not exceed 1.8h ( just roughly divided the kernel time limit 9h by 5). These things really make us anxious, so why not test it before the second-stage ! ~\n\nHere is our kernel, the function that fast generates 5 times size public test set:\n\n<h3><strong><a href=\"https://www.kaggle.com/luyujia/fast-5-times-test-data-generator\">fast 5 times test_data generator</a></strong></h3>\n\nWe can insert it into prediction kernel and use it to test time limit and memory limit !\n\n<p>Hope it is helpful to you.\nDon't forget delete it in the end : ) </p>",
  "messages": [
    {
      "id": "539368",
      "postDate": "05/30/2019 01:49:43",
      "content": "In Data Description it said that <strong>The second-stage test set is approximately five times the size of the first</strong> , which means more strict time limit, disk limit and memory limit we have to face. Your final kernel for prediction may not exceed 1.8h ( just roughly divided the kernel time limit 9h by 5). These things really make us anxious, so why not test it before the second-stage ! ~\n\nHere is our kernel, the function that fast generates 5 times size public test set:\n\n<h3><strong><a href=\"https://www.kaggle.com/luyujia/fast-5-times-test-data-generator\">fast 5 times test_data generator</a></strong></h3>\n\nWe can insert it into prediction kernel and use it to test time limit and memory limit !\n\n<p>Hope it is helpful to you.\nDon't forget delete it in the end : ) </p>",
      "rawMarkdown": "##### In Data Description it said that **The second-stage test set is approximately five times the size of the first** , which means more strict time limit, disk limit and memory limit we have to face. Your final kernel for prediction may not exceed 1.8h ( just roughly divided the kernel time limit 9h by 5). These things really make us anxious, so why not test it before the second-stage ! ~\n\n##### Here is our kernel, the function that fast generates 5 times size public test set:\n\n### **[fast 5 times test_data generator](https://www.kaggle.com/luyujia/fast-5-times-test-data-generator)**\n\n\n\n##### We can insert it into prediction kernel and use it to test time limit and memory limit !\n\n\nHope it is helpful to you.\nDon't forget delete it in the end : )",
      "votes": null
    },
    {
      "id": "539587",
      "postDate": "05/30/2019 08:45:27",
      "content": "<p>Oh dear, why not just <code>test_df = pd.concat([test_df] * 5, axis=0)</code> ?</p>",
      "rawMarkdown": "Oh dear, why not just `test_df = pd.concat([test_df] * 5, axis=0)` ?",
      "votes": null
    },
    {
      "id": "539713",
      "postDate": "05/30/2019 12:36:36",
      "content": "<p>Well, to be on the safe side, we should generates 6 times size test set. </p>\n\n<p>Because “The second-stage test set is <strong>approximately five times</strong> the size of the first. ”  We don't know what the \"approximately\" means, may be 5.499 times.</p>",
      "rawMarkdown": "Well, to be on the safe side, we should generates 6 times size test set. \n\nBecause “The second-stage test set is **approximately five times** the size of the first. ”  We don't know what the \"approximately\" means, may be 5.499 times.",
      "votes": null
    },
    {
      "id": "539803",
      "postDate": "05/30/2019 14:43:08",
      "content": "<p>good advice!</p>",
      "rawMarkdown": "good advice!",
      "votes": null
    },
    {
      "id": "539813",
      "postDate": "05/30/2019 14:56:17",
      "content": "<p>In this kernel we duplicated both labels and images with unique name, which is more like a real one. yeah, the advice is great!  thank you for offering a simple way !  </p>",
      "rawMarkdown": "In this kernel we duplicated both labels and images with unique name, which is more like a real one. yeah, the advice is great!  thank you for offering a simple way !",
      "votes": null
    },
    {
      "id": "540166",
      "postDate": "05/31/2019 05:50:20",
      "content": "<p>Thanks for the reminder! If the selected kernels both fail in second-stage, does it mean the efforts are all in vain?</p>",
      "rawMarkdown": "Thanks for the reminder! If the selected kernels both fail in second-stage, does it mean the efforts are all in vain?",
      "votes": null
    },
    {
      "id": "540208",
      "postDate": "05/31/2019 06:39:11",
      "content": "<p>There is not official reply yet, but I think it most likely to be that~  better safe than sorry :) </p>",
      "rawMarkdown": "There is not official reply yet, but I think it most likely to be that~  better safe than sorry :)",
      "votes": null
    },
    {
      "id": "540422",
      "postDate": "05/31/2019 13:36:01",
      "content": "<p>Yes, if the selected kernels will fail on second stage, you will be just disqualified from the competition. It already happened to some kagglers earlier on similar competitions. So choose your submissions wisely.</p>",
      "rawMarkdown": "Yes, if the selected kernels will fail on second stage, you will be just disqualified from the competition. It already happened to some kagglers earlier on similar competitions. So choose your submissions wisely.",
      "votes": null
    },
    {
      "id": "541010",
      "postDate": "06/01/2019 14:43:52",
      "content": "<p>I think its a very important info. I missed it.. Thank you!</p>",
      "rawMarkdown": "I think its a very important info. I missed it.. Thank you!",
      "votes": null
    },
    {
      "id": "541013",
      "postDate": "06/01/2019 14:51:53",
      "content": "<p>thanks Lu</p>",
      "rawMarkdown": "thanks Lu",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 539587,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "05/30/2019 08:45:27",
      "content": "<p>Oh dear, why not just <code>test_df = pd.concat([test_df] * 5, axis=0)</code> ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 539813,
          "author_name": "luyujia",
          "author_url": "",
          "post_date": "05/30/2019 14:56:17",
          "content": "<p>In this kernel we duplicated both labels and images with unique name, which is more like a real one. yeah, the advice is great!  thank you for offering a simple way !  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 539713,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "05/30/2019 12:36:36",
      "content": "<p>Well, to be on the safe side, we should generates 6 times size test set. </p>\n\n<p>Because “The second-stage test set is <strong>approximately five times</strong> the size of the first. ”  We don't know what the \"approximately\" means, may be 5.499 times.</p>",
      "votes": null,
      "replies": [
        {
          "id": 539803,
          "author_name": "luyujia",
          "author_url": "",
          "post_date": "05/30/2019 14:43:08",
          "content": "<p>good advice!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 540166,
      "author_name": "vivaroma",
      "author_url": "",
      "post_date": "05/31/2019 05:50:20",
      "content": "<p>Thanks for the reminder! If the selected kernels both fail in second-stage, does it mean the efforts are all in vain?</p>",
      "votes": null,
      "replies": [
        {
          "id": 540208,
          "author_name": "luyujia",
          "author_url": "",
          "post_date": "05/31/2019 06:39:11",
          "content": "<p>There is not official reply yet, but I think it most likely to be that~  better safe than sorry :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540422,
          "author_name": "demonplus",
          "author_url": "",
          "post_date": "05/31/2019 13:36:01",
          "content": "<p>Yes, if the selected kernels will fail on second stage, you will be just disqualified from the competition. It already happened to some kagglers earlier on similar competitions. So choose your submissions wisely.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 541010,
      "author_name": "dhaqui",
      "author_url": "",
      "post_date": "06/01/2019 14:43:52",
      "content": "<p>I think its a very important info. I missed it.. Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 541013,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "06/01/2019 14:51:53",
      "content": "<p>thanks Lu</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "539368": "##### In Data Description it said that **The second-stage test set is approximately five times the size of the first** , which means more strict time limit, disk limit and memory limit we have to face. Your final kernel for prediction may not exceed 1.8h ( just roughly divided the kernel time limit 9h by 5). These things really make us anxious, so why not test it before the second-stage ! ~\n\n##### Here is our kernel, the function that fast generates 5 times size public test set:\n\n### **[fast 5 times test_data generator](https://www.kaggle.com/luyujia/fast-5-times-test-data-generator)**\n\n\n\n##### We can insert it into prediction kernel and use it to test time limit and memory limit !\n\n\nHope it is helpful to you.\nDon't forget delete it in the end : )",
    "539587": "Oh dear, why not just `test_df = pd.concat([test_df] * 5, axis=0)` ?",
    "539713": "Well, to be on the safe side, we should generates 6 times size test set. \n\nBecause “The second-stage test set is **approximately five times** the size of the first. ”  We don't know what the \"approximately\" means, may be 5.499 times.",
    "539803": "good advice!",
    "539813": "In this kernel we duplicated both labels and images with unique name, which is more like a real one. yeah, the advice is great!  thank you for offering a simple way !",
    "540166": "Thanks for the reminder! If the selected kernels both fail in second-stage, does it mean the efforts are all in vain?",
    "540208": "There is not official reply yet, but I think it most likely to be that~  better safe than sorry :)",
    "540422": "Yes, if the selected kernels will fail on second stage, you will be just disqualified from the competition. It already happened to some kagglers earlier on similar competitions. So choose your submissions wisely.",
    "541010": "I think its a very important info. I missed it.. Thank you!",
    "541013": "thanks Lu"
  },
  "source": "meta"
}