{
  "id": 87675,
  "title": "A few notes for Kernel Competitions",
  "url": "/competitions/imet-2019-fgvc6/discussion/87675",
  "author_name": "cab",
  "post_date": "2019-04-02T13:25:28.057000",
  "votes": 43,
  "comment_count": 13,
  "views": 0,
  "content": "<p>It seems that kernel competition is going to be more popular in the future. There are some lesson learned I got from previous competitions (<a href=\"https://www.kaggle.com/c/petfinder-adoption-prediction\">PetFinder.my Adoption Prediction</a> , <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\">Quora Insincere Questions Classification</a> ). Here are a few notes: </p>\n\n<h1>1. Make sure your kernel can be re-procedured.</h1>\n\n<p>This is the most headache thing happened to me during the last two competitions. My score is unstable and can be changed everyday when I re-run it without any changes. To reduce it, you can set fixed <code>seed</code> at the top of kernel. </p>\n\n<p><code>\ndef seed_everything(seed=1234):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    tf.set_random_seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n</code> \nThanks @bminixhofer </p>\n\n<h1>2. Notebooks or scripts?</h1>\n\n<p>Personally, I prefer script for final submission. <br>\nNotebook is better at early stage since you need to brainstorming, test ideas, fix bugs, ... As long as competition run, it becomes longer and quite hard to maintain, control the modules. \nLast competition, I used notebook and felt crazy when adding more features. I got the bugs and unseen them for a long time. I saw the disadvantages and decided to port into script style. This work made me recognize errors, bugs inside. Finally, It took me one week to make sure the new script one could be reprocedured. \nSo, I recommend to use script as you can. </p>\n\n<p><img src=\"https://i.imgur.com/P0sWqgu.png\" alt=\"not_use_jupyter\"></p>\n\n<p>Source: <a href=\"https://www.youtube.com/watch?v=oxikDxBzKUU\">Kaggle Days Paris - „How to win competitions, spending the minimum amount of time”</a> </p>\n\n<h1>3. Local development</h1>\n\n<p>You can make your code (running, testing, check, ...) on your local computer then push code to run on Kaggle latter. You might need to fork <a href=\"https://github.com/Kaggle/docker-python\">Kaggle docker</a> and develope on this environment to avoid confliction of version and make your local work to be reliable. </p>\n\n<h1>4. Measure running time of stage2</h1>\n\n<p>Since test data of stage2 is different from the number of test cases, you should measure run time of <code>inference/prediction</code> and estimate how long will it take for stage2. In this competition, \n<code>The second-stage test set is approximately five times the size of the first</code> , you can upsampling the test set five times and measure the total time. </p>\n\n<h1>5. Your comments/experiments.</h1>\n\n<p>GLHF, </p>",
  "messages": [
    {
      "id": 505730,
      "postDate": "2019-04-02T13:25:28.057Z",
      "content": "<p>It seems that kernel competition is going to be more popular in the future. There are some lesson learned I got from previous competitions (<a href=\"https://www.kaggle.com/c/petfinder-adoption-prediction\">PetFinder.my Adoption Prediction</a> , <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\">Quora Insincere Questions Classification</a> ). Here are a few notes: </p>\n\n<h1>1. Make sure your kernel can be re-procedured.</h1>\n\n<p>This is the most headache thing happened to me during the last two competitions. My score is unstable and can be changed everyday when I re-run it without any changes. To reduce it, you can set fixed <code>seed</code> at the top of kernel. </p>\n\n<p><code>\ndef seed_everything(seed=1234):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    tf.set_random_seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n</code> \nThanks @bminixhofer </p>\n\n<h1>2. Notebooks or scripts?</h1>\n\n<p>Personally, I prefer script for final submission. <br>\nNotebook is better at early stage since you need to brainstorming, test ideas, fix bugs, ... As long as competition run, it becomes longer and quite hard to maintain, control the modules. \nLast competition, I used notebook and felt crazy when adding more features. I got the bugs and unseen them for a long time. I saw the disadvantages and decided to port into script style. This work made me recognize errors, bugs inside. Finally, It took me one week to make sure the new script one could be reprocedured. \nSo, I recommend to use script as you can. </p>\n\n<p><img src=\"https://i.imgur.com/P0sWqgu.png\" alt=\"not_use_jupyter\"></p>\n\n<p>Source: <a href=\"https://www.youtube.com/watch?v=oxikDxBzKUU\">Kaggle Days Paris - „How to win competitions, spending the minimum amount of time”</a> </p>\n\n<h1>3. Local development</h1>\n\n<p>You can make your code (running, testing, check, ...) on your local computer then push code to run on Kaggle latter. You might need to fork <a href=\"https://github.com/Kaggle/docker-python\">Kaggle docker</a> and develope on this environment to avoid confliction of version and make your local work to be reliable. </p>\n\n<h1>4. Measure running time of stage2</h1>\n\n<p>Since test data of stage2 is different from the number of test cases, you should measure run time of <code>inference/prediction</code> and estimate how long will it take for stage2. In this competition, \n<code>The second-stage test set is approximately five times the size of the first</code> , you can upsampling the test set five times and measure the total time. </p>\n\n<h1>5. Your comments/experiments.</h1>\n\n<p>GLHF, </p>",
      "rawMarkdown": "It seems that kernel competition is going to be more popular in the future. There are some lesson learned I got from previous competitions ([PetFinder.my Adoption Prediction](https://www.kaggle.com/c/petfinder-adoption-prediction) , [Quora Insincere Questions Classification](https://www.kaggle.com/c/quora-insincere-questions-classification) ). Here are a few notes: \n\n#1. Make sure your kernel can be re-procedured. \n\nThis is the most headache thing happened to me during the last two competitions. My score is unstable and can be changed everyday when I re-run it without any changes. To reduce it, you can set fixed `seed` at the top of kernel. \n\n``` \ndef seed_everything(seed=1234):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    tf.set_random_seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n``` \nThanks @bminixhofer \n\n#2. Notebooks or scripts? \nPersonally, I prefer script for final submission.  \nNotebook is better at early stage since you need to brainstorming, test ideas, fix bugs, ... As long as competition run, it becomes longer and quite hard to maintain, control the modules. \nLast competition, I used notebook and felt crazy when adding more features. I got the bugs and unseen them for a long time. I saw the disadvantages and decided to port into script style. This work made me recognize errors, bugs inside. Finally, It took me one week to make sure the new script one could be reprocedured. \nSo, I recommend to use script as you can. \n\n![not_use_jupyter](https://i.imgur.com/P0sWqgu.png)\n\nSource: [Kaggle Days Paris - „How to win competitions, spending the minimum amount of time”](https://www.youtube.com/watch?v=oxikDxBzKUU) \n\n#3. Local development \nYou can make your code (running, testing, check, ...) on your local computer then push code to run on Kaggle latter. You might need to fork [Kaggle docker](https://github.com/Kaggle/docker-python) and develope on this environment to avoid confliction of version and make your local work to be reliable. \n\n#4. Measure running time of stage2 \nSince test data of stage2 is different from the number of test cases, you should measure run time of `inference/prediction` and estimate how long will it take for stage2. In this competition, \n`The second-stage test set is approximately five times the size of the first` , you can upsampling the test set five times and measure the total time. \n\n#5. Your comments/experiments. \n\n\n\nGLHF, ",
      "votes": 43
    },
    {
      "id": 505998,
      "postDate": "2019-04-02T21:11:53.833Z",
      "content": "<p>Solid advice on using scripts for submissions. I also use a build script to convert a python package into one script, so that you can organize you code nicely for local development, instead of putting everything into one file: <a href=\"https://github.com/lopuhin/kaggle-script-template\">https://github.com/lopuhin/kaggle-script-template</a></p>",
      "rawMarkdown": "Solid advice on using scripts for submissions. I also use a build script to convert a python package into one script, so that you can organize you code nicely for local development, instead of putting everything into one file: https://github.com/lopuhin/kaggle-script-template",
      "votes": 13,
      "replies": [
        {
          "id": 506031,
          "postDate": "2019-04-02T22:35:04.050Z",
          "content": "<p>Brilliant idea!. \nThanks for sharing. </p>",
          "rawMarkdown": "Brilliant idea!. \nThanks for sharing. "
        },
        {
          "id": 510818,
          "postDate": "2019-04-09T13:53:40.990Z",
          "content": "<p>This is amazing :)</p>",
          "rawMarkdown": "This is amazing :)"
        }
      ]
    },
    {
      "id": 509052,
      "postDate": "2019-04-07T09:00:02.880Z",
      "content": "<p>Also be sure that you won't get out of memory error with bigger test set in case you keep images in ram.</p>",
      "rawMarkdown": "Also be sure that you won't get out of memory error with bigger test set in case you keep images in ram.",
      "votes": 3
    },
    {
      "id": 508673,
      "postDate": "2019-04-06T15:17:50.800Z",
      "content": "<p>Oh, I have a problem, do different seeds cause different results ?</p>",
      "rawMarkdown": "Oh, I have a problem, do different seeds cause different results ?",
      "votes": 1,
      "replies": [
        {
          "id": 509056,
          "postDate": "2019-04-07T09:04:29.557Z",
          "content": "<p>Yes. It may cause different results.</p>",
          "rawMarkdown": "Yes. It may cause different results."
        }
      ]
    },
    {
      "id": 506744,
      "postDate": "2019-04-03T20:21:15.617Z",
      "content": "<p>I like notebooks just for debugging purposes. But as you said, as your code grows it's hard to manage cell by cell, and the scripts become the most viable options. Nice text.</p>",
      "rawMarkdown": "I like notebooks just for debugging purposes. But as you said, as your code grows it's hard to manage cell by cell, and the scripts become the most viable options. Nice text.",
      "votes": 1
    },
    {
      "id": 505796,
      "postDate": "2019-04-02T14:43:58.963Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 516375,
      "postDate": "2019-04-14T04:21:23.193Z",
      "content": "<p>I want to know how to avoid the endless growth of memory usage.</p>",
      "rawMarkdown": "I want to know how to avoid the endless growth of memory usage."
    },
    {
      "id": 506631,
      "postDate": "2019-04-03T17:49:30.880Z",
      "content": "<p>what</p>",
      "rawMarkdown": "what"
    },
    {
      "id": 506612,
      "postDate": "2019-04-03T17:28:41.747Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 505830,
      "postDate": "2019-04-02T15:39:45.090Z",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks",
      "votes": 1
    },
    {
      "id": 515157,
      "postDate": "2019-04-12T08:54:35.910Z",
      "content": "<p>thanks\ncan i use caffe in kernel?</p>",
      "rawMarkdown": "thanks\ncan i use caffe in kernel?"
    }
  ],
  "comments": [
    {
      "id": 505998,
      "author_name": "Konstantin Lopukhin",
      "author_url": "",
      "post_date": "2019-04-02T21:11:53.833000",
      "content": "<p>Solid advice on using scripts for submissions. I also use a build script to convert a python package into one script, so that you can organize you code nicely for local development, instead of putting everything into one file: <a href=\"https://github.com/lopuhin/kaggle-script-template\">https://github.com/lopuhin/kaggle-script-template</a></p>",
      "votes": 13,
      "replies": [
        {
          "id": 506031,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-04-02T22:35:04.050000",
          "content": "<p>Brilliant idea!. \nThanks for sharing. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 510818,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2019-04-09T13:53:40.990000",
          "content": "<p>This is amazing :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 509052,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-04-07T09:00:02.880000",
      "content": "<p>Also be sure that you won't get out of memory error with bigger test set in case you keep images in ram.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 508673,
      "author_name": "xutao",
      "author_url": "",
      "post_date": "2019-04-06T15:17:50.800000",
      "content": "<p>Oh, I have a problem, do different seeds cause different results ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 509056,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-04-07T09:04:29.557000",
          "content": "<p>Yes. It may cause different results.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 506744,
      "author_name": "Ayrton Willian Casella",
      "author_url": "",
      "post_date": "2019-04-03T20:21:15.617000",
      "content": "<p>I like notebooks just for debugging purposes. But as you said, as your code grows it's hard to manage cell by cell, and the scripts become the most viable options. Nice text.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 505796,
      "author_name": "Chenyang Zhang",
      "author_url": "",
      "post_date": "2019-04-02T14:43:58.963000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 516375,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-14T04:21:23.193000",
      "content": "<p>I want to know how to avoid the endless growth of memory usage.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 506631,
      "author_name": "MUKUL  S ANAND",
      "author_url": "",
      "post_date": "2019-04-03T17:49:30.880000",
      "content": "<p>what</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 506612,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-03T17:28:41.747000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 505830,
      "author_name": "汪元洪(Wang Yuanhong)",
      "author_url": "",
      "post_date": "2019-04-02T15:39:45.090000",
      "content": "<p>thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 515157,
      "author_name": "LuteLuthern",
      "author_url": "",
      "post_date": "2019-04-12T08:54:35.910000",
      "content": "<p>thanks\ncan i use caffe in kernel?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "505730": "It seems that kernel competition is going to be more popular in the future. There are some lesson learned I got from previous competitions ([PetFinder.my Adoption Prediction](https://www.kaggle.com/c/petfinder-adoption-prediction) , [Quora Insincere Questions Classification](https://www.kaggle.com/c/quora-insincere-questions-classification) ). Here are a few notes: \n\n#1. Make sure your kernel can be re-procedured. \n\nThis is the most headache thing happened to me during the last two competitions. My score is unstable and can be changed everyday when I re-run it without any changes. To reduce it, you can set fixed `seed` at the top of kernel. \n\n``` \ndef seed_everything(seed=1234):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    tf.set_random_seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n``` \nThanks @bminixhofer \n\n#2. Notebooks or scripts? \nPersonally, I prefer script for final submission.  \nNotebook is better at early stage since you need to brainstorming, test ideas, fix bugs, ... As long as competition run, it becomes longer and quite hard to maintain, control the modules. \nLast competition, I used notebook and felt crazy when adding more features. I got the bugs and unseen them for a long time. I saw the disadvantages and decided to port into script style. This work made me recognize errors, bugs inside. Finally, It took me one week to make sure the new script one could be reprocedured. \nSo, I recommend to use script as you can. \n\n![not_use_jupyter](https://i.imgur.com/P0sWqgu.png)\n\nSource: [Kaggle Days Paris - „How to win competitions, spending the minimum amount of time”](https://www.youtube.com/watch?v=oxikDxBzKUU) \n\n#3. Local development \nYou can make your code (running, testing, check, ...) on your local computer then push code to run on Kaggle latter. You might need to fork [Kaggle docker](https://github.com/Kaggle/docker-python) and develope on this environment to avoid confliction of version and make your local work to be reliable. \n\n#4. Measure running time of stage2 \nSince test data of stage2 is different from the number of test cases, you should measure run time of `inference/prediction` and estimate how long will it take for stage2. In this competition, \n`The second-stage test set is approximately five times the size of the first` , you can upsampling the test set five times and measure the total time. \n\n#5. Your comments/experiments. \n\n\n\nGLHF, ",
    "505998": "Solid advice on using scripts for submissions. I also use a build script to convert a python package into one script, so that you can organize you code nicely for local development, instead of putting everything into one file: https://github.com/lopuhin/kaggle-script-template",
    "509052": "Also be sure that you won't get out of memory error with bigger test set in case you keep images in ram.",
    "508673": "Oh, I have a problem, do different seeds cause different results ?",
    "506744": "I like notebooks just for debugging purposes. But as you said, as your code grows it's hard to manage cell by cell, and the scripts become the most viable options. Nice text.",
    "505796": "Thanks for sharing!",
    "516375": "I want to know how to avoid the endless growth of memory usage.",
    "506631": "what",
    "506612": "",
    "505830": "thanks",
    "515157": "thanks\ncan i use caffe in kernel?"
  }
}