{
  "id": 35176,
  "title": "94th solution, with 4G RAM and GPU 740M .  Extract ResNet feature using Keras",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/writeups/kuhung-94th-solution-with-4g-ram-and-gpu-740m-extr",
  "author_name": "",
  "post_date": "2017-06-23T12:37:06.051630Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Not good enough, but I would like to share my solution.</p>\n\n<p>With intel i5, 4G RAM and 740M GPU, after several times memory error，I finished this comp and got my first kaggle medal.</p>\n\n<p>Method:  I used pretrianed ResNet50 to extract features and give it to xgboost, so thats my score :  0.95030.\n<a href=\"https://github.com/kuhung/Tips_for_Data_Scientist/blob/master/Feature_Engineering/Extract_ResNet_Feature_using_Keras.py\">Extract_ResNet_Feature_using_Keras.py</a> according to <a href=\"https://www.kaggle.com/kelexu/extract-resnet-feature-using-keras?scriptVersionId=1244781\">script from Planet</a> and <a href=\"https://github.com/fchollet/deep-learning-models\">fchollet/deep-learning-models</a></p>\n\n<p>Improvement：\nThe result can be improved by using Vgg19, but limited by my 4G RAM when merge the features. Dimensionality reduction can help 0.003 and Ensemble is a good way too. </p>\n\n<p>The next time you may try it in a CV competition:) </p>",
  "messages": [
    {
      "id": "195373",
      "postDate": "06/23/2017 12:37:06",
      "content": "<p>Not good enough, but I would like to share my solution.</p>\n\n<p>With intel i5, 4G RAM and 740M GPU, after several times memory error，I finished this comp and got my first kaggle medal.</p>\n\n<p>Method:  I used pretrianed ResNet50 to extract features and give it to xgboost, so thats my score :  0.95030.\n<a href=\"https://github.com/kuhung/Tips_for_Data_Scientist/blob/master/Feature_Engineering/Extract_ResNet_Feature_using_Keras.py\">Extract_ResNet_Feature_using_Keras.py</a> according to <a href=\"https://www.kaggle.com/kelexu/extract-resnet-feature-using-keras?scriptVersionId=1244781\">script from Planet</a> and <a href=\"https://github.com/fchollet/deep-learning-models\">fchollet/deep-learning-models</a></p>\n\n<p>Improvement：\nThe result can be improved by using Vgg19, but limited by my 4G RAM when merge the features. Dimensionality reduction can help 0.003 and Ensemble is a good way too. </p>\n\n<p>The next time you may try it in a CV competition:) </p>",
      "rawMarkdown": "Not good enough, but I would like to share my solution.\n\nWith intel i5, 4G RAM and 740M GPU, after several times memory error，I finished this comp and got my first kaggle medal.\n\nMethod:  I used pretrianed ResNet50 to extract features and give it to xgboost, so thats my score :  0.95030.\n[Extract_ResNet_Feature_using_Keras.py][1] according to [script from Planet][2] and [fchollet/deep-learning-models][3]\n\n\n  [1]: https://github.com/kuhung/Tips_for_Data_Scientist/blob/master/Feature_Engineering/Extract_ResNet_Feature_using_Keras.py\n  [2]: https://www.kaggle.com/kelexu/extract-resnet-feature-using-keras?scriptVersionId=1244781\n  [3]: https://github.com/fchollet/deep-learning-models\n\nImprovement：\nThe result can be improved by using Vgg19, but limited by my 4G RAM when merge the features. Dimensionality reduction can help 0.003 and Ensemble is a good way too. \n\nThe next time you may try it in a CV competition:)",
      "votes": null
    },
    {
      "id": "195382",
      "postDate": "06/23/2017 12:49:31",
      "content": "<p>Thanks for sharing! Did you just use all images raw or did you preprocess then in any way? Did you use only the train images or did you also use additional images? Is 0.95 your private leaderboard score or local CV.</p>\n\n<p>As I wrote in another post, I used these kind of models as the first level in a stacked system. My best (based on local CV) first level model was actually from from vgg16 features trained with ExtraTrees form sklearn. However my local CV had the famous leak...</p>",
      "rawMarkdown": "Thanks for sharing! Did you just use all images raw or did you preprocess then in any way? Did you use only the train images or did you also use additional images? Is 0.95 your private leaderboard score or local CV.\n\nAs I wrote in another post, I used these kind of models as the first level in a stacked system. My best (based on local CV) first level model was actually from from vgg16 features trained with ExtraTrees form sklearn. However my local CV had the famous leak...",
      "votes": null
    },
    {
      "id": "195390",
      "postDate": "06/23/2017 13:37:59",
      "content": "<p>Hi, Johansen:</p>\n\n<p>I used all the data, include the additional because I find the more data,the more stable.  Then resized them to 224x224 as the model needed. 0.95 is my private LB.</p>\n\n<p>Here is my score records(Not Accurate)：</p>\n\n<p>PL： 0.89xx (train only)   0.83xx (add additioin) </p>\n\n<p>By the way, your sharing is really instructive, I learned a lot from it.</p>",
      "rawMarkdown": "Hi, Johansen:\n\nI used all the data, include the additional because I find the more data,the more stable.  Then resized them to 224x224 as the model needed. 0.95 is my private LB.\n\nHere is my score records(Not Accurate)：\n\nPL： 0.89xx (train only)   0.83xx (add additioin) \n\nBy the way, your sharing is really instructive, I learned a lot from it.",
      "votes": null
    },
    {
      "id": "195402",
      "postDate": "06/23/2017 14:04:05",
      "content": "<p>I just read your post again and feel confused about one point:</p>\n\n<p>I find you used several base models like KNN and RF.However, in my test, RF overfitted seriously and changed  a lot when I set different seed in split. That means the model have much confidence on his predict. In the condition, If 'luckily' split one patient in 5, is it the possible reason lead to leak on CV, for I don't notice you changed the seed to split or test it on a small data?</p>",
      "rawMarkdown": "I just read your post again and feel confused about one point:\n\nI find you used several base models like KNN and RF.However, in my test, RF overfitted seriously and changed  a lot when I set different seed in split. That means the model have much confidence on his predict. In the condition, If 'luckily' split one patient in 5, is it the possible reason lead to leak on CV, for I don't notice you changed the seed to split or test it on a small data?",
      "votes": null
    },
    {
      "id": "195419",
      "postDate": "06/23/2017 14:44:12",
      "content": "<p>I think you've hit the same problem as me. Cross fold leaking. The problem is that images from the same patient is appearing in more than one fold. If you have a lucky shuffle with a seed, all images of the same cervix are all put in the same fold. The smart guys, like Darius, Russ W, Gilberto and the other GM's understood this while it still was time to sort it out. I didn't see this until it was to late.</p>\n\n<p>When comparing models with each other using K-folded CV, you should use the same seed for creating the folds. How do the models compare if you use same seed?  </p>",
      "rawMarkdown": "I think you've hit the same problem as me. Cross fold leaking. The problem is that images from the same patient is appearing in more than one fold. If you have a lucky shuffle with a seed, all images of the same cervix are all put in the same fold. The smart guys, like Darius, Russ W, Gilberto and the other GM's understood this while it still was time to sort it out. I didn't see this until it was to late.\n\nWhen comparing models with each other using K-folded CV, you should use the same seed for creating the folds. How do the models compare if you use same seed?",
      "votes": null
    },
    {
      "id": "195423",
      "postDate": "06/23/2017 14:58:23",
      "content": "<p>Yes, you are right, Same seed for model campare, I agree with you. After this leaking, the next time we will perform better, won't we? :P</p>",
      "rawMarkdown": "Yes, you are right, Same seed for model campare, I agree with you. After this leaking, the next time we will perform better, won't we? :P",
      "votes": null
    },
    {
      "id": "195576",
      "postDate": "06/24/2017 03:18:08",
      "content": "<p>Thanks for sharing!\nWhat's \"train_label.csv\" and what's \"test_stg1_label.csv\". I don't see that in the Data.</p>",
      "rawMarkdown": "Thanks for sharing!\nWhat's \"train_label.csv\" and what's \"test_stg1_label.csv\". I don't see that in the Data.",
      "votes": null
    },
    {
      "id": "195578",
      "postDate": "06/24/2017 03:34:36",
      "content": "<p>Sorry, I will attach it here. </p>\n\n<pre><code>import pandas as pd\nimport os\nfrom tqdm import *\n\nID = []\nType = []\nfor count in range (1,4):\n    path = 'train/Type_%s'%count\n    for file in tqdm(os.listdir(path)):\n        if 'DS' not in file:\n            ID.append(file)\n            Type.append('Type_%s'%count)\n    train_label = pd.DataFrame({'id':ID,'label':Type})\ntrain_label.to_csv('train_label.csv',index=False)\n\nstg1 = pd.read_csv('solution_stg1_release.csv')\nstg1['Type_2'] = stg1['Type_2'].apply(lambda x: x*2)\nstg1['Type_3'] = stg1['Type_3'].apply(lambda x: x*3)\nstg1['label'] = stg1['Type_1']+stg1['Type_2']+stg1['Type_3']\n\ndef label_convert(x):\n    if x==1:\n        return 'Type_1'\n    if x==2:\n        return 'Type_2'\n    if x==3:\n        return 'Type_3'\n\nstg1['label'] = stg1['label'].apply(lambda x: label_convert(x) )\nstg1.drop(['Type_1','Type_2','Type_3'],axis=1,inplace=True)\nstg1.columns = ['id','label']\nstg1.to_csv('test_stg1_label.csv',index=False)\n</code></pre>",
      "rawMarkdown": "Sorry, I will attach it here. \n\n    import pandas as pd\n    import os\n    from tqdm import *\n\n    ID = []\n    Type = []\n    for count in range (1,4):\n        path = 'train/Type_%s'%count\n        for file in tqdm(os.listdir(path)):\n            if 'DS' not in file:\n                ID.append(file)\n                Type.append('Type_%s'%count)\n        train_label = pd.DataFrame({'id':ID,'label':Type})\n    train_label.to_csv('train_label.csv',index=False)\n\n    stg1 = pd.read_csv('solution_stg1_release.csv')\n    stg1['Type_2'] = stg1['Type_2'].apply(lambda x: x*2)\n    stg1['Type_3'] = stg1['Type_3'].apply(lambda x: x*3)\n    stg1['label'] = stg1['Type_1']+stg1['Type_2']+stg1['Type_3']\n\n    def label_convert(x):\n        if x==1:\n            return 'Type_1'\n        if x==2:\n            return 'Type_2'\n        if x==3:\n            return 'Type_3'\n\n    stg1['label'] = stg1['label'].apply(lambda x: label_convert(x) )\n    stg1.drop(['Type_1','Type_2','Type_3'],axis=1,inplace=True)\n    stg1.columns = ['id','label']\n    stg1.to_csv('test_stg1_label.csv',index=False)",
      "votes": null
    },
    {
      "id": "195579",
      "postDate": "06/24/2017 04:06:29",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 195382,
      "author_name": "oysteijo",
      "author_url": "",
      "post_date": "06/23/2017 12:49:31",
      "content": "<p>Thanks for sharing! Did you just use all images raw or did you preprocess then in any way? Did you use only the train images or did you also use additional images? Is 0.95 your private leaderboard score or local CV.</p>\n\n<p>As I wrote in another post, I used these kind of models as the first level in a stacked system. My best (based on local CV) first level model was actually from from vgg16 features trained with ExtraTrees form sklearn. However my local CV had the famous leak...</p>",
      "votes": null,
      "replies": [
        {
          "id": 195390,
          "author_name": "badoun",
          "author_url": "",
          "post_date": "06/23/2017 13:37:59",
          "content": "<p>Hi, Johansen:</p>\n\n<p>I used all the data, include the additional because I find the more data,the more stable.  Then resized them to 224x224 as the model needed. 0.95 is my private LB.</p>\n\n<p>Here is my score records(Not Accurate)：</p>\n\n<p>PL： 0.89xx (train only)   0.83xx (add additioin) </p>\n\n<p>By the way, your sharing is really instructive, I learned a lot from it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195402,
          "author_name": "badoun",
          "author_url": "",
          "post_date": "06/23/2017 14:04:05",
          "content": "<p>I just read your post again and feel confused about one point:</p>\n\n<p>I find you used several base models like KNN and RF.However, in my test, RF overfitted seriously and changed  a lot when I set different seed in split. That means the model have much confidence on his predict. In the condition, If 'luckily' split one patient in 5, is it the possible reason lead to leak on CV, for I don't notice you changed the seed to split or test it on a small data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195419,
          "author_name": "oysteijo",
          "author_url": "",
          "post_date": "06/23/2017 14:44:12",
          "content": "<p>I think you've hit the same problem as me. Cross fold leaking. The problem is that images from the same patient is appearing in more than one fold. If you have a lucky shuffle with a seed, all images of the same cervix are all put in the same fold. The smart guys, like Darius, Russ W, Gilberto and the other GM's understood this while it still was time to sort it out. I didn't see this until it was to late.</p>\n\n<p>When comparing models with each other using K-folded CV, you should use the same seed for creating the folds. How do the models compare if you use same seed?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195423,
          "author_name": "badoun",
          "author_url": "",
          "post_date": "06/23/2017 14:58:23",
          "content": "<p>Yes, you are right, Same seed for model campare, I agree with you. After this leaking, the next time we will perform better, won't we? :P</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 195576,
      "author_name": "carlosaguayo",
      "author_url": "",
      "post_date": "06/24/2017 03:18:08",
      "content": "<p>Thanks for sharing!\nWhat's \"train_label.csv\" and what's \"test_stg1_label.csv\". I don't see that in the Data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 195578,
          "author_name": "badoun",
          "author_url": "",
          "post_date": "06/24/2017 03:34:36",
          "content": "<p>Sorry, I will attach it here. </p>\n\n<pre><code>import pandas as pd\nimport os\nfrom tqdm import *\n\nID = []\nType = []\nfor count in range (1,4):\n    path = 'train/Type_%s'%count\n    for file in tqdm(os.listdir(path)):\n        if 'DS' not in file:\n            ID.append(file)\n            Type.append('Type_%s'%count)\n    train_label = pd.DataFrame({'id':ID,'label':Type})\ntrain_label.to_csv('train_label.csv',index=False)\n\nstg1 = pd.read_csv('solution_stg1_release.csv')\nstg1['Type_2'] = stg1['Type_2'].apply(lambda x: x*2)\nstg1['Type_3'] = stg1['Type_3'].apply(lambda x: x*3)\nstg1['label'] = stg1['Type_1']+stg1['Type_2']+stg1['Type_3']\n\ndef label_convert(x):\n    if x==1:\n        return 'Type_1'\n    if x==2:\n        return 'Type_2'\n    if x==3:\n        return 'Type_3'\n\nstg1['label'] = stg1['label'].apply(lambda x: label_convert(x) )\nstg1.drop(['Type_1','Type_2','Type_3'],axis=1,inplace=True)\nstg1.columns = ['id','label']\nstg1.to_csv('test_stg1_label.csv',index=False)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195579,
          "author_name": "carlosaguayo",
          "author_url": "",
          "post_date": "06/24/2017 04:06:29",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "195373": "Not good enough, but I would like to share my solution.\n\nWith intel i5, 4G RAM and 740M GPU, after several times memory error，I finished this comp and got my first kaggle medal.\n\nMethod:  I used pretrianed ResNet50 to extract features and give it to xgboost, so thats my score :  0.95030.\n[Extract_ResNet_Feature_using_Keras.py][1] according to [script from Planet][2] and [fchollet/deep-learning-models][3]\n\n\n  [1]: https://github.com/kuhung/Tips_for_Data_Scientist/blob/master/Feature_Engineering/Extract_ResNet_Feature_using_Keras.py\n  [2]: https://www.kaggle.com/kelexu/extract-resnet-feature-using-keras?scriptVersionId=1244781\n  [3]: https://github.com/fchollet/deep-learning-models\n\nImprovement：\nThe result can be improved by using Vgg19, but limited by my 4G RAM when merge the features. Dimensionality reduction can help 0.003 and Ensemble is a good way too. \n\nThe next time you may try it in a CV competition:)",
    "195382": "Thanks for sharing! Did you just use all images raw or did you preprocess then in any way? Did you use only the train images or did you also use additional images? Is 0.95 your private leaderboard score or local CV.\n\nAs I wrote in another post, I used these kind of models as the first level in a stacked system. My best (based on local CV) first level model was actually from from vgg16 features trained with ExtraTrees form sklearn. However my local CV had the famous leak...",
    "195390": "Hi, Johansen:\n\nI used all the data, include the additional because I find the more data,the more stable.  Then resized them to 224x224 as the model needed. 0.95 is my private LB.\n\nHere is my score records(Not Accurate)：\n\nPL： 0.89xx (train only)   0.83xx (add additioin) \n\nBy the way, your sharing is really instructive, I learned a lot from it.",
    "195402": "I just read your post again and feel confused about one point:\n\nI find you used several base models like KNN and RF.However, in my test, RF overfitted seriously and changed  a lot when I set different seed in split. That means the model have much confidence on his predict. In the condition, If 'luckily' split one patient in 5, is it the possible reason lead to leak on CV, for I don't notice you changed the seed to split or test it on a small data?",
    "195419": "I think you've hit the same problem as me. Cross fold leaking. The problem is that images from the same patient is appearing in more than one fold. If you have a lucky shuffle with a seed, all images of the same cervix are all put in the same fold. The smart guys, like Darius, Russ W, Gilberto and the other GM's understood this while it still was time to sort it out. I didn't see this until it was to late.\n\nWhen comparing models with each other using K-folded CV, you should use the same seed for creating the folds. How do the models compare if you use same seed?",
    "195423": "Yes, you are right, Same seed for model campare, I agree with you. After this leaking, the next time we will perform better, won't we? :P",
    "195576": "Thanks for sharing!\nWhat's \"train_label.csv\" and what's \"test_stg1_label.csv\". I don't see that in the Data.",
    "195578": "Sorry, I will attach it here. \n\n    import pandas as pd\n    import os\n    from tqdm import *\n\n    ID = []\n    Type = []\n    for count in range (1,4):\n        path = 'train/Type_%s'%count\n        for file in tqdm(os.listdir(path)):\n            if 'DS' not in file:\n                ID.append(file)\n                Type.append('Type_%s'%count)\n        train_label = pd.DataFrame({'id':ID,'label':Type})\n    train_label.to_csv('train_label.csv',index=False)\n\n    stg1 = pd.read_csv('solution_stg1_release.csv')\n    stg1['Type_2'] = stg1['Type_2'].apply(lambda x: x*2)\n    stg1['Type_3'] = stg1['Type_3'].apply(lambda x: x*3)\n    stg1['label'] = stg1['Type_1']+stg1['Type_2']+stg1['Type_3']\n\n    def label_convert(x):\n        if x==1:\n            return 'Type_1'\n        if x==2:\n            return 'Type_2'\n        if x==3:\n            return 'Type_3'\n\n    stg1['label'] = stg1['label'].apply(lambda x: label_convert(x) )\n    stg1.drop(['Type_1','Type_2','Type_3'],axis=1,inplace=True)\n    stg1.columns = ['id','label']\n    stg1.to_csv('test_stg1_label.csv',index=False)",
    "195579": "Thanks!"
  },
  "source": "meta"
}