{
  "id": 20336,
  "title": "What is the best score using pre-trained model?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20336",
  "author_name": "",
  "post_date": "2016-04-22T13:52:30.167Z",
  "votes": 2,
  "comment_count": 19,
  "views": 4031,
  "content": "<p>Just be curious about the best score when using pre-trained model. </p>",
  "messages": [
    {
      "id": "116200",
      "postDate": "04/22/2016 13:52:30",
      "content": "<p>Just be curious about the best score when using pre-trained model. </p>",
      "rawMarkdown": "Just be curious about the best score when using pre-trained model.",
      "votes": null
    },
    {
      "id": "116204",
      "postDate": "04/22/2016 14:04:18",
      "content": "<p>I think most of the top-10 scores are being obtained using pre-trained models.</p>",
      "rawMarkdown": "I think most of the top-10 scores are being obtained using pre-trained models.",
      "votes": null
    },
    {
      "id": "116411",
      "postDate": "04/24/2016 01:13:09",
      "content": "<p>I have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..</p>",
      "rawMarkdown": "I have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..",
      "votes": null
    },
    {
      "id": "116440",
      "postDate": "04/24/2016 09:04:15",
      "content": "<p>[quote=woshialex;116411]</p>\n\n<p>I have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..</p>\n\n<p>[/quote]</p>\n\n<p>What's your accuracy woshialex? Maybe you can boost the logloss tuning the model's confidence</p>",
      "rawMarkdown": "[quote=woshialex;116411]\r\n\r\nI have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..\r\n\r\n[/quote]\r\n\r\nWhat's your accuracy woshialex? Maybe you can boost the logloss tuning the model's confidence",
      "votes": null
    },
    {
      "id": "116467",
      "postDate": "04/24/2016 16:53:32",
      "content": "<p>[quote=woshialex;116411]</p>\n\n<p>... the best LB score I can get at the moment is only about 0.46..</p>\n\n<p>[/quote]\nThis is with just a CNN running on the whole image? Care to share some details of your model and parameters? </p>",
      "rawMarkdown": "[quote=woshialex;116411]\r\n\r\n... the best LB score I can get at the moment is only about 0.46..\r\n\r\n[/quote]\r\nThis is with just a CNN running on the whole image? Care to share some details of your model and parameters?",
      "votes": null
    },
    {
      "id": "116513",
      "postDate": "04/24/2016 22:35:04",
      "content": "<p>[quote]</p>\n\n<p>This is with just a CNN running on the whole image? Care to share some details of your model and parameters? </p>\n\n<p>[/quote]</p>\n\n<p>I trained Torch based Resnet-50 &quot;as is&quot; on competition training images as they are structured very similar to Imagenet and got 0.59924 LB score. Tried the same with Resnet-101 and 152 and got worst results.</p>\n\n<p>PS. I am using single GeForce 980 Ti GPU for training so my batch size floats from 10 to 30 depending on network size.</p>\n\n<p>One of the drawbacks of Torch version of Resnet - it crops image from the center and, in many cases, trained network cannot distinguish between driving (category C0) and operating the radio (category C5).</p>",
      "rawMarkdown": "[quote]\r\n\r\n\r\nThis is with just a CNN running on the whole image? Care to share some details of your model and parameters? \r\n\r\n[/quote]\r\n\r\nI trained Torch based Resnet-50 \"as is\" on competition training images as they are structured very similar to Imagenet and got 0.59924 LB score. Tried the same with Resnet-101 and 152 and got worst results.\r\n\r\nPS. I am using single GeForce 980 Ti GPU for training so my batch size floats from 10 to 30 depending on network size.\r\n\r\nOne of the drawbacks of Torch version of Resnet - it crops image from the center and, in many cases, trained network cannot distinguish between driving (category C0) and operating the radio (category C5).",
      "votes": null
    },
    {
      "id": "116608",
      "postDate": "04/25/2016 11:48:38",
      "content": "<p>can you share the confusion matrix?</p>",
      "rawMarkdown": "can you share the confusion matrix?",
      "votes": null
    },
    {
      "id": "116637",
      "postDate": "04/25/2016 14:14:25",
      "content": "<p>Can anyone clarify what do we mean by <strong>PRE-Trained model</strong>? </p>\n\n<p>1) You use already trained CNN model, which were trained on other pictures not related to this problem</p>\n\n<p>2) You train model on pictures from this problem using CNN structure from other method</p>",
      "rawMarkdown": "Can anyone clarify what do we mean by **PRE-Trained model**? \r\n\r\n1) You use already trained CNN model, which were trained on other pictures not related to this problem\r\n\r\n2) You train model on pictures from this problem using CNN structure from other method",
      "votes": null
    },
    {
      "id": "116644",
      "postDate": "04/25/2016 14:44:44",
      "content": "<p>[quote=ZFTurbo;116637]</p>\n\n<p>Can anyone clarify what do we mean by <strong>PRE-Trained model</strong>? </p>\n\n<p>1) You use already trained CNN model, which were trained on other pictures not related to this problem</p>\n\n<p>2) You train model on pictures from this problem using CNN structure from other method</p>\n\n<p>[/quote]</p>\n\n<p>1) Yes, I downloaded Resnet-200 model/network trained on Imagenet data.\n2) Yes, I am fine-tuning/retraining downloaded pretrained network for only 10 classes instead of original 1000 classes for Imagenet. Basically, last layer of 1000 neurons got removed and replaced with 10 neurons layer.</p>",
      "rawMarkdown": "[quote=ZFTurbo;116637]\r\n\r\nCan anyone clarify what do we mean by **PRE-Trained model**? \r\n\r\n1) You use already trained CNN model, which were trained on other pictures not related to this problem\r\n\r\n2) You train model on pictures from this problem using CNN structure from other method\r\n\r\n\r\n[/quote]\r\n\r\n1) Yes, I downloaded Resnet-200 model/network trained on Imagenet data.\r\n2) Yes, I am fine-tuning/retraining downloaded pretrained network for only 10 classes instead of original 1000 classes for Imagenet. Basically, last layer of 1000 neurons got removed and replaced with 10 neurons layer.",
      "votes": null
    },
    {
      "id": "116685",
      "postDate": "04/25/2016 16:46:10",
      "content": "<p>[quote=ZFTurbo;116637]</p>\n\n<p>Can anyone clarify what do we mean by <strong>PRE-Trained model</strong>? </p>\n\n<p>1) You use already trained CNN model, which were trained on other pictures not related to this problem</p>\n\n<p>2) You train model on pictures from this problem using CNN structure from other method</p>\n\n<p>[/quote]</p>\n\n<p>The 1) item is pretraining.</p>\n\n<p>General algorithm is \n1) train model on large dataset (or download already trained model) \n2) initialize same network by obtained weights and train on other dataset.</p>\n\n<p>One of the ideas of pretraining is to prevent overfitting on small datasets via making good weights initialization. Futhermore, retrain is faster than traing from scratch.\nYou can find more details in <a href=\"http://cs231n.github.io/transfer-learning/\">http://cs231n.github.io/transfer-learning/</a></p>",
      "rawMarkdown": "[quote=ZFTurbo;116637]\r\n\r\nCan anyone clarify what do we mean by **PRE-Trained model**? \r\n\r\n1) You use already trained CNN model, which were trained on other pictures not related to this problem\r\n\r\n2) You train model on pictures from this problem using CNN structure from other method\r\n\r\n\r\n[/quote]\r\n\r\nThe 1) item is pretraining.\r\n\r\nGeneral algorithm is \r\n1) train model on large dataset (or download already trained model) \r\n2) initialize same network by obtained weights and train on other dataset.\r\n\r\nOne of the ideas of pretraining is to prevent overfitting on small datasets via making good weights initialization. Futhermore, retrain is faster than traing from scratch.\r\nYou can find more details in http://cs231n.github.io/transfer-learning/",
      "votes": null
    },
    {
      "id": "116717",
      "postDate": "04/25/2016 18:19:09",
      "content": "<p><strong>Oleksandr Baiev</strong>, thanks for useful info. I tried 1) and 2) separate, but I see there are good methods which allow to use both at once.</p>",
      "rawMarkdown": "**Oleksandr Baiev**, thanks for useful info. I tried 1) and 2) separate, but I see there are good methods which allow to use both at once.",
      "votes": null
    },
    {
      "id": "116753",
      "postDate": "04/25/2016 23:10:25",
      "content": "<p>Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43</p>",
      "rawMarkdown": "Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43",
      "votes": null
    },
    {
      "id": "116760",
      "postDate": "04/25/2016 23:41:12",
      "content": "<p>[quote=SecondPlan;116753]</p>\n\n<p>Pre-trained with resnet-34, got 0.52 LB</p>\n\n<p>[/quote]</p>\n\n<p>Interesting.</p>",
      "rawMarkdown": "[quote=SecondPlan;116753]\r\n\r\nPre-trained with resnet-34, got 0.52 LB\r\n\r\n[/quote]\r\n\r\nInteresting.",
      "votes": null
    },
    {
      "id": "117730",
      "postDate": "04/30/2016 15:41:39",
      "content": "<p>[quote=SecondPlan;116753]</p>\n\n<p>Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43</p>\n\n<p>[/quote]</p>\n\n<p>Which hyperparameters did you tune? Just the LR, or something else as well?</p>",
      "rawMarkdown": "[quote=SecondPlan;116753]\r\n\r\nPre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43\r\n\r\n[/quote]\r\n\r\nWhich hyperparameters did you tune? Just the LR, or something else as well?",
      "votes": null
    },
    {
      "id": "117737",
      "postDate": "04/30/2016 16:13:22",
      "content": "<p>I am wondering what the exact way to use pre-trained model. Is there any tutorial? Is it like using pre-trained model to extract features from images and then fitting by other model? </p>",
      "rawMarkdown": "I am wondering what the exact way to use pre-trained model. Is there any tutorial? Is it like using pre-trained model to extract features from images and then fitting by other model?",
      "votes": null
    },
    {
      "id": "117739",
      "postDate": "04/30/2016 16:23:29",
      "content": "<p>Here is a good example of extracting feature from a retrained model and using them for a new classification task:</p>\n\n<p><a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code\">https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code</a></p>",
      "rawMarkdown": "Here is a good example of extracting feature from a retrained model and using them for a new classification task:\r\n\r\nhttps://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code",
      "votes": null
    },
    {
      "id": "117745",
      "postDate": "04/30/2016 16:49:06",
      "content": "<p>[quote=Bojan Tunguz;117730]</p>\n\n<p>[quote=SecondPlan;116753]</p>\n\n<p>Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43</p>\n\n<p>[/quote]</p>\n\n<p>Which hyperparameters did you tune? Just the LR, or something else as well?</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, LR is the most important one, haven't tune others yet</p>",
      "rawMarkdown": "[quote=Bojan Tunguz;117730]\r\n\r\n[quote=SecondPlan;116753]\r\n\r\nPre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43\r\n\r\n[/quote]\r\n\r\nWhich hyperparameters did you tune? Just the LR, or something else as well?\r\n\r\n\r\n[/quote]\r\n\r\nYeah, LR is the most important one, haven't tune others yet",
      "votes": null
    },
    {
      "id": "117746",
      "postDate": "04/30/2016 16:53:03",
      "content": "<p>[quote=SecondPlan;117745]</p>\n\n<p>[quote=Bojan Tunguz;117730]</p>\n\n<p>[quote=SecondPlan;116753]</p>\n\n<p>Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, LR is the most important one, haven't tune others yet</p>\n\n<p>[/quote]\nDid you use center crop as in original Torch ResNet implementation or different image preprocessing?</p>",
      "rawMarkdown": "[quote=SecondPlan;117745]\r\n\r\n[quote=Bojan Tunguz;117730]\r\n\r\n[quote=SecondPlan;116753]\r\n\r\nPre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\nYeah, LR is the most important one, haven't tune others yet\r\n\r\n[/quote]\r\nDid you use center crop as in original Torch ResNet implementation or different image preprocessing?",
      "votes": null
    },
    {
      "id": "117752",
      "postDate": "04/30/2016 17:30:13",
      "content": "<p>@Bojan Thanks for the link. Will try to learn from it.</p>",
      "rawMarkdown": "Bojan Thanks for the link. Will try to learn from it.",
      "votes": null
    },
    {
      "id": "118795",
      "postDate": "05/05/2016 11:17:43",
      "content": "<p>[quote=woshialex;116411]</p>\n\n<p>I have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..</p>\n\n<p>[/quote]\nCan you share some ideas regarding your CNN structure? I am managing a score of only about 0.8 with my current structure.</p>",
      "rawMarkdown": "[quote=woshialex;116411]\r\n\r\nI have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..\r\n\r\n[/quote]\r\nCan you share some ideas regarding your CNN structure? I am managing a score of only about 0.8 with my current structure.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 116204,
      "author_name": "pennacchio",
      "author_url": "",
      "post_date": "04/22/2016 14:04:18",
      "content": "<p>I think most of the top-10 scores are being obtained using pre-trained models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116411,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "04/24/2016 01:13:09",
      "content": "<p>I have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116440,
      "author_name": "matteopresutto",
      "author_url": "",
      "post_date": "04/24/2016 09:04:15",
      "content": "<p>[quote=woshialex;116411]</p>\n\n<p>I have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..</p>\n\n<p>[/quote]</p>\n\n<p>What's your accuracy woshialex? Maybe you can boost the logloss tuning the model's confidence</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116467,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "04/24/2016 16:53:32",
      "content": "<p>[quote=woshialex;116411]</p>\n\n<p>... the best LB score I can get at the moment is only about 0.46..</p>\n\n<p>[/quote]\nThis is with just a CNN running on the whole image? Care to share some details of your model and parameters? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116513,
      "author_name": "serhiy",
      "author_url": "",
      "post_date": "04/24/2016 22:35:04",
      "content": "<p>[quote]</p>\n\n<p>This is with just a CNN running on the whole image? Care to share some details of your model and parameters? </p>\n\n<p>[/quote]</p>\n\n<p>I trained Torch based Resnet-50 &quot;as is&quot; on competition training images as they are structured very similar to Imagenet and got 0.59924 LB score. Tried the same with Resnet-101 and 152 and got worst results.</p>\n\n<p>PS. I am using single GeForce 980 Ti GPU for training so my batch size floats from 10 to 30 depending on network size.</p>\n\n<p>One of the drawbacks of Torch version of Resnet - it crops image from the center and, in many cases, trained network cannot distinguish between driving (category C0) and operating the radio (category C5).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116608,
      "author_name": "potamitis",
      "author_url": "",
      "post_date": "04/25/2016 11:48:38",
      "content": "<p>can you share the confusion matrix?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116637,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/25/2016 14:14:25",
      "content": "<p>Can anyone clarify what do we mean by <strong>PRE-Trained model</strong>? </p>\n\n<p>1) You use already trained CNN model, which were trained on other pictures not related to this problem</p>\n\n<p>2) You train model on pictures from this problem using CNN structure from other method</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116644,
      "author_name": "serhiy",
      "author_url": "",
      "post_date": "04/25/2016 14:44:44",
      "content": "<p>[quote=ZFTurbo;116637]</p>\n\n<p>Can anyone clarify what do we mean by <strong>PRE-Trained model</strong>? </p>\n\n<p>1) You use already trained CNN model, which were trained on other pictures not related to this problem</p>\n\n<p>2) You train model on pictures from this problem using CNN structure from other method</p>\n\n<p>[/quote]</p>\n\n<p>1) Yes, I downloaded Resnet-200 model/network trained on Imagenet data.\n2) Yes, I am fine-tuning/retraining downloaded pretrained network for only 10 classes instead of original 1000 classes for Imagenet. Basically, last layer of 1000 neurons got removed and replaced with 10 neurons layer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116685,
      "author_name": "obaiev",
      "author_url": "",
      "post_date": "04/25/2016 16:46:10",
      "content": "<p>[quote=ZFTurbo;116637]</p>\n\n<p>Can anyone clarify what do we mean by <strong>PRE-Trained model</strong>? </p>\n\n<p>1) You use already trained CNN model, which were trained on other pictures not related to this problem</p>\n\n<p>2) You train model on pictures from this problem using CNN structure from other method</p>\n\n<p>[/quote]</p>\n\n<p>The 1) item is pretraining.</p>\n\n<p>General algorithm is \n1) train model on large dataset (or download already trained model) \n2) initialize same network by obtained weights and train on other dataset.</p>\n\n<p>One of the ideas of pretraining is to prevent overfitting on small datasets via making good weights initialization. Futhermore, retrain is faster than traing from scratch.\nYou can find more details in <a href=\"http://cs231n.github.io/transfer-learning/\">http://cs231n.github.io/transfer-learning/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116717,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/25/2016 18:19:09",
      "content": "<p><strong>Oleksandr Baiev</strong>, thanks for useful info. I tried 1) and 2) separate, but I see there are good methods which allow to use both at once.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116753,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "04/25/2016 23:10:25",
      "content": "<p>Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116760,
      "author_name": "serhiy",
      "author_url": "",
      "post_date": "04/25/2016 23:41:12",
      "content": "<p>[quote=SecondPlan;116753]</p>\n\n<p>Pre-trained with resnet-34, got 0.52 LB</p>\n\n<p>[/quote]</p>\n\n<p>Interesting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117730,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "04/30/2016 15:41:39",
      "content": "<p>[quote=SecondPlan;116753]</p>\n\n<p>Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43</p>\n\n<p>[/quote]</p>\n\n<p>Which hyperparameters did you tune? Just the LR, or something else as well?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117737,
      "author_name": "zhugds",
      "author_url": "",
      "post_date": "04/30/2016 16:13:22",
      "content": "<p>I am wondering what the exact way to use pre-trained model. Is there any tutorial? Is it like using pre-trained model to extract features from images and then fitting by other model? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117739,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "04/30/2016 16:23:29",
      "content": "<p>Here is a good example of extracting feature from a retrained model and using them for a new classification task:</p>\n\n<p><a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code\">https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117745,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "04/30/2016 16:49:06",
      "content": "<p>[quote=Bojan Tunguz;117730]</p>\n\n<p>[quote=SecondPlan;116753]</p>\n\n<p>Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43</p>\n\n<p>[/quote]</p>\n\n<p>Which hyperparameters did you tune? Just the LR, or something else as well?</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, LR is the most important one, haven't tune others yet</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117746,
      "author_name": "serhiy",
      "author_url": "",
      "post_date": "04/30/2016 16:53:03",
      "content": "<p>[quote=SecondPlan;117745]</p>\n\n<p>[quote=Bojan Tunguz;117730]</p>\n\n<p>[quote=SecondPlan;116753]</p>\n\n<p>Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, LR is the most important one, haven't tune others yet</p>\n\n<p>[/quote]\nDid you use center crop as in original Torch ResNet implementation or different image preprocessing?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117752,
      "author_name": "zhugds",
      "author_url": "",
      "post_date": "04/30/2016 17:30:13",
      "content": "<p>@Bojan Thanks for the link. Will try to learn from it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118795,
      "author_name": "abhijayvuyyuru",
      "author_url": "",
      "post_date": "05/05/2016 11:17:43",
      "content": "<p>[quote=woshialex;116411]</p>\n\n<p>I have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..</p>\n\n<p>[/quote]\nCan you share some ideas regarding your CNN structure? I am managing a score of only about 0.8 with my current structure.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116200": "Just be curious about the best score when using pre-trained model.",
    "116204": "I think most of the top-10 scores are being obtained using pre-trained models.",
    "116411": "I have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..",
    "116440": "[quote=woshialex;116411]\r\n\r\nI have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..\r\n\r\n[/quote]\r\n\r\nWhat's your accuracy woshialex? Maybe you can boost the logloss tuning the model's confidence",
    "116467": "[quote=woshialex;116411]\r\n\r\n... the best LB score I can get at the moment is only about 0.46..\r\n\r\n[/quote]\r\nThis is with just a CNN running on the whole image? Care to share some details of your model and parameters?",
    "116513": "[quote]\r\n\r\n\r\nThis is with just a CNN running on the whole image? Care to share some details of your model and parameters? \r\n\r\n[/quote]\r\n\r\nI trained Torch based Resnet-50 \"as is\" on competition training images as they are structured very similar to Imagenet and got 0.59924 LB score. Tried the same with Resnet-101 and 152 and got worst results.\r\n\r\nPS. I am using single GeForce 980 Ti GPU for training so my batch size floats from 10 to 30 depending on network size.\r\n\r\nOne of the drawbacks of Torch version of Resnet - it crops image from the center and, in many cases, trained network cannot distinguish between driving (category C0) and operating the radio (category C5).",
    "116608": "can you share the confusion matrix?",
    "116637": "Can anyone clarify what do we mean by **PRE-Trained model**? \r\n\r\n1) You use already trained CNN model, which were trained on other pictures not related to this problem\r\n\r\n2) You train model on pictures from this problem using CNN structure from other method",
    "116644": "[quote=ZFTurbo;116637]\r\n\r\nCan anyone clarify what do we mean by **PRE-Trained model**? \r\n\r\n1) You use already trained CNN model, which were trained on other pictures not related to this problem\r\n\r\n2) You train model on pictures from this problem using CNN structure from other method\r\n\r\n\r\n[/quote]\r\n\r\n1) Yes, I downloaded Resnet-200 model/network trained on Imagenet data.\r\n2) Yes, I am fine-tuning/retraining downloaded pretrained network for only 10 classes instead of original 1000 classes for Imagenet. Basically, last layer of 1000 neurons got removed and replaced with 10 neurons layer.",
    "116685": "[quote=ZFTurbo;116637]\r\n\r\nCan anyone clarify what do we mean by **PRE-Trained model**? \r\n\r\n1) You use already trained CNN model, which were trained on other pictures not related to this problem\r\n\r\n2) You train model on pictures from this problem using CNN structure from other method\r\n\r\n\r\n[/quote]\r\n\r\nThe 1) item is pretraining.\r\n\r\nGeneral algorithm is \r\n1) train model on large dataset (or download already trained model) \r\n2) initialize same network by obtained weights and train on other dataset.\r\n\r\nOne of the ideas of pretraining is to prevent overfitting on small datasets via making good weights initialization. Futhermore, retrain is faster than traing from scratch.\r\nYou can find more details in http://cs231n.github.io/transfer-learning/",
    "116717": "**Oleksandr Baiev**, thanks for useful info. I tried 1) and 2) separate, but I see there are good methods which allow to use both at once.",
    "116753": "Pre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43",
    "116760": "[quote=SecondPlan;116753]\r\n\r\nPre-trained with resnet-34, got 0.52 LB\r\n\r\n[/quote]\r\n\r\nInteresting.",
    "117730": "[quote=SecondPlan;116753]\r\n\r\nPre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43\r\n\r\n[/quote]\r\n\r\nWhich hyperparameters did you tune? Just the LR, or something else as well?",
    "117737": "I am wondering what the exact way to use pre-trained model. Is there any tutorial? Is it like using pre-trained model to extract features from images and then fitting by other model?",
    "117739": "Here is a good example of extracting feature from a retrained model and using them for a new classification task:\r\n\r\nhttps://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/19206/deep-learning-starter-code",
    "117745": "[quote=Bojan Tunguz;117730]\r\n\r\n[quote=SecondPlan;116753]\r\n\r\nPre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43\r\n\r\n[/quote]\r\n\r\nWhich hyperparameters did you tune? Just the LR, or something else as well?\r\n\r\n\r\n[/quote]\r\n\r\nYeah, LR is the most important one, haven't tune others yet",
    "117746": "[quote=SecondPlan;117745]\r\n\r\n[quote=Bojan Tunguz;117730]\r\n\r\n[quote=SecondPlan;116753]\r\n\r\nPre-trained with resnet-34, got 0.52 LB. Update: tweaking hyperparameters on finetuning, leads to 0.43\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\nYeah, LR is the most important one, haven't tune others yet\r\n\r\n[/quote]\r\nDid you use center crop as in original Torch ResNet implementation or different image preprocessing?",
    "117752": "Bojan Thanks for the link. Will try to learn from it.",
    "118795": "[quote=woshialex;116411]\r\n\r\nI have the same wonder whether people are using pre-trained models to get higher score. I just used random initialization and tried many different CNN structures and image pre-processing but the best LB score I can get at the moment is only about 0.46..\r\n\r\n[/quote]\r\nCan you share some ideas regarding your CNN structure? I am managing a score of only about 0.8 with my current structure."
  },
  "source": "meta"
}