{
  "id": 46688,
  "title": "Sharing my approach",
  "url": "/competitions/sp-society-camera-model-identification/discussion/46688",
  "author_name": "",
  "post_date": "2018-01-01T17:24:58.276771900Z",
  "votes": 26,
  "comment_count": 80,
  "views": 0,
  "content": "<p>I just tried to fine-tune the Resnet50 model. I used 8 GPUS  and the batch size is 25. I fine-tuned the model using all the training data for 10000 iters. The model finetued for 7500 iters achieves the best LB now. In the prototxt, I set the crop_size 224.  I convert the tif files to jpg files for testing. \nI think this approach is simple. You can have a try.\nIf you have any advice, please tell me. Thanks!</p>",
  "messages": [
    {
      "id": "263932",
      "postDate": "01/01/2018 17:24:58",
      "content": "<p>I just tried to fine-tune the Resnet50 model. I used 8 GPUS  and the batch size is 25. I fine-tuned the model using all the training data for 10000 iters. The model finetued for 7500 iters achieves the best LB now. In the prototxt, I set the crop_size 224.  I convert the tif files to jpg files for testing. \nI think this approach is simple. You can have a try.\nIf you have any advice, please tell me. Thanks!</p>",
      "rawMarkdown": "I just tried to fine-tune the Resnet50 model. I used 8 GPUS  and the batch size is 25. I fine-tuned the model using all the training data for 10000 iters. The model finetued for 7500 iters achieves the best LB now. In the prototxt, I set the crop_size 224.  I convert the tif files to jpg files for testing. \nI think this approach is simple. You can have a try.\nIf you have any advice, please tell me. Thanks!",
      "votes": null
    },
    {
      "id": "264321",
      "postDate": "01/02/2018 19:56:18",
      "content": "<p>Hi Yong, thanks for the sharing! I am using similar approach, but my training loss and validation loss were stuck around 0.8 after 500 iters, do you mind share your thoughts in fine tuning optimizer part?</p>",
      "rawMarkdown": "Hi Yong, thanks for the sharing! I am using similar approach, but my training loss and validation loss were stuck around 0.8 after 500 iters, do you mind share your thoughts in fine tuning optimizer part?",
      "votes": null
    },
    {
      "id": "264414",
      "postDate": "01/03/2018 01:50:55",
      "content": "<p>Data augmentation using the possible processing operations may bring improvements.</p>",
      "rawMarkdown": "Data augmentation using the possible processing operations may bring improvements.",
      "votes": null
    },
    {
      "id": "264415",
      "postDate": "01/03/2018 01:51:21",
      "content": "<p>average_loss: 20 \nlr_policy: \"multistep\" \nbase_lr: 0.0002 \ngamma: 0.5 \nstepvalue: 5000 \nstepvalue: 10000 \nstepvalue: 15000 \nmax_iter:  20000 \ndisplay: 50 \nmomentum: 0.9 \nweight_decay: 0.0005 \nsnapshot: 2500</p>",
      "rawMarkdown": "average_loss: 20 \nlr_policy: \"multistep\" \nbase_lr: 0.0002 \ngamma: 0.5 \nstepvalue: 5000 \nstepvalue: 10000 \nstepvalue: 15000 \nmax_iter:  20000 \ndisplay: 50 \nmomentum: 0.9 \nweight_decay: 0.0005 \nsnapshot: 2500",
      "votes": null
    },
    {
      "id": "264949",
      "postDate": "01/04/2018 07:59:40",
      "content": "<p>Hi, Young!  I downloaded the train data set and moved the images of 10 models into one folders, therefore I miss the correct form of those labels. Would you please show me the right label names of these 10 models? Thanks.</p>",
      "rawMarkdown": "Hi, Young!  I downloaded the train data set and moved the images of 10 models into one folders, therefore I miss the correct form of those labels. Would you please show me the right label names of these 10 models? Thanks.",
      "votes": null
    },
    {
      "id": "265005",
      "postDate": "01/04/2018 10:52:02",
      "content": "<p>OK.\n'Motorola-X', 'Motorola-Nexus-6', 'Samsung-Galaxy-S4', 'Samsung-Galaxy-Note3', 'LG-Nexus-5x', \n            'iPhone-4s', 'Motorola-Droid-Maxx', 'HTC-1-M7', 'Sony-NEX-7', 'iPhone-6'</p>",
      "rawMarkdown": "OK.\n'Motorola-X', 'Motorola-Nexus-6', 'Samsung-Galaxy-S4', 'Samsung-Galaxy-Note3', 'LG-Nexus-5x', \n            'iPhone-4s', 'Motorola-Droid-Maxx', 'HTC-1-M7', 'Sony-NEX-7', 'iPhone-6'",
      "votes": null
    },
    {
      "id": "265278",
      "postDate": "01/05/2018 03:19:01",
      "content": "<p>Thanks a lot!!!</p>",
      "rawMarkdown": "Thanks a lot!!!",
      "votes": null
    },
    {
      "id": "265283",
      "postDate": "01/05/2018 03:36:29",
      "content": "<p>How do you train a neural network using multiple GPUs? Do you have any informative tutorials at hand?</p>",
      "rawMarkdown": "How do you train a neural network using multiple GPUs? Do you have any informative tutorials at hand?",
      "votes": null
    },
    {
      "id": "265284",
      "postDate": "01/05/2018 03:40:43",
      "content": "<p>Hi Young! You mentioned that the tif files were converted to jpg files for testing. How large is your jpeg depression quality in your conversion?</p>",
      "rawMarkdown": "Hi Young! You mentioned that the tif files were converted to jpg files for testing. How large is your jpeg depression quality in your conversion?",
      "votes": null
    },
    {
      "id": "265287",
      "postDate": "01/05/2018 03:59:16",
      "content": "<p>I think most kinds of deep learning platforms (like Tensorflow and Keras) will support multi-GPU training. The doc will be helpful :)</p>",
      "rawMarkdown": "I think most kinds of deep learning platforms (like Tensorflow and Keras) will support multi-GPU training. The doc will be helpful :)",
      "votes": null
    },
    {
      "id": "265288",
      "postDate": "01/05/2018 04:02:14",
      "content": "<p>Why converting tiff to jpeg? It may lead to some infomation loss.</p>",
      "rawMarkdown": "Why converting tiff to jpeg? It may lead to some infomation loss.",
      "votes": null
    },
    {
      "id": "265331",
      "postDate": "01/05/2018 08:11:21",
      "content": "<p>95.\nHowever, I later found that using tif files to test is better.</p>",
      "rawMarkdown": "95.\nHowever, I later found that using tif files to test is better.",
      "votes": null
    },
    {
      "id": "265332",
      "postDate": "01/05/2018 08:11:49",
      "content": "<p>Yes, you are right.</p>",
      "rawMarkdown": "Yes, you are right.",
      "votes": null
    },
    {
      "id": "265955",
      "postDate": "01/07/2018 08:00:52",
      "content": "<p>InceptionResNetV2 may be better.</p>",
      "rawMarkdown": "InceptionResNetV2 may be better.",
      "votes": null
    },
    {
      "id": "265957",
      "postDate": "01/07/2018 08:05:09",
      "content": "<p>Did you just finetune the FC layers or even more? </p>",
      "rawMarkdown": "Did you just finetune the FC layers or even more?",
      "votes": null
    },
    {
      "id": "265977",
      "postDate": "01/07/2018 10:08:48",
      "content": "<p>Just FC layers now.</p>",
      "rawMarkdown": "Just FC layers now.",
      "votes": null
    },
    {
      "id": "266312",
      "postDate": "01/08/2018 12:49:14",
      "content": "<p>For a single Resnet50 model , you can get LB of around 0.85, depending on how well you train.</p>",
      "rawMarkdown": "For a single Resnet50 model , you can get LB of around 0.85, depending on how well you train.",
      "votes": null
    },
    {
      "id": "266330",
      "postDate": "01/08/2018 13:59:20",
      "content": "<p>Hi! why did you use such a little batch size?</p>",
      "rawMarkdown": "Hi! why did you use such a little batch size?",
      "votes": null
    },
    {
      "id": "266506",
      "postDate": "01/09/2018 01:17:41",
      "content": "<p>First I applied some pre processing (gama correction, jpeg compression, resizing) and cropped all the training data in 512x512 blocks. Then I resized to 256x256 and tried to fine-tune the InceptionV2 model. I do not have a powerful gpu so I'm suffering with the delay of results. It seems like I'm not doing something very different, but I can not achieve more than 0.35 in final test. Any tips? </p>",
      "rawMarkdown": "First I applied some pre processing (gama correction, jpeg compression, resizing) and cropped all the training data in 512x512 blocks. Then I resized to 256x256 and tried to fine-tune the InceptionV2 model. I do not have a powerful gpu so I'm suffering with the delay of results. It seems like I'm not doing something very different, but I can not achieve more than 0.35 in final test. Any tips?",
      "votes": null
    },
    {
      "id": "266543",
      "postDate": "01/09/2018 03:29:34",
      "content": "<p>I just crop the data while I'm training and testing. You may try not to resize the picture. Train for more epochs.</p>",
      "rawMarkdown": "I just crop the data while I'm training and testing. You may try not to resize the picture. Train for more epochs.",
      "votes": null
    },
    {
      "id": "266576",
      "postDate": "01/09/2018 04:47:36",
      "content": "<p>This batch size is for a single GPU. </p>",
      "rawMarkdown": "This batch size is for a single GPU.",
      "votes": null
    },
    {
      "id": "266972",
      "postDate": "01/10/2018 08:40:51",
      "content": "<p>your gpus doesn't allow to make bigger batches?</p>",
      "rawMarkdown": "your gpus doesn't allow to make bigger batches?",
      "votes": null
    },
    {
      "id": "267077",
      "postDate": "01/10/2018 14:33:07",
      "content": "<p>Hi, it is not a good idea to resize the images. The cameras can be identified based on some proprietary color interpolation patterns and sensor noise patterns which can be detected on the raw pixels of the images. Resizing the images alters this information and this is basically why this competition exists. </p>",
      "rawMarkdown": "Hi, it is not a good idea to resize the images. The cameras can be identified based on some proprietary color interpolation patterns and sensor noise patterns which can be detected on the raw pixels of the images. Resizing the images alters this information and this is basically why this competition exists.",
      "votes": null
    },
    {
      "id": "267243",
      "postDate": "01/10/2018 22:26:22",
      "content": "<p>Thanks for the comments. You are right, it's better not to resize. But with full images in the training dataset I'll have to use a very small batch size. I'll try train for more epochs too. Thanks again.</p>",
      "rawMarkdown": "Thanks for the comments. You are right, it's better not to resize. But with full images in the training dataset I'll have to use a very small batch size. I'll try train for more epochs too. Thanks again.",
      "votes": null
    },
    {
      "id": "267363",
      "postDate": "01/11/2018 06:31:28",
      "content": "<p>Hi IgorMuniz, I was using the same approach as you. The highest I can get was 0.42. I also did fine-tuning on pre-trained ResNet50. I doubt the difference lays on the method. I think Young was using Caffe for the training and I was implementing things in Keras.</p>",
      "rawMarkdown": "Hi IgorMuniz, I was using the same approach as you. The highest I can get was 0.42. I also did fine-tuning on pre-trained ResNet50. I doubt the difference lays on the method. I think Young was using Caffe for the training and I was implementing things in Keras.",
      "votes": null
    },
    {
      "id": "268545",
      "postDate": "01/14/2018 20:30:52",
      "content": "<p>Hi guys,</p>\n\n<p>So am I getting this right that this competition, it's not about what's in the image, but more about the fundamental underlying qualities about the image itself.</p>\n\n<p>Also, how long and what hardware are you guys using to train the model?</p>\n\n<p>Best.</p>",
      "rawMarkdown": "Hi guys,\n\nSo am I getting this right that this competition, it's not about what's in the image, but more about the fundamental underlying qualities about the image itself.\n\nAlso, how long and what hardware are you guys using to train the model?\n\nBest.",
      "votes": null
    },
    {
      "id": "268597",
      "postDate": "01/15/2018 00:19:11",
      "content": "<p>Hi Shunjia Ding, \nI can see you've improved your acc. I did some improvements using ResNet too (in Keras). Are you guys normalizing the data? I can't find how the input data should be in Keras (RGB or BGR/ normalizing or not) </p>\n\n<p>Bo Peng, \nI do not have a powerfull GPU (gtx 930m), it takes around 1000s/epoch to fine-tuning ResNet50</p>",
      "rawMarkdown": "Hi Shunjia Ding, \nI can see you've improved your acc. I did some improvements using ResNet too (in Keras). Are you guys normalizing the data? I can't find how the input data should be in Keras (RGB or BGR/ normalizing or not) \n\nBo Peng, \nI do not have a powerfull GPU (gtx 930m), it takes around 1000s/epoch to fine-tuning ResNet50",
      "votes": null
    },
    {
      "id": "268649",
      "postDate": "01/15/2018 05:21:16",
      "content": "<p>Hi IgorMuniz, I managed to improve my score. I implemented the random cropping in Keras which was similar to Caffe. The more training examples, the better LB score you would get. In the end, I found it's pretty difficult to minimize the val_loss because I only used center crop for validation set. Learning rate was very important, it took me a while to find a good strategy to get a good learning rate that eventually lowered the loss.</p>\n\n<p>Not only you need a good GPU to train the ResNet50, you also need a SSD for speeding up (image reading). In the end, I ran my python code to generate random patches from the original images up to 30GB.</p>",
      "rawMarkdown": "Hi IgorMuniz, I managed to improve my score. I implemented the random cropping in Keras which was similar to Caffe. The more training examples, the better LB score you would get. In the end, I found it's pretty difficult to minimize the val_loss because I only used center crop for validation set. Learning rate was very important, it took me a while to find a good strategy to get a good learning rate that eventually lowered the loss.\n\nNot only you need a good GPU to train the ResNet50, you also need a SSD for speeding up (image reading). In the end, I ran my python code to generate random patches from the original images up to 30GB.",
      "votes": null
    },
    {
      "id": "268755",
      "postDate": "01/15/2018 13:38:49",
      "content": "<p>Hello ,Young. I appreciate ur good work ..Also i have some confused questions to ask u ...Did u use the extra dataset from additional data of Gleb's ? If so , which photos did u use to train, just good_jpgs or all photos? And how did u deal with the problem of the different number of each camera?  if u are in convenience , thanks for ur reply...thank u again..</p>",
      "rawMarkdown": "Hello ,Young. I appreciate ur good work ..Also i have some confused questions to ask u ...Did u use the extra dataset from additional data of Gleb's ? If so , which photos did u use to train, just good_jpgs or all photos? And how did u deal with the problem of the different number of each camera?  if u are in convenience , thanks for ur reply...thank u again..",
      "votes": null
    },
    {
      "id": "268763",
      "postDate": "01/15/2018 14:13:49",
      "content": "<p>Yes, I am using the extra data from Gleb to train the model and it has brought big improvements. I use the good_jpgs. I don't take the problem of the different number of each camera into consideration now.</p>",
      "rawMarkdown": "Yes, I am using the extra data from Gleb to train the model and it has brought big improvements. I use the good_jpgs. I don't take the problem of the different number of each camera into consideration now.",
      "votes": null
    },
    {
      "id": "268766",
      "postDate": "01/15/2018 14:29:04",
      "content": "<p>thanks...</p>",
      "rawMarkdown": "thanks...",
      "votes": null
    },
    {
      "id": "268806",
      "postDate": "01/15/2018 17:48:46",
      "content": "<p>I'm new here so this question can be kind of stupid... \nIt is allowed to use extra data?</p>",
      "rawMarkdown": "I'm new here so this question can be kind of stupid... \nIt is allowed to use extra data?",
      "votes": null
    },
    {
      "id": "268965",
      "postDate": "01/16/2018 01:43:07",
      "content": "<p>Maybe. I am not sure.</p>",
      "rawMarkdown": "Maybe. I am not sure.",
      "votes": null
    },
    {
      "id": "269017",
      "postDate": "01/16/2018 04:07:20",
      "content": "<p>In most of Kaggle competitions use of extra data is allowed as long as you post your source in the forum, so it is available for every one. However, in this competition, it is unknown, as no one specifically allowed nor denied use of external data.</p>",
      "rawMarkdown": "In most of Kaggle competitions use of extra data is allowed as long as you post your source in the forum, so it is available for every one. However, in this competition, it is unknown, as no one specifically allowed nor denied use of external data.",
      "votes": null
    },
    {
      "id": "269212",
      "postDate": "01/16/2018 13:30:21",
      "content": "<p>The answer is in the rules : \n\"External Data Use: The following provision supersedes General Rules Section 7.B. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; <strong>provided, you have the right</strong> and authority to use such external data for the purposes of the Competition, and <strong>to share such data with Sponsor and Kaggle</strong> as may be required.\"</p>\n\n<p>So, if you can use and share to the organizers the data, it is allowed</p>",
      "rawMarkdown": "The answer is in the rules : \n\"External Data Use: The following provision supersedes General Rules Section 7.B. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; **provided, you have the right** and authority to use such external data for the purposes of the Competition, and **to share such data with Sponsor and Kaggle** as may be required.\"\n\nSo, if you can use and share to the organizers the data, it is allowed",
      "votes": null
    },
    {
      "id": "269220",
      "postDate": "01/16/2018 13:43:10",
      "content": "<p>Hi Max Diebold, </p>\n\n<p>Where did you find this?\nIn the rules of this competition I found the following:</p>\n\n<p>General Competition Rules\n7. Competition Data\nC. External Data: Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions. Competition Sponsor reserves the right in its sole discretion to disqualify any Participant who Competition Sponsor discovers has undertaken or attempted to undertake the use of data other than the Competition Data, or who uses the Competition Data other than as permitted by the Competition Website and these Rules.</p>",
      "rawMarkdown": "Hi Max Diebold, \n\nWhere did you find this?\nIn the rules of this competition I found the following:\n\nGeneral Competition Rules\n7. Competition Data\nC. External Data: Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions. Competition Sponsor reserves the right in its sole discretion to disqualify any Participant who Competition Sponsor discovers has undertaken or attempted to undertake the use of data other than the Competition Data, or who uses the Competition Data other than as permitted by the Competition Website and these Rules.",
      "votes": null
    },
    {
      "id": "269226",
      "postDate": "01/16/2018 13:51:03",
      "content": "<p>In the A. SPECIFIC COMPETITION RULES section of this competition's rules, on this page</p>",
      "rawMarkdown": "In the A. SPECIFIC COMPETITION RULES section of this competition's rules, on this page",
      "votes": null
    },
    {
      "id": "269247",
      "postDate": "01/16/2018 14:24:08",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "269491",
      "postDate": "01/16/2018 23:25:47",
      "content": "<p>Thanks Young (and to Gleb of course for creating the data set)!</p>\n\n<p>I'd still like to get an official OK from the competition organizers/Kaggle - while Gleb was conscientious enough to filter for only Creative Commons images from Flickr,  there are different varieties of such licenses. I noticed skimming through the images that some of them were categorized under \"non-commercial\" and others under \"non-derivative'. And most of them were under \"attribution required\".</p>\n\n<p>I don't want to be nit-picky but would using it here fall under those use-cases? </p>",
      "rawMarkdown": "Thanks Young (and to Gleb of course for creating the data set)!\n\nI'd still like to get an official OK from the competition organizers/Kaggle - while Gleb was conscientious enough to filter for only Creative Commons images from Flickr,  there are different varieties of such licenses. I noticed skimming through the images that some of them were categorized under \"non-commercial\" and others under \"non-derivative'. And most of them were under \"attribution required\".\n\nI don't want to be nit-picky but would using it here fall under those use-cases?",
      "votes": null
    },
    {
      "id": "269904",
      "postDate": "01/17/2018 14:03:41",
      "content": "<p>May I ask u a question? Do u use these extra data? I want to use, but I'm afraid of the rejection of the competition.</p>",
      "rawMarkdown": "May I ask u a question? Do u use these extra data? I want to use, but I'm afraid of the rejection of the competition.",
      "votes": null
    },
    {
      "id": "269953",
      "postDate": "01/17/2018 15:14:15",
      "content": "<p>No, we're not using the Flickr images. We wouldn't mind having more training data though :)\nI would be a little surprised if the use of it is banned but there's no harm waiting for confirmation. </p>",
      "rawMarkdown": "No, we're not using the Flickr images. We wouldn't mind having more training data though :)\nI would be a little surprised if the use of it is banned but there's no harm waiting for confirmation.",
      "votes": null
    },
    {
      "id": "269962",
      "postDate": "01/17/2018 15:28:37",
      "content": "<p>Do you just fine-tune the pre-trained models? Do you use any method that is specified for the cameral problem?</p>",
      "rawMarkdown": "Do you just fine-tune the pre-trained models? Do you use any method that is specified for the cameral problem?",
      "votes": null
    },
    {
      "id": "270302",
      "postDate": "01/18/2018 03:11:32",
      "content": "<p>Just finetune the pre-trained models now.</p>",
      "rawMarkdown": "Just finetune the pre-trained models now.",
      "votes": null
    },
    {
      "id": "270311",
      "postDate": "01/18/2018 03:36:45",
      "content": "<p>Thanks, lao ge.</p>",
      "rawMarkdown": "Thanks, lao ge.",
      "votes": null
    },
    {
      "id": "270329",
      "postDate": "01/18/2018 05:09:51",
      "content": "<p>You are welcome, xiongdi.</p>",
      "rawMarkdown": "You are welcome, xiongdi.",
      "votes": null
    },
    {
      "id": "270825",
      "postDate": "01/19/2018 03:21:53",
      "content": "<p>How long did it take to train 10000 iters? How many iters in each epoch?</p>",
      "rawMarkdown": "How long did it take to train 10000 iters? How many iters in each epoch?",
      "votes": null
    },
    {
      "id": "270856",
      "postDate": "01/19/2018 05:08:28",
      "content": "<p>Hi !  Mr. Lee ..I am confused that if u use extra dataset...Or u just use the original data from official dataset  and use some domain knowledge ?  Of course , i just curious about it .u don't need to show the details ..thanks.</p>",
      "rawMarkdown": "Hi !  Mr. Lee ..I am confused that if u use extra dataset...Or u just use the original data from official dataset  and use some domain knowledge ?  Of course , i just curious about it .u don't need to show the details ..thanks.",
      "votes": null
    },
    {
      "id": "270861",
      "postDate": "01/19/2018 05:24:41",
      "content": "<p>Hello! Young! Have you tried to manipulate the train data (jpeg compression,gamma correction ,resizing ) to minimize the difference between the train data and the test data? If not, would you please explain the reason why your result comes out so good with the huge difference between the train data and the test data? What is the main factor  your model working so well?</p>",
      "rawMarkdown": "Hello! Young! Have you tried to manipulate the train data (jpeg compression,gamma correction ,resizing ) to minimize the difference between the train data and the test data? If not, would you please explain the reason why your result comes out so good with the huge difference between the train data and the test data? What is the main factor  your model working so well?",
      "votes": null
    },
    {
      "id": "270862",
      "postDate": "01/19/2018 05:28:01",
      "content": "<p>We're not using the Flickr images. </p>\n\n<p>You do not need to use the extra dataset to achieve a good standing in this competition.</p>",
      "rawMarkdown": "We're not using the Flickr images. \n\nYou do not need to use the extra dataset to achieve a good standing in this competition.",
      "votes": null
    },
    {
      "id": "270883",
      "postDate": "01/19/2018 06:36:28",
      "content": "<p>Sorry, I don't record the time.</p>",
      "rawMarkdown": "Sorry, I don't record the time.",
      "votes": null
    },
    {
      "id": "270884",
      "postDate": "01/19/2018 06:38:02",
      "content": "<p>Yes, I said \"Data augmentation using the possible processing operations may bring improvements\".</p>",
      "rawMarkdown": "Yes, I said \"Data augmentation using the possible processing operations may bring improvements\".",
      "votes": null
    },
    {
      "id": "270909",
      "postDate": "01/19/2018 07:45:00",
      "content": "<p>Hey Young, </p>\n\n<p>Are you treating the sp cup data and the flickr data from gleb equally? </p>\n\n<p>I am struggling adding the flickr data to our model. </p>\n\n<p>Funny thing is that our model with 90+ test accuracy can recognize the flickr data with 80-90+ acc but don't train on these as expected. </p>\n\n<p>And our model is failing horribly correctly predicting the Sony Nex 7 from Gleb. <br>\nI am confused about whether this is showing the weakness of our model or the random processing these sony images went through before being uploaded to flickr. </p>",
      "rawMarkdown": "Hey Young, \n\nAre you treating the sp cup data and the flickr data from gleb equally? \n\nI am struggling adding the flickr data to our model. \n\nFunny thing is that our model with 90+ test accuracy can recognize the flickr data with 80-90+ acc but don't train on these as expected. \n\nAnd our model is failing horribly correctly predicting the Sony Nex 7 from Gleb.  \nI am confused about whether this is showing the weakness of our model or the random processing these sony images went through before being uploaded to flickr.",
      "votes": null
    },
    {
      "id": "270933",
      "postDate": "01/19/2018 09:32:15",
      "content": "<p>Hi. I treated the sp cup data and the flickr data from gleb equally. May there is a gap between them but I didn't check them. </p>",
      "rawMarkdown": "Hi. I treated the sp cup data and the flickr data from gleb equally. May there is a gap between them but I didn't check them.",
      "votes": null
    },
    {
      "id": "270946",
      "postDate": "01/19/2018 09:58:35",
      "content": "<p>Thanks for your sharing. May I ask some questions?</p>\n\n<blockquote>\n  <p>I convert the tif files to jpg files for testing.</p>\n</blockquote>\n\n<p>Why did you convert the TIFF into JPG? Besides, did you make some special handling about the manipulated images in the test set? Lastly, why can you have so many GPUs, haha?</p>",
      "rawMarkdown": "Thanks for your sharing. May I ask some questions?\n\n&gt; I convert the tif files to jpg files for testing.\n\nWhy did you convert the TIFF into JPG? Besides, did you make some special handling about the manipulated images in the test set? Lastly, why can you have so many GPUs, haha?",
      "votes": null
    },
    {
      "id": "270952",
      "postDate": "01/19/2018 10:24:30",
      "content": "<p>It is not a good idea to convert TIFF to JPG. You can see the discussion in some comments above</p>",
      "rawMarkdown": "It is not a good idea to convert TIFF to JPG. You can see the discussion in some comments above",
      "votes": null
    },
    {
      "id": "270956",
      "postDate": "01/19/2018 10:32:35",
      "content": "<p>Got it. Thank you, IgorMuniz.</p>",
      "rawMarkdown": "Got it. Thank you, IgorMuniz.",
      "votes": null
    },
    {
      "id": "270959",
      "postDate": "01/19/2018 10:36:17",
      "content": "<p>My validation accuracy is stuck around 0.8 even with data augmentation. I keep trying fine tuning Resnet50 (for now it has brought better results). One question...</p>\n\n<p>How many epochs have you trained? I'm using EarlyStopping with patience 10 so maybe I just have to train for more epochs...</p>",
      "rawMarkdown": "My validation accuracy is stuck around 0.8 even with data augmentation. I keep trying fine tuning Resnet50 (for now it has brought better results). One question...\n\nHow many epochs have you trained? I'm using EarlyStopping with patience 10 so maybe I just have to train for more epochs...",
      "votes": null
    },
    {
      "id": "270970",
      "postDate": "01/19/2018 10:43:55",
      "content": "<p>relaunch the same task you did, disable the early stopping and you will see if you stopped too early or not. It depends of the minibatch size / data augmentation etc... My early stopping is at 100 epochs for example (Not using resnet)</p>",
      "rawMarkdown": "relaunch the same task you did, disable the early stopping and you will see if you stopped too early or not. It depends of the minibatch size / data augmentation etc... My early stopping is at 100 epochs for example (Not using resnet)",
      "votes": null
    },
    {
      "id": "271286",
      "postDate": "01/20/2018 03:09:39",
      "content": "<p>Hi!  Lee .. u mean that u don't use any extra data which also can get good result..so i just  wonder if u change some structures of network or build a new CNN by yourselves?</p>",
      "rawMarkdown": "Hi!  Lee .. u mean that u don't use any extra data which also can get good result..so i just  wonder if u change some structures of network or build a new CNN by yourselves?",
      "votes": null
    },
    {
      "id": "271295",
      "postDate": "01/20/2018 04:02:40",
      "content": "<p>Hi Zhuo Long, perhaps if you detailed your team's approach, we and other competitors could help you improve it? :)</p>\n\n<p>More seriously, we aren't doing anything funky. I'd suggest you search for older image-related Kaggle competitions and read up the winners' solutions that are posted in the forums (e.g., CDiscount). There's a lot of insight and it's a good way of getting more ideas if you're in a rut. </p>",
      "rawMarkdown": "Hi Zhuo Long, perhaps if you detailed your team's approach, we and other competitors could help you improve it? :)\n\nMore seriously, we aren't doing anything funky. I'd suggest you search for older image-related Kaggle competitions and read up the winners' solutions that are posted in the forums (e.g., CDiscount). There's a lot of insight and it's a good way of getting more ideas if you're in a rut.",
      "votes": null
    },
    {
      "id": "271698",
      "postDate": "01/21/2018 07:06:10",
      "content": "<p>Excuse me, to obtain 0.85, is it required to fine-tune the whole ResNet50 network? Or can we use feature extractor based on imagenet weights and train only the fully-connected layer?\nHave you tried any smaller model (such as DenseNet121 or MobileNet)? I am running low power GPU (laptop GPU 940MX 2GB).\nBesides, what optimizer do you use (SGD, Adam, etc)?</p>\n\n<p>Thank you very much.</p>",
      "rawMarkdown": "Excuse me, to obtain 0.85, is it required to fine-tune the whole ResNet50 network? Or can we use feature extractor based on imagenet weights and train only the fully-connected layer?\nHave you tried any smaller model (such as DenseNet121 or MobileNet)? I am running low power GPU (laptop GPU 940MX 2GB).\nBesides, what optimizer do you use (SGD, Adam, etc)?\n\nThank you very much.",
      "votes": null
    },
    {
      "id": "271976",
      "postDate": "01/22/2018 02:03:27",
      "content": "<p>Young, thanks for your sharing, may I ask u that how large is u train dataset. Now I crop the image of train dataset to produce lots of images that aren't overlap, but I think that the dataset is so large and take a lot of time to train. </p>",
      "rawMarkdown": "Young, thanks for your sharing, may I ask u that how large is u train dataset. Now I crop the image of train dataset to produce lots of images that aren't overlap, but I think that the dataset is so large and take a lot of time to train.",
      "votes": null
    },
    {
      "id": "271992",
      "postDate": "01/22/2018 03:01:03",
      "content": "<p>I use the eight possible processing operations for data augmentation and the network transforms the data by cropping randomly.</p>",
      "rawMarkdown": "I use the eight possible processing operations for data augmentation and the network transforms the data by cropping randomly.",
      "votes": null
    },
    {
      "id": "271996",
      "postDate": "01/22/2018 03:06:30",
      "content": "<p>So it means that u pull the full image into the network and use the random crop? I thought before that put such high resolution image into the network may train very slow, and the batch size should be small. The GPU I use is 1080Ti</p>",
      "rawMarkdown": "So it means that u pull the full image into the network and use the random crop? I thought before that put such high resolution image into the network may train very slow, and the batch size should be small. The GPU I use is 1080Ti",
      "votes": null
    },
    {
      "id": "272038",
      "postDate": "01/22/2018 05:29:07",
      "content": "<p>Yes. You can have a try.</p>",
      "rawMarkdown": "Yes. You can have a try.",
      "votes": null
    },
    {
      "id": "272040",
      "postDate": "01/22/2018 05:33:14",
      "content": "<p>thx and the last question is that may I ask u that what is the best acc u can get by using the single model?Now I use the single model, the best acc is 86.4, and I find it hard to improve.</p>",
      "rawMarkdown": "thx and the last question is that may I ask u that what is the best acc u can get by using the single model?Now I use the single model, the best acc is 86.4, and I find it hard to improve.",
      "votes": null
    },
    {
      "id": "272054",
      "postDate": "01/22/2018 06:42:30",
      "content": "<p>Hi, Chun Ming Lee, I admire your great score in this competition and thanks for your guide above. It is seemed that you have learned a lot from other competitions.  I am new  to images competitions in kaggle. Would you please recommend some more classic image-related Kaggle  competitions where I can learn  some ideas?</p>",
      "rawMarkdown": "Hi, Chun Ming Lee, I admire your great score in this competition and thanks for your guide above. It is seemed that you have learned a lot from other competitions.  I am new  to images competitions in kaggle. Would you please recommend some more classic image-related Kaggle  competitions where I can learn  some ideas?",
      "votes": null
    },
    {
      "id": "272062",
      "postDate": "01/22/2018 07:07:22",
      "content": "<p>Off the top of my head, I thought the winning solutions for the Carvana and CDiscount competitions were brilliant. You can also search under the Competitions bar for competitions with the tag \"image data\". </p>\n\n<p>I'd also recommend reading the posts of a image \"pro\", Heng CherKeng (<a href=\"https://www.kaggle.com/hengck23\">https://www.kaggle.com/hengck23</a>). He's very generous in detailing his approach and models in the middle of a competition.</p>\n\n<p>Lastly and this is not directed at you, I'd say having a mindset of learning (whether reading research papers, going through past competition solutions etc.) is better long-term than trying to fish for optimal hyper-parameters from others in a competition (especially if you're not sharing anything in return). Again, thank you to Young for being so open with his approach as a top competitor. </p>",
      "rawMarkdown": "Off the top of my head, I thought the winning solutions for the Carvana and CDiscount competitions were brilliant. You can also search under the Competitions bar for competitions with the tag \"image data\". \n\nI'd also recommend reading the posts of a image \"pro\", Heng CherKeng (https://www.kaggle.com/hengck23). He's very generous in detailing his approach and models in the middle of a competition.\n\nLastly and this is not directed at you, I'd say having a mindset of learning (whether reading research papers, going through past competition solutions etc.) is better long-term than trying to fish for optimal hyper-parameters from others in a competition (especially if you're not sharing anything in return). Again, thank you to Young for being so open with his approach as a top competitor.",
      "votes": null
    },
    {
      "id": "272063",
      "postDate": "01/22/2018 07:10:07",
      "content": "<p>0.964.</p>",
      "rawMarkdown": "0.964.",
      "votes": null
    },
    {
      "id": "272065",
      "postDate": "01/22/2018 07:17:00",
      "content": "<p>thx for your reply and your share. bangbangda</p>",
      "rawMarkdown": "thx for your reply and your share. bangbangda",
      "votes": null
    },
    {
      "id": "272074",
      "postDate": "01/22/2018 07:54:58",
      "content": "<p>I just finetuned the fc layer. I didn't try any smaller model. SGD.</p>",
      "rawMarkdown": "I just finetuned the fc layer. I didn't try any smaller model. SGD.",
      "votes": null
    },
    {
      "id": "272109",
      "postDate": "01/22/2018 09:34:42",
      "content": "<p>Thanks for your advice and I  agree with   your independent mindset!!</p>",
      "rawMarkdown": "Thanks for your advice and I  agree with   your independent mindset!!",
      "votes": null
    },
    {
      "id": "272223",
      "postDate": "01/22/2018 15:15:06",
      "content": "<p>It looks like I can get 0.85 on a single fine-tuned resnet50 if I use the gleb data. Without the gleb data I have only gotten to 0.72 using a single resnet50.</p>",
      "rawMarkdown": "It looks like I can get 0.85 on a single fine-tuned resnet50 if I use the gleb data. Without the gleb data I have only gotten to 0.72 using a single resnet50.",
      "votes": null
    },
    {
      "id": "272227",
      "postDate": "01/22/2018 15:19:45",
      "content": "<p>Do you augment the data ?</p>",
      "rawMarkdown": "Do you augment the data ?",
      "votes": null
    },
    {
      "id": "272232",
      "postDate": "01/22/2018 15:36:21",
      "content": "<p>I just fit resnet to random 224 x 224 crops. I haven't done any augmentation yet. I will be doing augmentation to predict the _manip images.</p>",
      "rawMarkdown": "I just fit resnet to random 224 x 224 crops. I haven't done any augmentation yet. I will be doing augmentation to predict the _manip images.",
      "votes": null
    },
    {
      "id": "272482",
      "postDate": "01/23/2018 05:40:16",
      "content": "<p>Hi, James, Can I ask you a question? how do you use the Gleb data? Adding them to the SP cup training data or other method?</p>",
      "rawMarkdown": "Hi, James, Can I ask you a question? how do you use the Gleb data? Adding them to the SP cup training data or other method?",
      "votes": null
    },
    {
      "id": "272673",
      "postDate": "01/23/2018 13:15:25",
      "content": "<p>I just combine them with the training data.</p>",
      "rawMarkdown": "I just combine them with the training data.",
      "votes": null
    },
    {
      "id": "273126",
      "postDate": "01/24/2018 07:02:52",
      "content": "<p>Hi. I would like to know how many different models you have used for your ensemble and also which ensemble techniques you have use? Many thanks in advance.</p>",
      "rawMarkdown": "Hi. I would like to know how many different models you have used for your ensemble and also which ensemble techniques you have use? Many thanks in advance.",
      "votes": null
    },
    {
      "id": "273168",
      "postDate": "01/24/2018 08:07:41",
      "content": "<p>A single model got LB score 0.964. I trained many models, such as res50, res101, inceptionResnetV2, SE-ResNext-50 etc. I'm trying to ensemble some models, but I haven't got a score above 0.964. I'm trying to use XGB or train fc layers using feature extracted from the models I have trained.</p>",
      "rawMarkdown": "A single model got LB score 0.964. I trained many models, such as res50, res101, inceptionResnetV2, SE-ResNext-50 etc. I'm trying to ensemble some models, but I haven't got a score above 0.964. I'm trying to use XGB or train fc layers using feature extracted from the models I have trained.",
      "votes": null
    },
    {
      "id": "273261",
      "postDate": "01/24/2018 11:33:30",
      "content": "<p>Thanks for your reply. Just another question: How do you predict for a single input image for a single model during test? Do you use <code>center crop</code>, one <code>random crop</code>, many random crops, <code>tiling</code> or something else. It will a be great help for me if you answer this question too.</p>",
      "rawMarkdown": "Thanks for your reply. Just another question: How do you predict for a single input image for a single model during test? Do you use `center crop`, one `random crop`, many random crops, `tiling` or something else. It will a be great help for me if you answer this question too.",
      "votes": null
    },
    {
      "id": "273336",
      "postDate": "01/24/2018 13:20:35",
      "content": "<p>Center crop.</p>",
      "rawMarkdown": "Center crop.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 264321,
      "author_name": "shunjiading",
      "author_url": "",
      "post_date": "01/02/2018 19:56:18",
      "content": "<p>Hi Yong, thanks for the sharing! I am using similar approach, but my training loss and validation loss were stuck around 0.8 after 500 iters, do you mind share your thoughts in fine tuning optimizer part?</p>",
      "votes": null,
      "replies": [
        {
          "id": 264415,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/03/2018 01:51:21",
          "content": "<p>average_loss: 20 \nlr_policy: \"multistep\" \nbase_lr: 0.0002 \ngamma: 0.5 \nstepvalue: 5000 \nstepvalue: 10000 \nstepvalue: 15000 \nmax_iter:  20000 \ndisplay: 50 \nmomentum: 0.9 \nweight_decay: 0.0005 \nsnapshot: 2500</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 264414,
      "author_name": "youngkl",
      "author_url": "",
      "post_date": "01/03/2018 01:50:55",
      "content": "<p>Data augmentation using the possible processing operations may bring improvements.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 264949,
      "author_name": "dachang",
      "author_url": "",
      "post_date": "01/04/2018 07:59:40",
      "content": "<p>Hi, Young!  I downloaded the train data set and moved the images of 10 models into one folders, therefore I miss the correct form of those labels. Would you please show me the right label names of these 10 models? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 265005,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/04/2018 10:52:02",
          "content": "<p>OK.\n'Motorola-X', 'Motorola-Nexus-6', 'Samsung-Galaxy-S4', 'Samsung-Galaxy-Note3', 'LG-Nexus-5x', \n            'iPhone-4s', 'Motorola-Droid-Maxx', 'HTC-1-M7', 'Sony-NEX-7', 'iPhone-6'</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265278,
          "author_name": "dachang",
          "author_url": "",
          "post_date": "01/05/2018 03:19:01",
          "content": "<p>Thanks a lot!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 265283,
      "author_name": "xiaokangwang",
      "author_url": "",
      "post_date": "01/05/2018 03:36:29",
      "content": "<p>How do you train a neural network using multiple GPUs? Do you have any informative tutorials at hand?</p>",
      "votes": null,
      "replies": [
        {
          "id": 265287,
          "author_name": "yanxiangyi",
          "author_url": "",
          "post_date": "01/05/2018 03:59:16",
          "content": "<p>I think most kinds of deep learning platforms (like Tensorflow and Keras) will support multi-GPU training. The doc will be helpful :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 265284,
      "author_name": "dachang",
      "author_url": "",
      "post_date": "01/05/2018 03:40:43",
      "content": "<p>Hi Young! You mentioned that the tif files were converted to jpg files for testing. How large is your jpeg depression quality in your conversion?</p>",
      "votes": null,
      "replies": [
        {
          "id": 265331,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/05/2018 08:11:21",
          "content": "<p>95.\nHowever, I later found that using tif files to test is better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 265288,
      "author_name": "yanxiangyi",
      "author_url": "",
      "post_date": "01/05/2018 04:02:14",
      "content": "<p>Why converting tiff to jpeg? It may lead to some infomation loss.</p>",
      "votes": null,
      "replies": [
        {
          "id": 265332,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/05/2018 08:11:49",
          "content": "<p>Yes, you are right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 265955,
      "author_name": "youngkl",
      "author_url": "",
      "post_date": "01/07/2018 08:00:52",
      "content": "<p>InceptionResNetV2 may be better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 265957,
          "author_name": "yanxiangyi",
          "author_url": "",
          "post_date": "01/07/2018 08:05:09",
          "content": "<p>Did you just finetune the FC layers or even more? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265977,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/07/2018 10:08:48",
          "content": "<p>Just FC layers now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 266312,
      "author_name": "youngkl",
      "author_url": "",
      "post_date": "01/08/2018 12:49:14",
      "content": "<p>For a single Resnet50 model , you can get LB of around 0.85, depending on how well you train.</p>",
      "votes": null,
      "replies": [
        {
          "id": 271698,
          "author_name": "kuntoro",
          "author_url": "",
          "post_date": "01/21/2018 07:06:10",
          "content": "<p>Excuse me, to obtain 0.85, is it required to fine-tune the whole ResNet50 network? Or can we use feature extractor based on imagenet weights and train only the fully-connected layer?\nHave you tried any smaller model (such as DenseNet121 or MobileNet)? I am running low power GPU (laptop GPU 940MX 2GB).\nBesides, what optimizer do you use (SGD, Adam, etc)?</p>\n\n<p>Thank you very much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272074,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/22/2018 07:54:58",
          "content": "<p>I just finetuned the fc layer. I didn't try any smaller model. SGD.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272223,
          "author_name": "jfkingiii",
          "author_url": "",
          "post_date": "01/22/2018 15:15:06",
          "content": "<p>It looks like I can get 0.85 on a single fine-tuned resnet50 if I use the gleb data. Without the gleb data I have only gotten to 0.72 using a single resnet50.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272227,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/22/2018 15:19:45",
          "content": "<p>Do you augment the data ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272232,
          "author_name": "jfkingiii",
          "author_url": "",
          "post_date": "01/22/2018 15:36:21",
          "content": "<p>I just fit resnet to random 224 x 224 crops. I haven't done any augmentation yet. I will be doing augmentation to predict the _manip images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272482,
          "author_name": "wuzuping",
          "author_url": "",
          "post_date": "01/23/2018 05:40:16",
          "content": "<p>Hi, James, Can I ask you a question? how do you use the Gleb data? Adding them to the SP cup training data or other method?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272673,
          "author_name": "jfkingiii",
          "author_url": "",
          "post_date": "01/23/2018 13:15:25",
          "content": "<p>I just combine them with the training data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273126,
          "author_name": "hamyadlab",
          "author_url": "",
          "post_date": "01/24/2018 07:02:52",
          "content": "<p>Hi. I would like to know how many different models you have used for your ensemble and also which ensemble techniques you have use? Many thanks in advance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273168,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/24/2018 08:07:41",
          "content": "<p>A single model got LB score 0.964. I trained many models, such as res50, res101, inceptionResnetV2, SE-ResNext-50 etc. I'm trying to ensemble some models, but I haven't got a score above 0.964. I'm trying to use XGB or train fc layers using feature extracted from the models I have trained.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273261,
          "author_name": "hamyadlab",
          "author_url": "",
          "post_date": "01/24/2018 11:33:30",
          "content": "<p>Thanks for your reply. Just another question: How do you predict for a single input image for a single model during test? Do you use <code>center crop</code>, one <code>random crop</code>, many random crops, <code>tiling</code> or something else. It will a be great help for me if you answer this question too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 273336,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/24/2018 13:20:35",
          "content": "<p>Center crop.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 266330,
      "author_name": "ywi4ebyrawi",
      "author_url": "",
      "post_date": "01/08/2018 13:59:20",
      "content": "<p>Hi! why did you use such a little batch size?</p>",
      "votes": null,
      "replies": [
        {
          "id": 266576,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/09/2018 04:47:36",
          "content": "<p>This batch size is for a single GPU. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266972,
          "author_name": "ywi4ebyrawi",
          "author_url": "",
          "post_date": "01/10/2018 08:40:51",
          "content": "<p>your gpus doesn't allow to make bigger batches?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 266506,
      "author_name": "igormunizims",
      "author_url": "",
      "post_date": "01/09/2018 01:17:41",
      "content": "<p>First I applied some pre processing (gama correction, jpeg compression, resizing) and cropped all the training data in 512x512 blocks. Then I resized to 256x256 and tried to fine-tune the InceptionV2 model. I do not have a powerful gpu so I'm suffering with the delay of results. It seems like I'm not doing something very different, but I can not achieve more than 0.35 in final test. Any tips? </p>",
      "votes": null,
      "replies": [
        {
          "id": 266543,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/09/2018 03:29:34",
          "content": "<p>I just crop the data while I'm training and testing. You may try not to resize the picture. Train for more epochs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267077,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "01/10/2018 14:33:07",
          "content": "<p>Hi, it is not a good idea to resize the images. The cameras can be identified based on some proprietary color interpolation patterns and sensor noise patterns which can be detected on the raw pixels of the images. Resizing the images alters this information and this is basically why this competition exists. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267243,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/10/2018 22:26:22",
          "content": "<p>Thanks for the comments. You are right, it's better not to resize. But with full images in the training dataset I'll have to use a very small batch size. I'll try train for more epochs too. Thanks again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267363,
          "author_name": "shunjiading",
          "author_url": "",
          "post_date": "01/11/2018 06:31:28",
          "content": "<p>Hi IgorMuniz, I was using the same approach as you. The highest I can get was 0.42. I also did fine-tuning on pre-trained ResNet50. I doubt the difference lays on the method. I think Young was using Caffe for the training and I was implementing things in Keras.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268545,
          "author_name": "bopengiowa",
          "author_url": "",
          "post_date": "01/14/2018 20:30:52",
          "content": "<p>Hi guys,</p>\n\n<p>So am I getting this right that this competition, it's not about what's in the image, but more about the fundamental underlying qualities about the image itself.</p>\n\n<p>Also, how long and what hardware are you guys using to train the model?</p>\n\n<p>Best.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268597,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/15/2018 00:19:11",
          "content": "<p>Hi Shunjia Ding, \nI can see you've improved your acc. I did some improvements using ResNet too (in Keras). Are you guys normalizing the data? I can't find how the input data should be in Keras (RGB or BGR/ normalizing or not) </p>\n\n<p>Bo Peng, \nI do not have a powerfull GPU (gtx 930m), it takes around 1000s/epoch to fine-tuning ResNet50</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268649,
          "author_name": "shunjiading",
          "author_url": "",
          "post_date": "01/15/2018 05:21:16",
          "content": "<p>Hi IgorMuniz, I managed to improve my score. I implemented the random cropping in Keras which was similar to Caffe. The more training examples, the better LB score you would get. In the end, I found it's pretty difficult to minimize the val_loss because I only used center crop for validation set. Learning rate was very important, it took me a while to find a good strategy to get a good learning rate that eventually lowered the loss.</p>\n\n<p>Not only you need a good GPU to train the ResNet50, you also need a SSD for speeding up (image reading). In the end, I ran my python code to generate random patches from the original images up to 30GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270959,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/19/2018 10:36:17",
          "content": "<p>My validation accuracy is stuck around 0.8 even with data augmentation. I keep trying fine tuning Resnet50 (for now it has brought better results). One question...</p>\n\n<p>How many epochs have you trained? I'm using EarlyStopping with patience 10 so maybe I just have to train for more epochs...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270970,
          "author_name": "mxdbld",
          "author_url": "",
          "post_date": "01/19/2018 10:43:55",
          "content": "<p>relaunch the same task you did, disable the early stopping and you will see if you stopped too early or not. It depends of the minibatch size / data augmentation etc... My early stopping is at 100 epochs for example (Not using resnet)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 268755,
      "author_name": "zolo1997",
      "author_url": "",
      "post_date": "01/15/2018 13:38:49",
      "content": "<p>Hello ,Young. I appreciate ur good work ..Also i have some confused questions to ask u ...Did u use the extra dataset from additional data of Gleb's ? If so , which photos did u use to train, just good_jpgs or all photos? And how did u deal with the problem of the different number of each camera?  if u are in convenience , thanks for ur reply...thank u again..</p>",
      "votes": null,
      "replies": [
        {
          "id": 268763,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/15/2018 14:13:49",
          "content": "<p>Yes, I am using the extra data from Gleb to train the model and it has brought big improvements. I use the good_jpgs. I don't take the problem of the different number of each camera into consideration now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268766,
          "author_name": "zolo1997",
          "author_url": "",
          "post_date": "01/15/2018 14:29:04",
          "content": "<p>thanks...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268806,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/15/2018 17:48:46",
          "content": "<p>I'm new here so this question can be kind of stupid... \nIt is allowed to use extra data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268965,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/16/2018 01:43:07",
          "content": "<p>Maybe. I am not sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269017,
          "author_name": "serhiy",
          "author_url": "",
          "post_date": "01/16/2018 04:07:20",
          "content": "<p>In most of Kaggle competitions use of extra data is allowed as long as you post your source in the forum, so it is available for every one. However, in this competition, it is unknown, as no one specifically allowed nor denied use of external data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269212,
          "author_name": "mxdbld",
          "author_url": "",
          "post_date": "01/16/2018 13:30:21",
          "content": "<p>The answer is in the rules : \n\"External Data Use: The following provision supersedes General Rules Section 7.B. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; <strong>provided, you have the right</strong> and authority to use such external data for the purposes of the Competition, and <strong>to share such data with Sponsor and Kaggle</strong> as may be required.\"</p>\n\n<p>So, if you can use and share to the organizers the data, it is allowed</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269220,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/16/2018 13:43:10",
          "content": "<p>Hi Max Diebold, </p>\n\n<p>Where did you find this?\nIn the rules of this competition I found the following:</p>\n\n<p>General Competition Rules\n7. Competition Data\nC. External Data: Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions. Competition Sponsor reserves the right in its sole discretion to disqualify any Participant who Competition Sponsor discovers has undertaken or attempted to undertake the use of data other than the Competition Data, or who uses the Competition Data other than as permitted by the Competition Website and these Rules.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269226,
          "author_name": "mxdbld",
          "author_url": "",
          "post_date": "01/16/2018 13:51:03",
          "content": "<p>In the A. SPECIFIC COMPETITION RULES section of this competition's rules, on this page</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269247,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/16/2018 14:24:08",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269491,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/16/2018 23:25:47",
          "content": "<p>Thanks Young (and to Gleb of course for creating the data set)!</p>\n\n<p>I'd still like to get an official OK from the competition organizers/Kaggle - while Gleb was conscientious enough to filter for only Creative Commons images from Flickr,  there are different varieties of such licenses. I noticed skimming through the images that some of them were categorized under \"non-commercial\" and others under \"non-derivative'. And most of them were under \"attribution required\".</p>\n\n<p>I don't want to be nit-picky but would using it here fall under those use-cases? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269904,
          "author_name": "oujiayu",
          "author_url": "",
          "post_date": "01/17/2018 14:03:41",
          "content": "<p>May I ask u a question? Do u use these extra data? I want to use, but I'm afraid of the rejection of the competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269953,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/17/2018 15:14:15",
          "content": "<p>No, we're not using the Flickr images. We wouldn't mind having more training data though :)\nI would be a little surprised if the use of it is banned but there's no harm waiting for confirmation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270856,
          "author_name": "zolo1997",
          "author_url": "",
          "post_date": "01/19/2018 05:08:28",
          "content": "<p>Hi !  Mr. Lee ..I am confused that if u use extra dataset...Or u just use the original data from official dataset  and use some domain knowledge ?  Of course , i just curious about it .u don't need to show the details ..thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270862,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/19/2018 05:28:01",
          "content": "<p>We're not using the Flickr images. </p>\n\n<p>You do not need to use the extra dataset to achieve a good standing in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271286,
          "author_name": "zolo1997",
          "author_url": "",
          "post_date": "01/20/2018 03:09:39",
          "content": "<p>Hi!  Lee .. u mean that u don't use any extra data which also can get good result..so i just  wonder if u change some structures of network or build a new CNN by yourselves?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271295,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/20/2018 04:02:40",
          "content": "<p>Hi Zhuo Long, perhaps if you detailed your team's approach, we and other competitors could help you improve it? :)</p>\n\n<p>More seriously, we aren't doing anything funky. I'd suggest you search for older image-related Kaggle competitions and read up the winners' solutions that are posted in the forums (e.g., CDiscount). There's a lot of insight and it's a good way of getting more ideas if you're in a rut. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272054,
          "author_name": "dachang",
          "author_url": "",
          "post_date": "01/22/2018 06:42:30",
          "content": "<p>Hi, Chun Ming Lee, I admire your great score in this competition and thanks for your guide above. It is seemed that you have learned a lot from other competitions.  I am new  to images competitions in kaggle. Would you please recommend some more classic image-related Kaggle  competitions where I can learn  some ideas?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272062,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/22/2018 07:07:22",
          "content": "<p>Off the top of my head, I thought the winning solutions for the Carvana and CDiscount competitions were brilliant. You can also search under the Competitions bar for competitions with the tag \"image data\". </p>\n\n<p>I'd also recommend reading the posts of a image \"pro\", Heng CherKeng (<a href=\"https://www.kaggle.com/hengck23\">https://www.kaggle.com/hengck23</a>). He's very generous in detailing his approach and models in the middle of a competition.</p>\n\n<p>Lastly and this is not directed at you, I'd say having a mindset of learning (whether reading research papers, going through past competition solutions etc.) is better long-term than trying to fish for optimal hyper-parameters from others in a competition (especially if you're not sharing anything in return). Again, thank you to Young for being so open with his approach as a top competitor. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272109,
          "author_name": "dachang",
          "author_url": "",
          "post_date": "01/22/2018 09:34:42",
          "content": "<p>Thanks for your advice and I  agree with   your independent mindset!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269962,
      "author_name": "wuzuping",
      "author_url": "",
      "post_date": "01/17/2018 15:28:37",
      "content": "<p>Do you just fine-tune the pre-trained models? Do you use any method that is specified for the cameral problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 270302,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/18/2018 03:11:32",
          "content": "<p>Just finetune the pre-trained models now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270311,
          "author_name": "wuzuping",
          "author_url": "",
          "post_date": "01/18/2018 03:36:45",
          "content": "<p>Thanks, lao ge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270329,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/18/2018 05:09:51",
          "content": "<p>You are welcome, xiongdi.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270825,
      "author_name": "xiaokangwang",
      "author_url": "",
      "post_date": "01/19/2018 03:21:53",
      "content": "<p>How long did it take to train 10000 iters? How many iters in each epoch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 270883,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/19/2018 06:36:28",
          "content": "<p>Sorry, I don't record the time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270861,
      "author_name": "dachang",
      "author_url": "",
      "post_date": "01/19/2018 05:24:41",
      "content": "<p>Hello! Young! Have you tried to manipulate the train data (jpeg compression,gamma correction ,resizing ) to minimize the difference between the train data and the test data? If not, would you please explain the reason why your result comes out so good with the huge difference between the train data and the test data? What is the main factor  your model working so well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 270884,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/19/2018 06:38:02",
          "content": "<p>Yes, I said \"Data augmentation using the possible processing operations may bring improvements\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270909,
      "author_name": "muntakimrafi",
      "author_url": "",
      "post_date": "01/19/2018 07:45:00",
      "content": "<p>Hey Young, </p>\n\n<p>Are you treating the sp cup data and the flickr data from gleb equally? </p>\n\n<p>I am struggling adding the flickr data to our model. </p>\n\n<p>Funny thing is that our model with 90+ test accuracy can recognize the flickr data with 80-90+ acc but don't train on these as expected. </p>\n\n<p>And our model is failing horribly correctly predicting the Sony Nex 7 from Gleb. <br>\nI am confused about whether this is showing the weakness of our model or the random processing these sony images went through before being uploaded to flickr. </p>",
      "votes": null,
      "replies": [
        {
          "id": 270933,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/19/2018 09:32:15",
          "content": "<p>Hi. I treated the sp cup data and the flickr data from gleb equally. May there is a gap between them but I didn't check them. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270946,
      "author_name": "shuhuagao",
      "author_url": "",
      "post_date": "01/19/2018 09:58:35",
      "content": "<p>Thanks for your sharing. May I ask some questions?</p>\n\n<blockquote>\n  <p>I convert the tif files to jpg files for testing.</p>\n</blockquote>\n\n<p>Why did you convert the TIFF into JPG? Besides, did you make some special handling about the manipulated images in the test set? Lastly, why can you have so many GPUs, haha?</p>",
      "votes": null,
      "replies": [
        {
          "id": 270952,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/19/2018 10:24:30",
          "content": "<p>It is not a good idea to convert TIFF to JPG. You can see the discussion in some comments above</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270956,
          "author_name": "shuhuagao",
          "author_url": "",
          "post_date": "01/19/2018 10:32:35",
          "content": "<p>Got it. Thank you, IgorMuniz.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 271976,
      "author_name": "oujiayu",
      "author_url": "",
      "post_date": "01/22/2018 02:03:27",
      "content": "<p>Young, thanks for your sharing, may I ask u that how large is u train dataset. Now I crop the image of train dataset to produce lots of images that aren't overlap, but I think that the dataset is so large and take a lot of time to train. </p>",
      "votes": null,
      "replies": [
        {
          "id": 271992,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/22/2018 03:01:03",
          "content": "<p>I use the eight possible processing operations for data augmentation and the network transforms the data by cropping randomly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271996,
          "author_name": "oujiayu",
          "author_url": "",
          "post_date": "01/22/2018 03:06:30",
          "content": "<p>So it means that u pull the full image into the network and use the random crop? I thought before that put such high resolution image into the network may train very slow, and the batch size should be small. The GPU I use is 1080Ti</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272038,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/22/2018 05:29:07",
          "content": "<p>Yes. You can have a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272040,
          "author_name": "oujiayu",
          "author_url": "",
          "post_date": "01/22/2018 05:33:14",
          "content": "<p>thx and the last question is that may I ask u that what is the best acc u can get by using the single model?Now I use the single model, the best acc is 86.4, and I find it hard to improve.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272063,
          "author_name": "youngkl",
          "author_url": "",
          "post_date": "01/22/2018 07:10:07",
          "content": "<p>0.964.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272065,
          "author_name": "oujiayu",
          "author_url": "",
          "post_date": "01/22/2018 07:17:00",
          "content": "<p>thx for your reply and your share. bangbangda</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "263932": "I just tried to fine-tune the Resnet50 model. I used 8 GPUS  and the batch size is 25. I fine-tuned the model using all the training data for 10000 iters. The model finetued for 7500 iters achieves the best LB now. In the prototxt, I set the crop_size 224.  I convert the tif files to jpg files for testing. \nI think this approach is simple. You can have a try.\nIf you have any advice, please tell me. Thanks!",
    "264321": "Hi Yong, thanks for the sharing! I am using similar approach, but my training loss and validation loss were stuck around 0.8 after 500 iters, do you mind share your thoughts in fine tuning optimizer part?",
    "264414": "Data augmentation using the possible processing operations may bring improvements.",
    "264415": "average_loss: 20 \nlr_policy: \"multistep\" \nbase_lr: 0.0002 \ngamma: 0.5 \nstepvalue: 5000 \nstepvalue: 10000 \nstepvalue: 15000 \nmax_iter:  20000 \ndisplay: 50 \nmomentum: 0.9 \nweight_decay: 0.0005 \nsnapshot: 2500",
    "264949": "Hi, Young!  I downloaded the train data set and moved the images of 10 models into one folders, therefore I miss the correct form of those labels. Would you please show me the right label names of these 10 models? Thanks.",
    "265005": "OK.\n'Motorola-X', 'Motorola-Nexus-6', 'Samsung-Galaxy-S4', 'Samsung-Galaxy-Note3', 'LG-Nexus-5x', \n            'iPhone-4s', 'Motorola-Droid-Maxx', 'HTC-1-M7', 'Sony-NEX-7', 'iPhone-6'",
    "265278": "Thanks a lot!!!",
    "265283": "How do you train a neural network using multiple GPUs? Do you have any informative tutorials at hand?",
    "265284": "Hi Young! You mentioned that the tif files were converted to jpg files for testing. How large is your jpeg depression quality in your conversion?",
    "265287": "I think most kinds of deep learning platforms (like Tensorflow and Keras) will support multi-GPU training. The doc will be helpful :)",
    "265288": "Why converting tiff to jpeg? It may lead to some infomation loss.",
    "265331": "95.\nHowever, I later found that using tif files to test is better.",
    "265332": "Yes, you are right.",
    "265955": "InceptionResNetV2 may be better.",
    "265957": "Did you just finetune the FC layers or even more?",
    "265977": "Just FC layers now.",
    "266312": "For a single Resnet50 model , you can get LB of around 0.85, depending on how well you train.",
    "266330": "Hi! why did you use such a little batch size?",
    "266506": "First I applied some pre processing (gama correction, jpeg compression, resizing) and cropped all the training data in 512x512 blocks. Then I resized to 256x256 and tried to fine-tune the InceptionV2 model. I do not have a powerful gpu so I'm suffering with the delay of results. It seems like I'm not doing something very different, but I can not achieve more than 0.35 in final test. Any tips?",
    "266543": "I just crop the data while I'm training and testing. You may try not to resize the picture. Train for more epochs.",
    "266576": "This batch size is for a single GPU.",
    "266972": "your gpus doesn't allow to make bigger batches?",
    "267077": "Hi, it is not a good idea to resize the images. The cameras can be identified based on some proprietary color interpolation patterns and sensor noise patterns which can be detected on the raw pixels of the images. Resizing the images alters this information and this is basically why this competition exists.",
    "267243": "Thanks for the comments. You are right, it's better not to resize. But with full images in the training dataset I'll have to use a very small batch size. I'll try train for more epochs too. Thanks again.",
    "267363": "Hi IgorMuniz, I was using the same approach as you. The highest I can get was 0.42. I also did fine-tuning on pre-trained ResNet50. I doubt the difference lays on the method. I think Young was using Caffe for the training and I was implementing things in Keras.",
    "268545": "Hi guys,\n\nSo am I getting this right that this competition, it's not about what's in the image, but more about the fundamental underlying qualities about the image itself.\n\nAlso, how long and what hardware are you guys using to train the model?\n\nBest.",
    "268597": "Hi Shunjia Ding, \nI can see you've improved your acc. I did some improvements using ResNet too (in Keras). Are you guys normalizing the data? I can't find how the input data should be in Keras (RGB or BGR/ normalizing or not) \n\nBo Peng, \nI do not have a powerfull GPU (gtx 930m), it takes around 1000s/epoch to fine-tuning ResNet50",
    "268649": "Hi IgorMuniz, I managed to improve my score. I implemented the random cropping in Keras which was similar to Caffe. The more training examples, the better LB score you would get. In the end, I found it's pretty difficult to minimize the val_loss because I only used center crop for validation set. Learning rate was very important, it took me a while to find a good strategy to get a good learning rate that eventually lowered the loss.\n\nNot only you need a good GPU to train the ResNet50, you also need a SSD for speeding up (image reading). In the end, I ran my python code to generate random patches from the original images up to 30GB.",
    "268755": "Hello ,Young. I appreciate ur good work ..Also i have some confused questions to ask u ...Did u use the extra dataset from additional data of Gleb's ? If so , which photos did u use to train, just good_jpgs or all photos? And how did u deal with the problem of the different number of each camera?  if u are in convenience , thanks for ur reply...thank u again..",
    "268763": "Yes, I am using the extra data from Gleb to train the model and it has brought big improvements. I use the good_jpgs. I don't take the problem of the different number of each camera into consideration now.",
    "268766": "thanks...",
    "268806": "I'm new here so this question can be kind of stupid... \nIt is allowed to use extra data?",
    "268965": "Maybe. I am not sure.",
    "269017": "In most of Kaggle competitions use of extra data is allowed as long as you post your source in the forum, so it is available for every one. However, in this competition, it is unknown, as no one specifically allowed nor denied use of external data.",
    "269212": "The answer is in the rules : \n\"External Data Use: The following provision supersedes General Rules Section 7.B. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; **provided, you have the right** and authority to use such external data for the purposes of the Competition, and **to share such data with Sponsor and Kaggle** as may be required.\"\n\nSo, if you can use and share to the organizers the data, it is allowed",
    "269220": "Hi Max Diebold, \n\nWhere did you find this?\nIn the rules of this competition I found the following:\n\nGeneral Competition Rules\n7. Competition Data\nC. External Data: Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions. Competition Sponsor reserves the right in its sole discretion to disqualify any Participant who Competition Sponsor discovers has undertaken or attempted to undertake the use of data other than the Competition Data, or who uses the Competition Data other than as permitted by the Competition Website and these Rules.",
    "269226": "In the A. SPECIFIC COMPETITION RULES section of this competition's rules, on this page",
    "269247": "Thanks!",
    "269491": "Thanks Young (and to Gleb of course for creating the data set)!\n\nI'd still like to get an official OK from the competition organizers/Kaggle - while Gleb was conscientious enough to filter for only Creative Commons images from Flickr,  there are different varieties of such licenses. I noticed skimming through the images that some of them were categorized under \"non-commercial\" and others under \"non-derivative'. And most of them were under \"attribution required\".\n\nI don't want to be nit-picky but would using it here fall under those use-cases?",
    "269904": "May I ask u a question? Do u use these extra data? I want to use, but I'm afraid of the rejection of the competition.",
    "269953": "No, we're not using the Flickr images. We wouldn't mind having more training data though :)\nI would be a little surprised if the use of it is banned but there's no harm waiting for confirmation.",
    "269962": "Do you just fine-tune the pre-trained models? Do you use any method that is specified for the cameral problem?",
    "270302": "Just finetune the pre-trained models now.",
    "270311": "Thanks, lao ge.",
    "270329": "You are welcome, xiongdi.",
    "270825": "How long did it take to train 10000 iters? How many iters in each epoch?",
    "270856": "Hi !  Mr. Lee ..I am confused that if u use extra dataset...Or u just use the original data from official dataset  and use some domain knowledge ?  Of course , i just curious about it .u don't need to show the details ..thanks.",
    "270861": "Hello! Young! Have you tried to manipulate the train data (jpeg compression,gamma correction ,resizing ) to minimize the difference between the train data and the test data? If not, would you please explain the reason why your result comes out so good with the huge difference between the train data and the test data? What is the main factor  your model working so well?",
    "270862": "We're not using the Flickr images. \n\nYou do not need to use the extra dataset to achieve a good standing in this competition.",
    "270883": "Sorry, I don't record the time.",
    "270884": "Yes, I said \"Data augmentation using the possible processing operations may bring improvements\".",
    "270909": "Hey Young, \n\nAre you treating the sp cup data and the flickr data from gleb equally? \n\nI am struggling adding the flickr data to our model. \n\nFunny thing is that our model with 90+ test accuracy can recognize the flickr data with 80-90+ acc but don't train on these as expected. \n\nAnd our model is failing horribly correctly predicting the Sony Nex 7 from Gleb.  \nI am confused about whether this is showing the weakness of our model or the random processing these sony images went through before being uploaded to flickr.",
    "270933": "Hi. I treated the sp cup data and the flickr data from gleb equally. May there is a gap between them but I didn't check them.",
    "270946": "Thanks for your sharing. May I ask some questions?\n\n&gt; I convert the tif files to jpg files for testing.\n\nWhy did you convert the TIFF into JPG? Besides, did you make some special handling about the manipulated images in the test set? Lastly, why can you have so many GPUs, haha?",
    "270952": "It is not a good idea to convert TIFF to JPG. You can see the discussion in some comments above",
    "270956": "Got it. Thank you, IgorMuniz.",
    "270959": "My validation accuracy is stuck around 0.8 even with data augmentation. I keep trying fine tuning Resnet50 (for now it has brought better results). One question...\n\nHow many epochs have you trained? I'm using EarlyStopping with patience 10 so maybe I just have to train for more epochs...",
    "270970": "relaunch the same task you did, disable the early stopping and you will see if you stopped too early or not. It depends of the minibatch size / data augmentation etc... My early stopping is at 100 epochs for example (Not using resnet)",
    "271286": "Hi!  Lee .. u mean that u don't use any extra data which also can get good result..so i just  wonder if u change some structures of network or build a new CNN by yourselves?",
    "271295": "Hi Zhuo Long, perhaps if you detailed your team's approach, we and other competitors could help you improve it? :)\n\nMore seriously, we aren't doing anything funky. I'd suggest you search for older image-related Kaggle competitions and read up the winners' solutions that are posted in the forums (e.g., CDiscount). There's a lot of insight and it's a good way of getting more ideas if you're in a rut.",
    "271698": "Excuse me, to obtain 0.85, is it required to fine-tune the whole ResNet50 network? Or can we use feature extractor based on imagenet weights and train only the fully-connected layer?\nHave you tried any smaller model (such as DenseNet121 or MobileNet)? I am running low power GPU (laptop GPU 940MX 2GB).\nBesides, what optimizer do you use (SGD, Adam, etc)?\n\nThank you very much.",
    "271976": "Young, thanks for your sharing, may I ask u that how large is u train dataset. Now I crop the image of train dataset to produce lots of images that aren't overlap, but I think that the dataset is so large and take a lot of time to train.",
    "271992": "I use the eight possible processing operations for data augmentation and the network transforms the data by cropping randomly.",
    "271996": "So it means that u pull the full image into the network and use the random crop? I thought before that put such high resolution image into the network may train very slow, and the batch size should be small. The GPU I use is 1080Ti",
    "272038": "Yes. You can have a try.",
    "272040": "thx and the last question is that may I ask u that what is the best acc u can get by using the single model?Now I use the single model, the best acc is 86.4, and I find it hard to improve.",
    "272054": "Hi, Chun Ming Lee, I admire your great score in this competition and thanks for your guide above. It is seemed that you have learned a lot from other competitions.  I am new  to images competitions in kaggle. Would you please recommend some more classic image-related Kaggle  competitions where I can learn  some ideas?",
    "272062": "Off the top of my head, I thought the winning solutions for the Carvana and CDiscount competitions were brilliant. You can also search under the Competitions bar for competitions with the tag \"image data\". \n\nI'd also recommend reading the posts of a image \"pro\", Heng CherKeng (https://www.kaggle.com/hengck23). He's very generous in detailing his approach and models in the middle of a competition.\n\nLastly and this is not directed at you, I'd say having a mindset of learning (whether reading research papers, going through past competition solutions etc.) is better long-term than trying to fish for optimal hyper-parameters from others in a competition (especially if you're not sharing anything in return). Again, thank you to Young for being so open with his approach as a top competitor.",
    "272063": "0.964.",
    "272065": "thx for your reply and your share. bangbangda",
    "272074": "I just finetuned the fc layer. I didn't try any smaller model. SGD.",
    "272109": "Thanks for your advice and I  agree with   your independent mindset!!",
    "272223": "It looks like I can get 0.85 on a single fine-tuned resnet50 if I use the gleb data. Without the gleb data I have only gotten to 0.72 using a single resnet50.",
    "272227": "Do you augment the data ?",
    "272232": "I just fit resnet to random 224 x 224 crops. I haven't done any augmentation yet. I will be doing augmentation to predict the _manip images.",
    "272482": "Hi, James, Can I ask you a question? how do you use the Gleb data? Adding them to the SP cup training data or other method?",
    "272673": "I just combine them with the training data.",
    "273126": "Hi. I would like to know how many different models you have used for your ensemble and also which ensemble techniques you have use? Many thanks in advance.",
    "273168": "A single model got LB score 0.964. I trained many models, such as res50, res101, inceptionResnetV2, SE-ResNext-50 etc. I'm trying to ensemble some models, but I haven't got a score above 0.964. I'm trying to use XGB or train fc layers using feature extracted from the models I have trained.",
    "273261": "Thanks for your reply. Just another question: How do you predict for a single input image for a single model during test? Do you use `center crop`, one `random crop`, many random crops, `tiling` or something else. It will a be great help for me if you answer this question too.",
    "273336": "Center crop."
  },
  "source": "meta"
}