{
  "id": 18066,
  "title": "pre-trained nets",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/18066",
  "author_name": "",
  "post_date": "2015-12-21T21:42:04.073Z",
  "votes": null,
  "comment_count": 18,
  "views": 4846,
  "content": "<p>allowed or not?</p>",
  "messages": [
    {
      "id": "102359",
      "postDate": "12/21/2015 21:42:04",
      "content": "<p>allowed or not?</p>",
      "rawMarkdown": "allowed or not?",
      "votes": null
    },
    {
      "id": "102373",
      "postDate": "12/21/2015 23:06:31",
      "content": "<p>External data is not allowed, so, looks clear to me that they are not allowed.</p>",
      "rawMarkdown": "External data is not allowed, so, looks clear to me that they are not allowed.",
      "votes": null
    },
    {
      "id": "102375",
      "postDate": "12/21/2015 23:09:47",
      "content": "<p>[quote=NxGTR;102373]</p>\n\n<p>External data is not allowed, so, looks clear to me that they are not allowed.</p>\n\n<p>[/quote]</p>\n\n<p>you would think so but I've been wrong with assuming this in the past.</p>",
      "rawMarkdown": "[quote=NxGTR;102373]\r\n\r\nExternal data is not allowed, so, looks clear to me that they are not allowed.\r\n\r\n[/quote]\r\n\r\nyou would think so but I've been wrong with assuming this in the past.",
      "votes": null
    },
    {
      "id": "102417",
      "postDate": "12/22/2015 13:43:53",
      "content": "<p>No external data rule prohibits the use of pretrained networks.</p>\n\n<p>I asked same question here:\n<a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17982/pretrained-networks\">https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17982/pretrained-networks</a></p>",
      "rawMarkdown": "No external data rule prohibits the use of pretrained networks.\r\n\r\nI asked same question here:\r\nhttps://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17982/pretrained-networks",
      "votes": null
    },
    {
      "id": "102491",
      "postDate": "12/23/2015 03:13:02",
      "content": "<p>So a competition that's going to basically require a deep network trained from scratch to win... and a no-cash-prize, no-teaming recruitment one at that.  I'm pretty sure Yelp will still get some good candidates, but I doubt there will be that many real competitors here! </p>",
      "rawMarkdown": "So a competition that's going to basically require a deep network trained from scratch to win... and a no-cash-prize, no-teaming recruitment one at that.  I'm pretty sure Yelp will still get some good candidates, but I doubt there will be that many real competitors here!",
      "votes": null
    },
    {
      "id": "102584",
      "postDate": "12/23/2015 18:53:19",
      "content": "<p>@all, </p>\n\n<p>Sorry for the delayed response. It's been a real debate internally whether or not to allow pre-trained models. Eventually it was decided that external data (including pre-trained nets) won't be allowed in this competition. </p>",
      "rawMarkdown": "all, \r\n\r\nSorry for the delayed response. It's been a real debate internally whether or not to allow pre-trained models. Eventually it was decided that external data (including pre-trained nets) won't be allowed in this competition.",
      "votes": null
    },
    {
      "id": "102596",
      "postDate": "12/23/2015 19:51:48",
      "content": "<p>@Wendy, </p>\n\n<p>I think for this problem, don't allow pre-trained network is not very reasonable by following reason:</p>\n\n<ol>\n<li><p>Transfer learning by using pre-trained model is very commonly used in industry, if I remember correctly, Pintrest, Yelp and a lot of companies are using AlexNet in their product pipeline.</p></li>\n<li><p>It will take long time to train a new network, then it becomes a battle of GPU. Eg, with tens of GPU and using MXNet it is able to train Inception-BN network in 10 minutes/epoch, but for most people, train this network with a single GPU will take a day for an epoch. And it is in situation to have the best GPU. In a word, training from scratch is waste power &amp; money and make our earth warmer :)</p></li>\n<li><p>I think in this problem, if use pretrained model, there is still a lot of space to improve the result. This is recruiting competition, prevent using public pretrained model is stopping major players to join. </p></li>\n</ol>\n\n<p>Bing</p>",
      "rawMarkdown": "Wendy, \r\n\r\nI think for this problem, don't allow pre-trained network is not very reasonable by following reason:\r\n\r\n1. Transfer learning by using pre-trained model is very commonly used in industry, if I remember correctly, Pintrest, Yelp and a lot of companies are using AlexNet in their product pipeline.\r\n\r\n2.  It will take long time to train a new network, then it becomes a battle of GPU. Eg, with tens of GPU and using MXNet it is able to train Inception-BN network in 10 minutes/epoch, but for most people, train this network with a single GPU will take a day for an epoch. And it is in situation to have the best GPU. In a word, training from scratch is waste power & money and make our earth warmer :)\r\n\r\n3. I think in this problem, if use pretrained model, there is still a lot of space to improve the result. This is recruiting competition, prevent using public pretrained model is stopping major players to join. \r\n\r\nBing",
      "votes": null
    },
    {
      "id": "102646",
      "postDate": "12/24/2015 05:54:21",
      "content": "<p>@Bing, </p>\n\n<p>This is exactly why the internal discussion took a long time :) All the researchers at Yelp agree that pre-trained models are great for the problem, but allowing pre-trained models means allowing external data, which means a lot more administrative labor to put into this on Kaggle/Yelp to pay close attention to the forum. </p>\n\n<p>However, if you feel very strongly about pre-trained model (using it or not using it), please comment on this thread, and we will review it. I can circle this back again for another discussion, but it will have to be after the holidays. In the mean time, thank you for following the rules and refraining from using pre-trained models. </p>",
      "rawMarkdown": "Bing, \r\n\r\nThis is exactly why the internal discussion took a long time :) All the researchers at Yelp agree that pre-trained models are great for the problem, but allowing pre-trained models means allowing external data, which means a lot more administrative labor to put into this on Kaggle/Yelp to pay close attention to the forum. \r\n\r\nHowever, if you feel very strongly about pre-trained model (using it or not using it), please comment on this thread, and we will review it. I can circle this back again for another discussion, but it will have to be after the holidays. In the mean time, thank you for following the rules and refraining from using pre-trained models.",
      "votes": null
    },
    {
      "id": "102648",
      "postDate": "12/24/2015 06:18:07",
      "content": "<p>@Wendy,</p>\n\n<p>In my view, I think allowing pretrained model won't require more attention:</p>\n\n<ol>\n<li><p>There are only handful pretrained ImageNet models in Caffe, MXNet and TensorFlow. I don't think any other pretrained model will be useful. Or we can constrained it to ImageNet/ILSVRC pretrained model, so it will be less than 10 choice.</p></li>\n<li><p>In NDSB with 200k cash prize, I know a lot of people were running on Amazon AWS. But for this recruiting competition without any money prize, let's compute the cost: Assume we are using g2.8x large, training Inception-BN network will cost 14,203 second per epoch on ILSVRC (MXNet). This dataset is in 1/5 size of ILSVRC, so we can expect to finish one epoch in   0.8 hour. g2.8x large's price is $2.6 per hour, usually we need at least 40 epoch in training ImageNet, so assume it is same to here, train one group parameter will at least cost:  $83. Prediction on 1 million images will also cost some money, let's assume $10. So roughly,  submit once costs $100.</p></li>\n<li><p>This problem is different to ILSVRC image classification, we have to investigate new method, let's assume we are super lucky, in 10 times we find an optimal way for it, then it will cost $1000. Similarly, buying a Titan X will cost $1000. </p></li>\n<li><p>Question is: Who will pay $1000 for getting interview faster? At least I think I won't, although last year when I was seeking job, Yelp rejected my CV without even a phone interview.</p></li>\n</ol>",
      "rawMarkdown": "Wendy,\r\n\r\nIn my view, I think allowing pretrained model won't require more attention:\r\n\r\n1. There are only handful pretrained ImageNet models in Caffe, MXNet and TensorFlow. I don't think any other pretrained model will be useful. Or we can constrained it to ImageNet/ILSVRC pretrained model, so it will be less than 10 choice.\r\n\r\n2. In NDSB with 200k cash prize, I know a lot of people were running on Amazon AWS. But for this recruiting competition without any money prize, let's compute the cost: Assume we are using g2.8x large, training Inception-BN network will cost 14,203 second per epoch on ILSVRC (MXNet). This dataset is in 1/5 size of ILSVRC, so we can expect to finish one epoch in   0.8 hour. g2.8x large's price is $2.6 per hour, usually we need at least 40 epoch in training ImageNet, so assume it is same to here, train one group parameter will at least cost:  $83. Prediction on 1 million images will also cost some money, let's assume $10. So roughly,  submit once costs $100.\r\n\r\n3. This problem is different to ILSVRC image classification, we have to investigate new method, let's assume we are super lucky, in 10 times we find an optimal way for it, then it will cost $1000. Similarly, buying a Titan X will cost $1000. \r\n\r\n4. Question is: Who will pay $1000 for getting interview faster? At least I think I won't, although last year when I was seeking job, Yelp rejected my CV without even a phone interview.",
      "votes": null
    },
    {
      "id": "102650",
      "postDate": "12/24/2015 06:37:18",
      "content": "<p>Agree with Bing: I won't spend 1000$ for an interview either, and Yelp rejected me without even looking at my resume.</p>",
      "rawMarkdown": "Agree with Bing: I won't spend 1000$ for an interview either, and Yelp rejected me without even looking at my resume.",
      "votes": null
    },
    {
      "id": "102681",
      "postDate": "12/24/2015 11:12:45",
      "content": "<p>I am conflicted - on one hand, this contest has a large enough dataset to support training a complex model like Inception from scratch, which is a fairly rare challenge.</p>\n\n<p>I agree with everyone though that approaching this as a real problem it makes no sense not to utilize pretrained models. And this problem is different enough that it probably won't just be a fine-tuning exercise. I think being able to incorporate semantic knowledge available from ImageNet models would  lead to more innovative and interesting approaches than simply training a deep network on the provided data.</p>\n\n<p>Also I agree with Bing that whitelisting a small group of models shouldn't require a lot of effort. </p>\n\n<p>Whatever you decide, please give us a definitive decision as soon as you can after the holidays. While I could enjoy either approach, to have the rules changed mid-contest would be very frustrating.</p>",
      "rawMarkdown": "I am conflicted - on one hand, this contest has a large enough dataset to support training a complex model like Inception from scratch, which is a fairly rare challenge.\r\n\r\nI agree with everyone though that approaching this as a real problem it makes no sense not to utilize pretrained models. And this problem is different enough that it probably won't just be a fine-tuning exercise. I think being able to incorporate semantic knowledge available from ImageNet models would  lead to more innovative and interesting approaches than simply training a deep network on the provided data.\r\n\r\nAlso I agree with Bing that whitelisting a small group of models shouldn't require a lot of effort. \r\n\r\nWhatever you decide, please give us a definitive decision as soon as you can after the holidays. While I could enjoy either approach, to have the rules changed mid-contest would be very frustrating.",
      "votes": null
    },
    {
      "id": "102765",
      "postDate": "12/25/2015 18:23:54",
      "content": "<p>One day soon, pre-trained models will be viewed like the commonly accepted list of stop-words.  Or the use of any off the shelf algorithm (we usually don't make those ourselves either).</p>\n\n<p>When you look at the low level filters learned on any large image conv net, they are quite similar from one net to another.  To me, this says that there is something we &quot;know&quot; about images in general versus something we can derive from a specific problem.  When we get to that point where we can say we &quot;know&quot; something, it is no longer external data...but something we can take for granted.  In any real-world problem, we already take them for granted.</p>\n\n<p>Happy holidays!</p>",
      "rawMarkdown": "One day soon, pre-trained models will be viewed like the commonly accepted list of stop-words.  Or the use of any off the shelf algorithm (we usually don't make those ourselves either).\r\n\r\nWhen you look at the low level filters learned on any large image conv net, they are quite similar from one net to another.  To me, this says that there is something we \"know\" about images in general versus something we can derive from a specific problem.  When we get to that point where we can say we \"know\" something, it is no longer external data...but something we can take for granted.  In any real-world problem, we already take them for granted.\r\n\r\nHappy holidays!",
      "votes": null
    },
    {
      "id": "103098",
      "postDate": "12/28/2015 22:26:09",
      "content": "<p>Any updates on this one?</p>",
      "rawMarkdown": "Any updates on this one?",
      "votes": null
    },
    {
      "id": "103114",
      "postDate": "12/29/2015 01:21:50",
      "content": "<p>We will discuss the issue again after the holiday. In the mean time, you are welcomed to leave your opinions here; we will take all your comments into account when we revisit the issue.</p>\n\n<p>Thanks for your patience and happy holiday!</p>",
      "rawMarkdown": "We will discuss the issue again after the holiday. In the mean time, you are welcomed to leave your opinions here; we will take all your comments into account when we revisit the issue.\r\n\r\nThanks for your patience and happy holiday!",
      "votes": null
    },
    {
      "id": "103322",
      "postDate": "12/31/2015 19:26:42",
      "content": "<p>Thanks for your input everyone! We've decided to allow pre-trained nets.</p>\n\n<p>As Fang-Chieh and Wendy said, we'll have to wait until after the holidays to update all of the rules, but we wanted to make the announcement as early as possible. Please follow the normal instructions for using external data: declare the source so that all contestants have an equal opportunity to use the data.</p>\n\n<p>One note about the contest: our main goal is to discover awesome engineers out there that can help us tackle challenges like this one. We'd love to see the engineering effort you put into the problem, not just a top spot on the leader board.</p>\n\n<p>Happy Holidays, and good luck!</p>",
      "rawMarkdown": "Thanks for your input everyone! We've decided to allow pre-trained nets.\r\n\r\n As Fang-Chieh and Wendy said, we'll have to wait until after the holidays to update all of the rules, but we wanted to make the announcement as early as possible. Please follow the normal instructions for using external data: declare the source so that all contestants have an equal opportunity to use the data.\r\n\r\nOne note about the contest: our main goal is to discover awesome engineers out there that can help us tackle challenges like this one. We'd love to see the engineering effort you put into the problem, not just a top spot on the leader board.\r\n\r\nHappy Holidays, and good luck!",
      "votes": null
    },
    {
      "id": "103385",
      "postDate": "01/02/2016 01:17:00",
      "content": "<p>I presume that if pretrained nets are allowed, (novel) networks trained on the same data are also permitted?</p>",
      "rawMarkdown": "I presume that if pretrained nets are allowed, (novel) networks trained on the same data are also permitted?",
      "votes": null
    },
    {
      "id": "103398",
      "postDate": "01/02/2016 06:10:29",
      "content": "<p>Hi Torgos, thanks for your question.  We'll be reviewing external data on a case-by-case basis, so unfortunately it is not safe to presume that any particular layer, network, or model will be permitted. Our main considerations are 1) licensing (contestants and Yelp are freely able to use the data), 2) fairness (everyone has equal access to the data), 3) esstentialness (we want to see contestants' ingenuity, not scripting around an entire pre-trained model).</p>\n\n<p>I've created a new thread for asking about the use of external data.</p>\n\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hi Torgos, thanks for your question.  We'll be reviewing external data on a case-by-case basis, so unfortunately it is not safe to presume that any particular layer, network, or model will be permitted. Our main considerations are 1) licensing (contestants and Yelp are freely able to use the data), 2) fairness (everyone has equal access to the data), 3) esstentialness (we want to see contestants' ingenuity, not scripting around an entire pre-trained model).\r\n\r\nI've created a new thread for asking about the use of external data.\r\n\r\nHope this helps!",
      "votes": null
    },
    {
      "id": "103399",
      "postDate": "01/02/2016 06:10:49",
      "content": "<p>Hi Torgos, thanks for your question.  We'll be reviewing external data on a case-by-case basis, so unfortunately it is not safe to presume that any particular layer, network, or model will be permitted. Our main considerations are 1) licensing (contestants and Yelp are freely able to use the data), 2) fairness (everyone has equal access to the data), 3) esstentialness (we want to see contestants' ingenuity, not scripting around an entire pre-trained model).</p>\n\n<p>I've created a new thread for asking about the use of external data.</p>\n\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hi Torgos, thanks for your question.  We'll be reviewing external data on a case-by-case basis, so unfortunately it is not safe to presume that any particular layer, network, or model will be permitted. Our main considerations are 1) licensing (contestants and Yelp are freely able to use the data), 2) fairness (everyone has equal access to the data), 3) esstentialness (we want to see contestants' ingenuity, not scripting around an entire pre-trained model).\r\n\r\nI've created a new thread for asking about the use of external data.\r\n\r\nHope this helps!",
      "votes": null
    },
    {
      "id": "107756",
      "postDate": "02/12/2016 09:51:50",
      "content": "<p>Hi!</p>\n\n<p>Can I use this VGG16 pretrained model for Keras?</p>\n\n<p><a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3</a></p>\n\n<p>Kindly regards,\nDmitry</p>",
      "rawMarkdown": "Hi!\r\n\r\nCan I use this VGG16 pretrained model for Keras?\r\n\r\nhttps://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n\r\nKindly regards,\r\nDmitry",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 102373,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "12/21/2015 23:06:31",
      "content": "<p>External data is not allowed, so, looks clear to me that they are not allowed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102375,
      "author_name": "zerozero",
      "author_url": "",
      "post_date": "12/21/2015 23:09:47",
      "content": "<p>[quote=NxGTR;102373]</p>\n\n<p>External data is not allowed, so, looks clear to me that they are not allowed.</p>\n\n<p>[/quote]</p>\n\n<p>you would think so but I've been wrong with assuming this in the past.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102417,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "12/22/2015 13:43:53",
      "content": "<p>No external data rule prohibits the use of pretrained networks.</p>\n\n<p>I asked same question here:\n<a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17982/pretrained-networks\">https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17982/pretrained-networks</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102491,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "12/23/2015 03:13:02",
      "content": "<p>So a competition that's going to basically require a deep network trained from scratch to win... and a no-cash-prize, no-teaming recruitment one at that.  I'm pretty sure Yelp will still get some good candidates, but I doubt there will be that many real competitors here! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102584,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "12/23/2015 18:53:19",
      "content": "<p>@all, </p>\n\n<p>Sorry for the delayed response. It's been a real debate internally whether or not to allow pre-trained models. Eventually it was decided that external data (including pre-trained nets) won't be allowed in this competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102596,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "12/23/2015 19:51:48",
      "content": "<p>@Wendy, </p>\n\n<p>I think for this problem, don't allow pre-trained network is not very reasonable by following reason:</p>\n\n<ol>\n<li><p>Transfer learning by using pre-trained model is very commonly used in industry, if I remember correctly, Pintrest, Yelp and a lot of companies are using AlexNet in their product pipeline.</p></li>\n<li><p>It will take long time to train a new network, then it becomes a battle of GPU. Eg, with tens of GPU and using MXNet it is able to train Inception-BN network in 10 minutes/epoch, but for most people, train this network with a single GPU will take a day for an epoch. And it is in situation to have the best GPU. In a word, training from scratch is waste power &amp; money and make our earth warmer :)</p></li>\n<li><p>I think in this problem, if use pretrained model, there is still a lot of space to improve the result. This is recruiting competition, prevent using public pretrained model is stopping major players to join. </p></li>\n</ol>\n\n<p>Bing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102646,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "12/24/2015 05:54:21",
      "content": "<p>@Bing, </p>\n\n<p>This is exactly why the internal discussion took a long time :) All the researchers at Yelp agree that pre-trained models are great for the problem, but allowing pre-trained models means allowing external data, which means a lot more administrative labor to put into this on Kaggle/Yelp to pay close attention to the forum. </p>\n\n<p>However, if you feel very strongly about pre-trained model (using it or not using it), please comment on this thread, and we will review it. I can circle this back again for another discussion, but it will have to be after the holidays. In the mean time, thank you for following the rules and refraining from using pre-trained models. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102648,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "12/24/2015 06:18:07",
      "content": "<p>@Wendy,</p>\n\n<p>In my view, I think allowing pretrained model won't require more attention:</p>\n\n<ol>\n<li><p>There are only handful pretrained ImageNet models in Caffe, MXNet and TensorFlow. I don't think any other pretrained model will be useful. Or we can constrained it to ImageNet/ILSVRC pretrained model, so it will be less than 10 choice.</p></li>\n<li><p>In NDSB with 200k cash prize, I know a lot of people were running on Amazon AWS. But for this recruiting competition without any money prize, let's compute the cost: Assume we are using g2.8x large, training Inception-BN network will cost 14,203 second per epoch on ILSVRC (MXNet). This dataset is in 1/5 size of ILSVRC, so we can expect to finish one epoch in   0.8 hour. g2.8x large's price is $2.6 per hour, usually we need at least 40 epoch in training ImageNet, so assume it is same to here, train one group parameter will at least cost:  $83. Prediction on 1 million images will also cost some money, let's assume $10. So roughly,  submit once costs $100.</p></li>\n<li><p>This problem is different to ILSVRC image classification, we have to investigate new method, let's assume we are super lucky, in 10 times we find an optimal way for it, then it will cost $1000. Similarly, buying a Titan X will cost $1000. </p></li>\n<li><p>Question is: Who will pay $1000 for getting interview faster? At least I think I won't, although last year when I was seeking job, Yelp rejected my CV without even a phone interview.</p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102650,
      "author_name": "phunter",
      "author_url": "",
      "post_date": "12/24/2015 06:37:18",
      "content": "<p>Agree with Bing: I won't spend 1000$ for an interview either, and Yelp rejected me without even looking at my resume.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102681,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "12/24/2015 11:12:45",
      "content": "<p>I am conflicted - on one hand, this contest has a large enough dataset to support training a complex model like Inception from scratch, which is a fairly rare challenge.</p>\n\n<p>I agree with everyone though that approaching this as a real problem it makes no sense not to utilize pretrained models. And this problem is different enough that it probably won't just be a fine-tuning exercise. I think being able to incorporate semantic knowledge available from ImageNet models would  lead to more innovative and interesting approaches than simply training a deep network on the provided data.</p>\n\n<p>Also I agree with Bing that whitelisting a small group of models shouldn't require a lot of effort. </p>\n\n<p>Whatever you decide, please give us a definitive decision as soon as you can after the holidays. While I could enjoy either approach, to have the rules changed mid-contest would be very frustrating.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102765,
      "author_name": "zerozero",
      "author_url": "",
      "post_date": "12/25/2015 18:23:54",
      "content": "<p>One day soon, pre-trained models will be viewed like the commonly accepted list of stop-words.  Or the use of any off the shelf algorithm (we usually don't make those ourselves either).</p>\n\n<p>When you look at the low level filters learned on any large image conv net, they are quite similar from one net to another.  To me, this says that there is something we &quot;know&quot; about images in general versus something we can derive from a specific problem.  When we get to that point where we can say we &quot;know&quot; something, it is no longer external data...but something we can take for granted.  In any real-world problem, we already take them for granted.</p>\n\n<p>Happy holidays!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103098,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "12/28/2015 22:26:09",
      "content": "<p>Any updates on this one?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103114,
      "author_name": "fcchou",
      "author_url": "",
      "post_date": "12/29/2015 01:21:50",
      "content": "<p>We will discuss the issue again after the holiday. In the mean time, you are welcomed to leave your opinions here; we will take all your comments into account when we revisit the issue.</p>\n\n<p>Thanks for your patience and happy holiday!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103322,
      "author_name": "",
      "author_url": "",
      "post_date": "12/31/2015 19:26:42",
      "content": "<p>Thanks for your input everyone! We've decided to allow pre-trained nets.</p>\n\n<p>As Fang-Chieh and Wendy said, we'll have to wait until after the holidays to update all of the rules, but we wanted to make the announcement as early as possible. Please follow the normal instructions for using external data: declare the source so that all contestants have an equal opportunity to use the data.</p>\n\n<p>One note about the contest: our main goal is to discover awesome engineers out there that can help us tackle challenges like this one. We'd love to see the engineering effort you put into the problem, not just a top spot on the leader board.</p>\n\n<p>Happy Holidays, and good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103385,
      "author_name": "telser",
      "author_url": "",
      "post_date": "01/02/2016 01:17:00",
      "content": "<p>I presume that if pretrained nets are allowed, (novel) networks trained on the same data are also permitted?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103398,
      "author_name": "",
      "author_url": "",
      "post_date": "01/02/2016 06:10:29",
      "content": "<p>Hi Torgos, thanks for your question.  We'll be reviewing external data on a case-by-case basis, so unfortunately it is not safe to presume that any particular layer, network, or model will be permitted. Our main considerations are 1) licensing (contestants and Yelp are freely able to use the data), 2) fairness (everyone has equal access to the data), 3) esstentialness (we want to see contestants' ingenuity, not scripting around an entire pre-trained model).</p>\n\n<p>I've created a new thread for asking about the use of external data.</p>\n\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103399,
      "author_name": "",
      "author_url": "",
      "post_date": "01/02/2016 06:10:49",
      "content": "<p>Hi Torgos, thanks for your question.  We'll be reviewing external data on a case-by-case basis, so unfortunately it is not safe to presume that any particular layer, network, or model will be permitted. Our main considerations are 1) licensing (contestants and Yelp are freely able to use the data), 2) fairness (everyone has equal access to the data), 3) esstentialness (we want to see contestants' ingenuity, not scripting around an entire pre-trained model).</p>\n\n<p>I've created a new thread for asking about the use of external data.</p>\n\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107756,
      "author_name": "deaddy",
      "author_url": "",
      "post_date": "02/12/2016 09:51:50",
      "content": "<p>Hi!</p>\n\n<p>Can I use this VGG16 pretrained model for Keras?</p>\n\n<p><a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3</a></p>\n\n<p>Kindly regards,\nDmitry</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "102359": "allowed or not?",
    "102373": "External data is not allowed, so, looks clear to me that they are not allowed.",
    "102375": "[quote=NxGTR;102373]\r\n\r\nExternal data is not allowed, so, looks clear to me that they are not allowed.\r\n\r\n[/quote]\r\n\r\nyou would think so but I've been wrong with assuming this in the past.",
    "102417": "No external data rule prohibits the use of pretrained networks.\r\n\r\nI asked same question here:\r\nhttps://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17982/pretrained-networks",
    "102491": "So a competition that's going to basically require a deep network trained from scratch to win... and a no-cash-prize, no-teaming recruitment one at that.  I'm pretty sure Yelp will still get some good candidates, but I doubt there will be that many real competitors here!",
    "102584": "all, \r\n\r\nSorry for the delayed response. It's been a real debate internally whether or not to allow pre-trained models. Eventually it was decided that external data (including pre-trained nets) won't be allowed in this competition.",
    "102596": "Wendy, \r\n\r\nI think for this problem, don't allow pre-trained network is not very reasonable by following reason:\r\n\r\n1. Transfer learning by using pre-trained model is very commonly used in industry, if I remember correctly, Pintrest, Yelp and a lot of companies are using AlexNet in their product pipeline.\r\n\r\n2.  It will take long time to train a new network, then it becomes a battle of GPU. Eg, with tens of GPU and using MXNet it is able to train Inception-BN network in 10 minutes/epoch, but for most people, train this network with a single GPU will take a day for an epoch. And it is in situation to have the best GPU. In a word, training from scratch is waste power & money and make our earth warmer :)\r\n\r\n3. I think in this problem, if use pretrained model, there is still a lot of space to improve the result. This is recruiting competition, prevent using public pretrained model is stopping major players to join. \r\n\r\nBing",
    "102646": "Bing, \r\n\r\nThis is exactly why the internal discussion took a long time :) All the researchers at Yelp agree that pre-trained models are great for the problem, but allowing pre-trained models means allowing external data, which means a lot more administrative labor to put into this on Kaggle/Yelp to pay close attention to the forum. \r\n\r\nHowever, if you feel very strongly about pre-trained model (using it or not using it), please comment on this thread, and we will review it. I can circle this back again for another discussion, but it will have to be after the holidays. In the mean time, thank you for following the rules and refraining from using pre-trained models.",
    "102648": "Wendy,\r\n\r\nIn my view, I think allowing pretrained model won't require more attention:\r\n\r\n1. There are only handful pretrained ImageNet models in Caffe, MXNet and TensorFlow. I don't think any other pretrained model will be useful. Or we can constrained it to ImageNet/ILSVRC pretrained model, so it will be less than 10 choice.\r\n\r\n2. In NDSB with 200k cash prize, I know a lot of people were running on Amazon AWS. But for this recruiting competition without any money prize, let's compute the cost: Assume we are using g2.8x large, training Inception-BN network will cost 14,203 second per epoch on ILSVRC (MXNet). This dataset is in 1/5 size of ILSVRC, so we can expect to finish one epoch in   0.8 hour. g2.8x large's price is $2.6 per hour, usually we need at least 40 epoch in training ImageNet, so assume it is same to here, train one group parameter will at least cost:  $83. Prediction on 1 million images will also cost some money, let's assume $10. So roughly,  submit once costs $100.\r\n\r\n3. This problem is different to ILSVRC image classification, we have to investigate new method, let's assume we are super lucky, in 10 times we find an optimal way for it, then it will cost $1000. Similarly, buying a Titan X will cost $1000. \r\n\r\n4. Question is: Who will pay $1000 for getting interview faster? At least I think I won't, although last year when I was seeking job, Yelp rejected my CV without even a phone interview.",
    "102650": "Agree with Bing: I won't spend 1000$ for an interview either, and Yelp rejected me without even looking at my resume.",
    "102681": "I am conflicted - on one hand, this contest has a large enough dataset to support training a complex model like Inception from scratch, which is a fairly rare challenge.\r\n\r\nI agree with everyone though that approaching this as a real problem it makes no sense not to utilize pretrained models. And this problem is different enough that it probably won't just be a fine-tuning exercise. I think being able to incorporate semantic knowledge available from ImageNet models would  lead to more innovative and interesting approaches than simply training a deep network on the provided data.\r\n\r\nAlso I agree with Bing that whitelisting a small group of models shouldn't require a lot of effort. \r\n\r\nWhatever you decide, please give us a definitive decision as soon as you can after the holidays. While I could enjoy either approach, to have the rules changed mid-contest would be very frustrating.",
    "102765": "One day soon, pre-trained models will be viewed like the commonly accepted list of stop-words.  Or the use of any off the shelf algorithm (we usually don't make those ourselves either).\r\n\r\nWhen you look at the low level filters learned on any large image conv net, they are quite similar from one net to another.  To me, this says that there is something we \"know\" about images in general versus something we can derive from a specific problem.  When we get to that point where we can say we \"know\" something, it is no longer external data...but something we can take for granted.  In any real-world problem, we already take them for granted.\r\n\r\nHappy holidays!",
    "103098": "Any updates on this one?",
    "103114": "We will discuss the issue again after the holiday. In the mean time, you are welcomed to leave your opinions here; we will take all your comments into account when we revisit the issue.\r\n\r\nThanks for your patience and happy holiday!",
    "103322": "Thanks for your input everyone! We've decided to allow pre-trained nets.\r\n\r\n As Fang-Chieh and Wendy said, we'll have to wait until after the holidays to update all of the rules, but we wanted to make the announcement as early as possible. Please follow the normal instructions for using external data: declare the source so that all contestants have an equal opportunity to use the data.\r\n\r\nOne note about the contest: our main goal is to discover awesome engineers out there that can help us tackle challenges like this one. We'd love to see the engineering effort you put into the problem, not just a top spot on the leader board.\r\n\r\nHappy Holidays, and good luck!",
    "103385": "I presume that if pretrained nets are allowed, (novel) networks trained on the same data are also permitted?",
    "103398": "Hi Torgos, thanks for your question.  We'll be reviewing external data on a case-by-case basis, so unfortunately it is not safe to presume that any particular layer, network, or model will be permitted. Our main considerations are 1) licensing (contestants and Yelp are freely able to use the data), 2) fairness (everyone has equal access to the data), 3) esstentialness (we want to see contestants' ingenuity, not scripting around an entire pre-trained model).\r\n\r\nI've created a new thread for asking about the use of external data.\r\n\r\nHope this helps!",
    "103399": "Hi Torgos, thanks for your question.  We'll be reviewing external data on a case-by-case basis, so unfortunately it is not safe to presume that any particular layer, network, or model will be permitted. Our main considerations are 1) licensing (contestants and Yelp are freely able to use the data), 2) fairness (everyone has equal access to the data), 3) esstentialness (we want to see contestants' ingenuity, not scripting around an entire pre-trained model).\r\n\r\nI've created a new thread for asking about the use of external data.\r\n\r\nHope this helps!",
    "107756": "Hi!\r\n\r\nCan I use this VGG16 pretrained model for Keras?\r\n\r\nhttps://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n\r\nKindly regards,\r\nDmitry"
  },
  "source": "meta"
}