{
  "id": 20136,
  "title": "Pre-trained models and external data ",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20136",
  "author_name": "",
  "post_date": "2016-04-14T18:06:04.857Z",
  "votes": 14,
  "comment_count": 26,
  "views": 7375,
  "content": "<p>Thanks for all the questions and requests. We have discussed the use of pre-trained techniques with the host and agreed that the interpretation line between a library vs. a pre-trained model is more ambiguous than we would like. To simplify the task and reduce ambiguity, we have decided to allow usage of external data and pre-trained models.</p>",
  "messages": [
    {
      "id": "114910",
      "postDate": "04/14/2016 18:06:04",
      "content": "<p>Thanks for all the questions and requests. We have discussed the use of pre-trained techniques with the host and agreed that the interpretation line between a library vs. a pre-trained model is more ambiguous than we would like. To simplify the task and reduce ambiguity, we have decided to allow usage of external data and pre-trained models.</p>",
      "rawMarkdown": "Thanks for all the questions and requests. We have discussed the use of pre-trained techniques with the host and agreed that the interpretation line between a library vs. a pre-trained model is more ambiguous than we would like. To simplify the task and reduce ambiguity, we have decided to allow usage of external data and pre-trained models.",
      "votes": null
    },
    {
      "id": "114912",
      "postDate": "04/14/2016 18:16:55",
      "content": "<p>Pre-trained models for a money competition? Don't know if that's a good idea....</p>",
      "rawMarkdown": "Pre-trained models for a money competition? Don't know if that's a good idea....",
      "votes": null
    },
    {
      "id": "114913",
      "postDate": "04/14/2016 18:17:36",
      "content": "<p>exited....</p>",
      "rawMarkdown": "exited....",
      "votes": null
    },
    {
      "id": "114939",
      "postDate": "04/14/2016 21:55:47",
      "content": "<p>I am also against the pre-trained models. I think that you could allow using sth like OpenCV Face Detection Cascades, but not pre-trained models (which are the key in Deep Learning).</p>\n\n<p>Additional, If I can use external data, now I could spend ~2k$ for collecting ~100k images. Then I would be probable to be at top 3. Or I could create a &quot;Super-Face-Detector&quot;, &quot;Super-CellPhone-Detector&quot; etc. using any available dataset.</p>\n\n<p>Please, rethink  you decision </p>\n\n<p>-&gt; Maybe enable pre-trained models for just some libraries but no all possible one.</p>\n\n<p>-&gt; Do not allow the external data (where I mean any possible images, which can be used for training)</p>\n\n<p>-&gt; If you will allow, pre-trained model or/and external data, you should do in the same way like in Yelp: <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification\">https://www.kaggle.com/c/yelp-restaurant-photo-classification</a> I mean exactly:</p>\n\n<p>External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.</p>\n\n<p>To sum-up, I think that allowing pre-trained models and external data will hurt this competition. </p>",
      "rawMarkdown": "I am also against the pre-trained models. I think that you could allow using sth like OpenCV Face Detection Cascades, but not pre-trained models (which are the key in Deep Learning).\r\n\r\nAdditional, If I can use external data, now I could spend ~2k$ for collecting ~100k images. Then I would be probable to be at top 3. Or I could create a \"Super-Face-Detector\", \"Super-CellPhone-Detector\" etc. using any available dataset.\r\n\r\nPlease, rethink  you decision \r\n\r\n-> Maybe enable pre-trained models for just some libraries but no all possible one.\r\n\r\n-> Do not allow the external data (where I mean any possible images, which can be used for training)\r\n\r\n-> If you will allow, pre-trained model or/and external data, you should do in the same way like in Yelp: https://www.kaggle.com/c/yelp-restaurant-photo-classification I mean exactly:\r\n\r\nExternal data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\r\n\r\n\r\nTo sum-up, I think that allowing pre-trained models and external data will hurt this competition.",
      "votes": null
    },
    {
      "id": "114940",
      "postDate": "04/14/2016 22:04:16",
      "content": "<p>@Bartek, </p>\n\n<p>Thanks for the question. We're exactly following the Yelp competition model. <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\">Rules</a> are updated now. I'll post an external data thread in a minute. </p>\n\n<ul>\n<li>Hand labeling images on the test set is strictly forbidden.</li>\n<li>External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.</li>\n</ul>",
      "rawMarkdown": "Bartek, \r\n\r\nThanks for the question. We're exactly following the Yelp competition model. [Rules][1] are updated now. I'll post an external data thread in a minute. \r\n\r\n - Hand labeling images on the test set is strictly forbidden.\r\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules",
      "votes": null
    },
    {
      "id": "114951",
      "postDate": "04/14/2016 23:40:07",
      "content": "<p>This sounds like a good compromise - the previous rules were too fuzzy to really know what was allowed, but the new ones seem fair and clear.</p>\n\n<p>Wendy, one question - is there some time limit within which teams must post pre-trained models they have used? Or could they in theory post their model one minute before the competition finishes?</p>",
      "rawMarkdown": "This sounds like a good compromise - the previous rules were too fuzzy to really know what was allowed, but the new ones seem fair and clear.\r\n\r\nWendy, one question - is there some time limit within which teams must post pre-trained models they have used? Or could they in theory post their model one minute before the competition finishes?",
      "votes": null
    },
    {
      "id": "114952",
      "postDate": "04/15/2016 00:14:22",
      "content": "<p>Well, the 4 hours I've invested (so far) labeling images is probably going to be less important now. Hmph.</p>\n\n<p>There's been enough image processing contests on kaggle that perhaps these discussions can happen up-front with the sponsors rather than post launch? (The questions are always the same.)</p>",
      "rawMarkdown": "Well, the 4 hours I've invested (so far) labeling images is probably going to be less important now. Hmph.\r\n\r\nThere's been enough image processing contests on kaggle that perhaps these discussions can happen up-front with the sponsors rather than post launch? (The questions are always the same.)",
      "votes": null
    },
    {
      "id": "114960",
      "postDate": "04/15/2016 03:46:21",
      "content": "<p>[quote=Wendy Kan;114940]\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\n[/quote]</p>\n\n<p>Not sure I understand this rule. So if I use pre-trained net, I must write about it on forums? After how much time? What if, as Bartek said, I paid $2k for additional data?</p>",
      "rawMarkdown": "[quote=Wendy Kan;114940]\r\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\r\n[/quote]\r\n\r\nNot sure I understand this rule. So if I use pre-trained net, I must write about it on forums? After how much time? What if, as Bartek said, I paid $2k for additional data?",
      "votes": null
    },
    {
      "id": "114968",
      "postDate": "04/15/2016 06:17:40",
      "content": "<p>[quote=Roman Ring;114960]</p>\n\n<p>[quote=Wendy Kan;114940]\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\n[/quote]</p>\n\n<p>Not sure I understand this rule. So if I use pre-trained net, I must write about it on forums? After how much time? What if, as Bartek said, I paid $2k for additional data?</p>\n\n<p>[/quote]</p>\n\n<p>I think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)</p>",
      "rawMarkdown": "[quote=Roman Ring;114960]\r\n\r\n[quote=Wendy Kan;114940]\r\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\r\n[/quote]\r\n\r\nNot sure I understand this rule. So if I use pre-trained net, I must write about it on forums? After how much time? What if, as Bartek said, I paid $2k for additional data?\r\n\r\n[/quote]\r\n\r\nI think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)",
      "votes": null
    },
    {
      "id": "114982",
      "postDate": "04/15/2016 09:08:29",
      "content": "<p>[quote=Neil Slater;114968]</p>\n\n<p>I think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)\n[/quote]</p>\n\n<p>It's not clear at all, there's multitude of ways rules can be bent with the way they're worded right now. </p>\n\n<p>For example, I can collect the additional data, but wait until the very last second before submitting to LB.</p>",
      "rawMarkdown": "[quote=Neil Slater;114968]\r\n\r\nI think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)\r\n[/quote]\r\n\r\nIt's not clear at all, there's multitude of ways rules can be bent with the way they're worded right now. \r\n\r\nFor example, I can collect the additional data, but wait until the very last second before submitting to LB.",
      "votes": null
    },
    {
      "id": "114983",
      "postDate": "04/15/2016 09:23:01",
      "content": "<p>[quote=Roman Ring;114982]</p>\n\n<p>[quote=Neil Slater;114968]</p>\n\n<p>I think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)\n[/quote]</p>\n\n<p>It's not clear at all, there's multitude of ways rules can be bent with the way they're worded right now. </p>\n\n<p>For example, I can collect the additional data, but wait until the very last second before submitting to LB.</p>\n\n<p>[/quote]</p>\n\n<p>You won't get iron-clad rulings on timings, because it will be case-by-case what is reasonable. Clearly there needs to be some lead time, because making use of new data sources or models in this competition takes time and effort. </p>\n\n<p>What you will get is if you skirt too close to the edge and behave badly enough towards other competitors, Kaggle (and/or the competition sponsors) will exercise its right to disqualify your entry. </p>\n\n<p>If you really want, you can take advantage of there not being a strict set of rules, and try playing a game where what you do is in a grey area and the offence to other competitors of you being disqualified &quot;unfairly&quot; would soften Kaggle's stance and merely cause a whole lot of grumpiness in the forums. In that case you could score a financial victory at the expense of goodwill from some other Kaggle users. I don't see that is worth it, but some people like the &quot;win at any cost&quot; approach, so I guess that route is open if you feel you want to gamble a bit of social capital to gain an edge.</p>",
      "rawMarkdown": "[quote=Roman Ring;114982]\r\n\r\n[quote=Neil Slater;114968]\r\n\r\nI think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)\r\n[/quote]\r\n\r\nIt's not clear at all, there's multitude of ways rules can be bent with the way they're worded right now. \r\n\r\nFor example, I can collect the additional data, but wait until the very last second before submitting to LB.\r\n\r\n[/quote]\r\n\r\nYou won't get iron-clad rulings on timings, because it will be case-by-case what is reasonable. Clearly there needs to be some lead time, because making use of new data sources or models in this competition takes time and effort. \r\n\r\nWhat you will get is if you skirt too close to the edge and behave badly enough towards other competitors, Kaggle (and/or the competition sponsors) will exercise its right to disqualify your entry. \r\n\r\nIf you really want, you can take advantage of there not being a strict set of rules, and try playing a game where what you do is in a grey area and the offence to other competitors of you being disqualified \"unfairly\" would soften Kaggle's stance and merely cause a whole lot of grumpiness in the forums. In that case you could score a financial victory at the expense of goodwill from some other Kaggle users. I don't see that is worth it, but some people like the \"win at any cost\" approach, so I guess that route is open if you feel you want to gamble a bit of social capital to gain an edge.",
      "votes": null
    },
    {
      "id": "114985",
      "postDate": "04/15/2016 10:22:58",
      "content": "<p>In my opinion there should be a deadline for informing about using any external-data or pretrained model. Maybe something 1 month before the end of competition.  It would be good for preventing the situation of waiting until the very last second before submitting to LB using any external data.</p>",
      "rawMarkdown": "In my opinion there should be a deadline for informing about using any external-data or pretrained model. Maybe something 1 month before the end of competition.  It would be good for preventing the situation of waiting until the very last second before submitting to LB using any external data.",
      "votes": null
    },
    {
      "id": "114987",
      "postDate": "04/15/2016 10:39:02",
      "content": "<p>In previous competitions, external data <a href=\"https://www.kaggle.com/c/rossmann-store-sales/forums/t/16905/clarification-on-future-external-data/95523#post95523\">had to be posted</a> &quot;within a reasonable timeframe of when you start to use it in your models, and definitely before the new entrants deadline.&quot; And further, &quot;Use common sense (don't post it 2 seconds before midnight UTC) and we'll apply the same common sense on the enforcement side.&quot;</p>",
      "rawMarkdown": "In previous competitions, external data [had to be posted][1] \"within a reasonable timeframe of when you start to use it in your models, and definitely before the new entrants deadline.\" And further, \"Use common sense (don't post it 2 seconds before midnight UTC) and we'll apply the same common sense on the enforcement side.\"\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/rossmann-store-sales/forums/t/16905/clarification-on-future-external-data/95523#post95523",
      "votes": null
    },
    {
      "id": "115011",
      "postDate": "04/15/2016 16:21:49",
      "content": "<p>Yes, as a rule of thumb for posting external data/model use, that <a href=\"https://www.kaggle.com/c/rossmann-store-sales/forums/t/16905/clarification-on-future-external-data/95523#post95523\">post</a> is a good resource. </p>\n\n<p>&quot;Data/model should be posted within a reasonable timeframe of <strong>when you start to use it</strong> in your models, and definitely <strong>before the new entrants deadline</strong>. Use common sense (don't post it 2 seconds before midnight UTC) and we'll apply the same common sense on the enforcement side.&quot;</p>",
      "rawMarkdown": "Yes, as a rule of thumb for posting external data/model use, that [post][1] is a good resource. \r\n\r\n\"Data/model should be posted within a reasonable timeframe of **when you start to use it** in your models, and definitely **before the new entrants deadline**. Use common sense (don't post it 2 seconds before midnight UTC) and we'll apply the same common sense on the enforcement side.\"\r\n\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/rossmann-store-sales/forums/t/16905/clarification-on-future-external-data/95523#post95523",
      "votes": null
    },
    {
      "id": "115035",
      "postDate": "04/15/2016 20:54:47",
      "content": "<p>Need some clarification on the definition of external data. I can clearly see discrepancies in the labels. Especially among the images tagged under driving safely and talking to passenger.  Cleaning the dataset would lead to different image-to-label mappings. Should this be considered as external data?</p>\n\n<p>Is it a good assumption to consider any info/labels used other than what was provided be considered as external data?</p>",
      "rawMarkdown": "Need some clarification on the definition of external data. I can clearly see discrepancies in the labels. Especially among the images tagged under driving safely and talking to passenger.  Cleaning the dataset would lead to different image-to-label mappings. Should this be considered as external data?\r\n\r\nIs it a good assumption to consider any info/labels used other than what was provided be considered as external data?",
      "votes": null
    },
    {
      "id": "115043",
      "postDate": "04/15/2016 21:17:49",
      "content": "<p>@&gt;_&lt;</p>\n\n<p>It's a pretty safe bet that anything you do to the <em>training</em> data is not considered external data. You can crop, label, sub label, add noise, transform, etc., etc., etc., whatever you want.</p>\n\n<p>(Note: You can't manually alter or manually label the test data.)</p>\n\n<p>Removing bad images or re-labeling them (again, on the training set only) is no different that, e.g., removing outliers from numerical training data. So it's fine.</p>\n\n<p>What would be considered external data?  Pretty much anything not provided by the competition, or that isn't generated from the content given in the competition. (This doesn't apply to &quot;common knowledge&quot;. For example, transforming a date feature into work-day/weekend is common knowledge.) Downloading images from the internet is external data. Using a library of facial emotions would be external data. Etc.</p>",
      "rawMarkdown": ">_<\r\n\r\nIt's a pretty safe bet that anything you do to the _training_ data is not considered external data. You can crop, label, sub label, add noise, transform, etc., etc., etc., whatever you want.\r\n\r\n(Note: You can't manually alter or manually label the test data.)\r\n\r\nRemoving bad images or re-labeling them (again, on the training set only) is no different that, e.g., removing outliers from numerical training data. So it's fine.\r\n\r\nWhat would be considered external data?  Pretty much anything not provided by the competition, or that isn't generated from the content given in the competition. (This doesn't apply to \"common knowledge\". For example, transforming a date feature into work-day/weekend is common knowledge.) Downloading images from the internet is external data. Using a library of facial emotions would be external data. Etc.",
      "votes": null
    },
    {
      "id": "115056",
      "postDate": "04/15/2016 23:07:27",
      "content": "<p>What extent can we use the test data?  I can't be bothered to count directly but seem's to be more drivers than in the test set and if we were to create a model and use test images for CV and just not say so, seem's devious but I imagine it'd be a good idea.</p>",
      "rawMarkdown": "What extent can we use the test data?  I can't be bothered to count directly but seem's to be more drivers than in the test set and if we were to create a model and use test images for CV and just not say so, seem's devious but I imagine it'd be a good idea.",
      "votes": null
    },
    {
      "id": "115059",
      "postDate": "04/15/2016 23:28:25",
      "content": "<p>[quote=inversion;114952]</p>\n\n<p>Well, the 4 hours I've invested (so far) labeling images is probably going to be less important now. Hmph.</p>\n\n<p>There's been enough image processing contests on kaggle that perhaps these discussions can happen up-front with the sponsors rather than post launch? (The questions are always the same.)</p>\n\n<p>[/quote]\nwhat do you labeled? I also spend many hours labeling keypoints of a driver , including the left right of many parts of arm and head</p>",
      "rawMarkdown": "[quote=inversion;114952]\r\n\r\nWell, the 4 hours I've invested (so far) labeling images is probably going to be less important now. Hmph.\r\n\r\nThere's been enough image processing contests on kaggle that perhaps these discussions can happen up-front with the sponsors rather than post launch? (The questions are always the same.)\r\n\r\n[/quote]\r\nwhat do you labeled? I also spend many hours labeling keypoints of a driver , including the left right of many parts of arm and head",
      "votes": null
    },
    {
      "id": "115097",
      "postDate": "04/16/2016 07:34:20",
      "content": "<p>[quote=hassiktir;115056]</p>\n\n<p>What extent can we use the test data?  I can't be bothered to count directly but seem's to be more drivers than in the test set and if we were to create a model and use test images for CV and just not say so, seem's devious but I imagine it'd be a good idea.</p>\n\n<p>[/quote]</p>\n\n<p>A semi-supervised run on the test data with auto-encoder or RBM is likely OK (it has been OK in previous competitions), and that might help a little.</p>",
      "rawMarkdown": "[quote=hassiktir;115056]\r\n\r\nWhat extent can we use the test data?  I can't be bothered to count directly but seem's to be more drivers than in the test set and if we were to create a model and use test images for CV and just not say so, seem's devious but I imagine it'd be a good idea.\r\n\r\n[/quote]\r\n\r\nA semi-supervised run on the test data with auto-encoder or RBM is likely OK (it has been OK in previous competitions), and that might help a little.",
      "votes": null
    },
    {
      "id": "115787",
      "postDate": "04/20/2016 04:51:46",
      "content": "<p>Michael, a participant of this competition, downloaded all the data and really spent $2k to label them. After that he sent the labeled data (~100k) to his buddy Lincoln.</p>\n\n<p>Now is the show time of Lincoln. He combined the ImageNet training data and the labeled competition data to train a powerful &#8216;object classification&#8217; ConvNet, and even wrote a paper as the only author to arXiv with title such as &#8216;balabala convolutional neural network for object classification&#8217;. Along with this paper, the powerful pre-trained model has also been released, and was even posted on Caffe model zoo (now publicly available). God, it is not the so called &#8216;object classification pre-trained model&#8217;. It is a distracted driver detection model, but nobody knows that.</p>\n\n<p>Michael received the model and posted the link in the kaggle forum to announce he used it together with VGG19, ResNet, GoogleNet. However, who will really notice that. If someone notice it, who dare to use the model in a zero citation arXiv paper (who is Lincoln, never heard). Nevertheless, with the powerful over fitted model, Michael could get any performance as long as he wants, and all he should concern is not too obvious. Michael got the first place finally, but all the other kagglers think he is suspect. No evidence at all.</p>",
      "rawMarkdown": "Michael, a participant of this competition, downloaded all the data and really spent $2k to label them. After that he sent the labeled data (~100k) to his buddy Lincoln.\r\n\r\nNow is the show time of Lincoln. He combined the ImageNet training data and the labeled competition data to train a powerful ‘object classification’ ConvNet, and even wrote a paper as the only author to arXiv with title such as ‘balabala convolutional neural network for object classification’. Along with this paper, the powerful pre-trained model has also been released, and was even posted on Caffe model zoo (now publicly available). God, it is not the so called ‘object classification pre-trained model’. It is a distracted driver detection model, but nobody knows that.\r\n\r\nMichael received the model and posted the link in the kaggle forum to announce he used it together with VGG19, ResNet, GoogleNet. However, who will really notice that. If someone notice it, who dare to use the model in a zero citation arXiv paper (who is Lincoln, never heard). Nevertheless, with the powerful over fitted model, Michael could get any performance as long as he wants, and all he should concern is not too obvious. Michael got the first place finally, but all the other kagglers think he is suspect. No evidence at all.",
      "votes": null
    },
    {
      "id": "117024",
      "postDate": "04/26/2016 22:46:23",
      "content": "<p>[quote=godknowsall;115787]</p>\n\n<p>Michael [...] Lincoln.</p>\n\n<p>[/quote]</p>\n\n<p>Interesting choice of names.</p>",
      "rawMarkdown": "[quote=godknowsall;115787]\r\n\r\nMichael [...] Lincoln.\r\n\r\n[/quote]\r\n\r\nInteresting choice of names.",
      "votes": null
    },
    {
      "id": "117025",
      "postDate": "04/26/2016 22:54:20",
      "content": "<p>Two questions/requests:</p>\n\n<ol>\n<li>To admins: Will some of the most frequently used pre-trained models be available in Kaggle scripts?</li>\n<li>To other Kagglers: I have no prior experience whatsoever in computer vision, does someone have a link to a quick tutorial that might help me get started? Googling has not helped me.</li>\n</ol>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Two questions/requests:\r\n\r\n 1. To admins: Will some of the most frequently used pre-trained models be available in Kaggle scripts?\r\n 2. To other Kagglers: I have no prior experience whatsoever in computer vision, does someone have a link to a quick tutorial that might help me get started? Googling has not helped me.\r\n\r\nThanks.",
      "votes": null
    },
    {
      "id": "117387",
      "postDate": "04/28/2016 21:02:07",
      "content": "<p><a href=\"https://solderspot.wordpress.com/2014/10/18/using-opencv-for-simple-object-detection/\">https://solderspot.wordpress.com/2014/10/18/using-opencv-for-simple-object-detection/</a>\nif you can detect how much the steering is turning from frame to frame (toyota logo orientation) you might get a useful feature.</p>",
      "rawMarkdown": "https://solderspot.wordpress.com/2014/10/18/using-opencv-for-simple-object-detection/\r\nif you can detect how much the steering is turning from frame to frame (toyota logo orientation) you might get a useful feature.",
      "votes": null
    },
    {
      "id": "119297",
      "postDate": "05/08/2016 23:30:10",
      "content": "<p>@ innerproduct question 2.</p>\n\n<p>Have a look here <a href=\"https://www.kaggle.com/c/facial-keypoints-detection/details/deep-learning-tutorial\">Kaggle Key Point detection deep learning tutoriall</a></p>",
      "rawMarkdown": "innerproduct question 2.\r\n\r\nHave a look here [Kaggle Key Point detection deep learning tutoriall][1]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/facial-keypoints-detection/details/deep-learning-tutorial",
      "votes": null
    },
    {
      "id": "119332",
      "postDate": "05/09/2016 09:13:13",
      "content": "<p>VGG16 and VGG19 available by CC BY-NC 4.0 <a href=\"http://creativecommons.org/licenses/by-nc/4.0/\">http://creativecommons.org/licenses/by-nc/4.0/</a> \nAnd as said by admin are not allowed to be used. Is there any other good Pre-Trained models which have implementation with weights for Keras?</p>",
      "rawMarkdown": "VGG16 and VGG19 available by CC BY-NC 4.0 http://creativecommons.org/licenses/by-nc/4.0/ \r\nAnd as said by admin are not allowed to be used. Is there any other good Pre-Trained models which have implementation with weights for Keras?",
      "votes": null
    },
    {
      "id": "119339",
      "postDate": "05/09/2016 09:35:05",
      "content": "<p>[quote=ZFTurbo;119332]</p>\n\n<p>VGG16 and VGG19 available by CC BY-NC 4.0 <a href=\"http://creativecommons.org/licenses/by-nc/4.0/\">http://creativecommons.org/licenses/by-nc/4.0/</a> \nAnd as said by admin are not allowed to be used. Is there any other good Pre-Trained models which have implementation with weights for Keras?</p>\n\n<p>[/quote]</p>\n\n<p>The discussion is ongoing. VGG models should be fine in general, but something is a bit wrong with the licensing on published conversions to Keras.</p>\n\n<p>As an aside, Creative Commons is designed for open sharing of text and media content, but is an awkward way to license code. I guess it may have been done in this case for attribution clause. I'm not sure if there is a good alternative that would achieve the same goal for the authors.</p>",
      "rawMarkdown": "[quote=ZFTurbo;119332]\r\n\r\nVGG16 and VGG19 available by CC BY-NC 4.0 http://creativecommons.org/licenses/by-nc/4.0/ \r\nAnd as said by admin are not allowed to be used. Is there any other good Pre-Trained models which have implementation with weights for Keras?\r\n\r\n[/quote]\r\n\r\nThe discussion is ongoing. VGG models should be fine in general, but something is a bit wrong with the licensing on published conversions to Keras.\r\n\r\nAs an aside, Creative Commons is designed for open sharing of text and media content, but is an awkward way to license code. I guess it may have been done in this case for attribution clause. I'm not sure if there is a good alternative that would achieve the same goal for the authors.",
      "votes": null
    },
    {
      "id": "658834",
      "postDate": "10/26/2019 15:47:56",
      "content": "<p>Thank you</p>",
      "rawMarkdown": "Thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 114912,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/14/2016 18:16:55",
      "content": "<p>Pre-trained models for a money competition? Don't know if that's a good idea....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114913,
      "author_name": "meanku",
      "author_url": "",
      "post_date": "04/14/2016 18:17:36",
      "content": "<p>exited....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114939,
      "author_name": "melgor",
      "author_url": "",
      "post_date": "04/14/2016 21:55:47",
      "content": "<p>I am also against the pre-trained models. I think that you could allow using sth like OpenCV Face Detection Cascades, but not pre-trained models (which are the key in Deep Learning).</p>\n\n<p>Additional, If I can use external data, now I could spend ~2k$ for collecting ~100k images. Then I would be probable to be at top 3. Or I could create a &quot;Super-Face-Detector&quot;, &quot;Super-CellPhone-Detector&quot; etc. using any available dataset.</p>\n\n<p>Please, rethink  you decision </p>\n\n<p>-&gt; Maybe enable pre-trained models for just some libraries but no all possible one.</p>\n\n<p>-&gt; Do not allow the external data (where I mean any possible images, which can be used for training)</p>\n\n<p>-&gt; If you will allow, pre-trained model or/and external data, you should do in the same way like in Yelp: <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification\">https://www.kaggle.com/c/yelp-restaurant-photo-classification</a> I mean exactly:</p>\n\n<p>External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.</p>\n\n<p>To sum-up, I think that allowing pre-trained models and external data will hurt this competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114940,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "04/14/2016 22:04:16",
      "content": "<p>@Bartek, </p>\n\n<p>Thanks for the question. We're exactly following the Yelp competition model. <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\">Rules</a> are updated now. I'll post an external data thread in a minute. </p>\n\n<ul>\n<li>Hand labeling images on the test set is strictly forbidden.</li>\n<li>External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114951,
      "author_name": "jhoward",
      "author_url": "",
      "post_date": "04/14/2016 23:40:07",
      "content": "<p>This sounds like a good compromise - the previous rules were too fuzzy to really know what was allowed, but the new ones seem fair and clear.</p>\n\n<p>Wendy, one question - is there some time limit within which teams must post pre-trained models they have used? Or could they in theory post their model one minute before the competition finishes?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114952,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/15/2016 00:14:22",
      "content": "<p>Well, the 4 hours I've invested (so far) labeling images is probably going to be less important now. Hmph.</p>\n\n<p>There's been enough image processing contests on kaggle that perhaps these discussions can happen up-front with the sponsors rather than post launch? (The questions are always the same.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114960,
      "author_name": "inoryy",
      "author_url": "",
      "post_date": "04/15/2016 03:46:21",
      "content": "<p>[quote=Wendy Kan;114940]\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\n[/quote]</p>\n\n<p>Not sure I understand this rule. So if I use pre-trained net, I must write about it on forums? After how much time? What if, as Bartek said, I paid $2k for additional data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114968,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/15/2016 06:17:40",
      "content": "<p>[quote=Roman Ring;114960]</p>\n\n<p>[quote=Wendy Kan;114940]\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\n[/quote]</p>\n\n<p>Not sure I understand this rule. So if I use pre-trained net, I must write about it on forums? After how much time? What if, as Bartek said, I paid $2k for additional data?</p>\n\n<p>[/quote]</p>\n\n<p>I think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114982,
      "author_name": "inoryy",
      "author_url": "",
      "post_date": "04/15/2016 09:08:29",
      "content": "<p>[quote=Neil Slater;114968]</p>\n\n<p>I think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)\n[/quote]</p>\n\n<p>It's not clear at all, there's multitude of ways rules can be bent with the way they're worded right now. </p>\n\n<p>For example, I can collect the additional data, but wait until the very last second before submitting to LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114983,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/15/2016 09:23:01",
      "content": "<p>[quote=Roman Ring;114982]</p>\n\n<p>[quote=Neil Slater;114968]</p>\n\n<p>I think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)\n[/quote]</p>\n\n<p>It's not clear at all, there's multitude of ways rules can be bent with the way they're worded right now. </p>\n\n<p>For example, I can collect the additional data, but wait until the very last second before submitting to LB.</p>\n\n<p>[/quote]</p>\n\n<p>You won't get iron-clad rulings on timings, because it will be case-by-case what is reasonable. Clearly there needs to be some lead time, because making use of new data sources or models in this competition takes time and effort. </p>\n\n<p>What you will get is if you skirt too close to the edge and behave badly enough towards other competitors, Kaggle (and/or the competition sponsors) will exercise its right to disqualify your entry. </p>\n\n<p>If you really want, you can take advantage of there not being a strict set of rules, and try playing a game where what you do is in a grey area and the offence to other competitors of you being disqualified &quot;unfairly&quot; would soften Kaggle's stance and merely cause a whole lot of grumpiness in the forums. In that case you could score a financial victory at the expense of goodwill from some other Kaggle users. I don't see that is worth it, but some people like the &quot;win at any cost&quot; approach, so I guess that route is open if you feel you want to gamble a bit of social capital to gain an edge.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114985,
      "author_name": "melgor",
      "author_url": "",
      "post_date": "04/15/2016 10:22:58",
      "content": "<p>In my opinion there should be a deadline for informing about using any external-data or pretrained model. Maybe something 1 month before the end of competition.  It would be good for preventing the situation of waiting until the very last second before submitting to LB using any external data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114987,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/15/2016 10:39:02",
      "content": "<p>In previous competitions, external data <a href=\"https://www.kaggle.com/c/rossmann-store-sales/forums/t/16905/clarification-on-future-external-data/95523#post95523\">had to be posted</a> &quot;within a reasonable timeframe of when you start to use it in your models, and definitely before the new entrants deadline.&quot; And further, &quot;Use common sense (don't post it 2 seconds before midnight UTC) and we'll apply the same common sense on the enforcement side.&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115011,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "04/15/2016 16:21:49",
      "content": "<p>Yes, as a rule of thumb for posting external data/model use, that <a href=\"https://www.kaggle.com/c/rossmann-store-sales/forums/t/16905/clarification-on-future-external-data/95523#post95523\">post</a> is a good resource. </p>\n\n<p>&quot;Data/model should be posted within a reasonable timeframe of <strong>when you start to use it</strong> in your models, and definitely <strong>before the new entrants deadline</strong>. Use common sense (don't post it 2 seconds before midnight UTC) and we'll apply the same common sense on the enforcement side.&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115035,
      "author_name": "lmldev",
      "author_url": "",
      "post_date": "04/15/2016 20:54:47",
      "content": "<p>Need some clarification on the definition of external data. I can clearly see discrepancies in the labels. Especially among the images tagged under driving safely and talking to passenger.  Cleaning the dataset would lead to different image-to-label mappings. Should this be considered as external data?</p>\n\n<p>Is it a good assumption to consider any info/labels used other than what was provided be considered as external data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115043,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/15/2016 21:17:49",
      "content": "<p>@&gt;_&lt;</p>\n\n<p>It's a pretty safe bet that anything you do to the <em>training</em> data is not considered external data. You can crop, label, sub label, add noise, transform, etc., etc., etc., whatever you want.</p>\n\n<p>(Note: You can't manually alter or manually label the test data.)</p>\n\n<p>Removing bad images or re-labeling them (again, on the training set only) is no different that, e.g., removing outliers from numerical training data. So it's fine.</p>\n\n<p>What would be considered external data?  Pretty much anything not provided by the competition, or that isn't generated from the content given in the competition. (This doesn't apply to &quot;common knowledge&quot;. For example, transforming a date feature into work-day/weekend is common knowledge.) Downloading images from the internet is external data. Using a library of facial emotions would be external data. Etc.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115056,
      "author_name": "hassiktir",
      "author_url": "",
      "post_date": "04/15/2016 23:07:27",
      "content": "<p>What extent can we use the test data?  I can't be bothered to count directly but seem's to be more drivers than in the test set and if we were to create a model and use test images for CV and just not say so, seem's devious but I imagine it'd be a good idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115059,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "04/15/2016 23:28:25",
      "content": "<p>[quote=inversion;114952]</p>\n\n<p>Well, the 4 hours I've invested (so far) labeling images is probably going to be less important now. Hmph.</p>\n\n<p>There's been enough image processing contests on kaggle that perhaps these discussions can happen up-front with the sponsors rather than post launch? (The questions are always the same.)</p>\n\n<p>[/quote]\nwhat do you labeled? I also spend many hours labeling keypoints of a driver , including the left right of many parts of arm and head</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115097,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/16/2016 07:34:20",
      "content": "<p>[quote=hassiktir;115056]</p>\n\n<p>What extent can we use the test data?  I can't be bothered to count directly but seem's to be more drivers than in the test set and if we were to create a model and use test images for CV and just not say so, seem's devious but I imagine it'd be a good idea.</p>\n\n<p>[/quote]</p>\n\n<p>A semi-supervised run on the test data with auto-encoder or RBM is likely OK (it has been OK in previous competitions), and that might help a little.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115787,
      "author_name": "godknowsall",
      "author_url": "",
      "post_date": "04/20/2016 04:51:46",
      "content": "<p>Michael, a participant of this competition, downloaded all the data and really spent $2k to label them. After that he sent the labeled data (~100k) to his buddy Lincoln.</p>\n\n<p>Now is the show time of Lincoln. He combined the ImageNet training data and the labeled competition data to train a powerful &#8216;object classification&#8217; ConvNet, and even wrote a paper as the only author to arXiv with title such as &#8216;balabala convolutional neural network for object classification&#8217;. Along with this paper, the powerful pre-trained model has also been released, and was even posted on Caffe model zoo (now publicly available). God, it is not the so called &#8216;object classification pre-trained model&#8217;. It is a distracted driver detection model, but nobody knows that.</p>\n\n<p>Michael received the model and posted the link in the kaggle forum to announce he used it together with VGG19, ResNet, GoogleNet. However, who will really notice that. If someone notice it, who dare to use the model in a zero citation arXiv paper (who is Lincoln, never heard). Nevertheless, with the powerful over fitted model, Michael could get any performance as long as he wants, and all he should concern is not too obvious. Michael got the first place finally, but all the other kagglers think he is suspect. No evidence at all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117024,
      "author_name": "innerproduct",
      "author_url": "",
      "post_date": "04/26/2016 22:46:23",
      "content": "<p>[quote=godknowsall;115787]</p>\n\n<p>Michael [...] Lincoln.</p>\n\n<p>[/quote]</p>\n\n<p>Interesting choice of names.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117025,
      "author_name": "innerproduct",
      "author_url": "",
      "post_date": "04/26/2016 22:54:20",
      "content": "<p>Two questions/requests:</p>\n\n<ol>\n<li>To admins: Will some of the most frequently used pre-trained models be available in Kaggle scripts?</li>\n<li>To other Kagglers: I have no prior experience whatsoever in computer vision, does someone have a link to a quick tutorial that might help me get started? Googling has not helped me.</li>\n</ol>\n\n<p>Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117387,
      "author_name": "harriken",
      "author_url": "",
      "post_date": "04/28/2016 21:02:07",
      "content": "<p><a href=\"https://solderspot.wordpress.com/2014/10/18/using-opencv-for-simple-object-detection/\">https://solderspot.wordpress.com/2014/10/18/using-opencv-for-simple-object-detection/</a>\nif you can detect how much the steering is turning from frame to frame (toyota logo orientation) you might get a useful feature.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119297,
      "author_name": "thibmain",
      "author_url": "",
      "post_date": "05/08/2016 23:30:10",
      "content": "<p>@ innerproduct question 2.</p>\n\n<p>Have a look here <a href=\"https://www.kaggle.com/c/facial-keypoints-detection/details/deep-learning-tutorial\">Kaggle Key Point detection deep learning tutoriall</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119332,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "05/09/2016 09:13:13",
      "content": "<p>VGG16 and VGG19 available by CC BY-NC 4.0 <a href=\"http://creativecommons.org/licenses/by-nc/4.0/\">http://creativecommons.org/licenses/by-nc/4.0/</a> \nAnd as said by admin are not allowed to be used. Is there any other good Pre-Trained models which have implementation with weights for Keras?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119339,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "05/09/2016 09:35:05",
      "content": "<p>[quote=ZFTurbo;119332]</p>\n\n<p>VGG16 and VGG19 available by CC BY-NC 4.0 <a href=\"http://creativecommons.org/licenses/by-nc/4.0/\">http://creativecommons.org/licenses/by-nc/4.0/</a> \nAnd as said by admin are not allowed to be used. Is there any other good Pre-Trained models which have implementation with weights for Keras?</p>\n\n<p>[/quote]</p>\n\n<p>The discussion is ongoing. VGG models should be fine in general, but something is a bit wrong with the licensing on published conversions to Keras.</p>\n\n<p>As an aside, Creative Commons is designed for open sharing of text and media content, but is an awkward way to license code. I guess it may have been done in this case for attribution clause. I'm not sure if there is a good alternative that would achieve the same goal for the authors.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 658834,
      "author_name": "ee15mtech11021",
      "author_url": "",
      "post_date": "10/26/2019 15:47:56",
      "content": "<p>Thank you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "114910": "Thanks for all the questions and requests. We have discussed the use of pre-trained techniques with the host and agreed that the interpretation line between a library vs. a pre-trained model is more ambiguous than we would like. To simplify the task and reduce ambiguity, we have decided to allow usage of external data and pre-trained models.",
    "114912": "Pre-trained models for a money competition? Don't know if that's a good idea....",
    "114913": "exited....",
    "114939": "I am also against the pre-trained models. I think that you could allow using sth like OpenCV Face Detection Cascades, but not pre-trained models (which are the key in Deep Learning).\r\n\r\nAdditional, If I can use external data, now I could spend ~2k$ for collecting ~100k images. Then I would be probable to be at top 3. Or I could create a \"Super-Face-Detector\", \"Super-CellPhone-Detector\" etc. using any available dataset.\r\n\r\nPlease, rethink  you decision \r\n\r\n-> Maybe enable pre-trained models for just some libraries but no all possible one.\r\n\r\n-> Do not allow the external data (where I mean any possible images, which can be used for training)\r\n\r\n-> If you will allow, pre-trained model or/and external data, you should do in the same way like in Yelp: https://www.kaggle.com/c/yelp-restaurant-photo-classification I mean exactly:\r\n\r\nExternal data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\r\n\r\n\r\nTo sum-up, I think that allowing pre-trained models and external data will hurt this competition.",
    "114940": "Bartek, \r\n\r\nThanks for the question. We're exactly following the Yelp competition model. [Rules][1] are updated now. I'll post an external data thread in a minute. \r\n\r\n - Hand labeling images on the test set is strictly forbidden.\r\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules",
    "114951": "This sounds like a good compromise - the previous rules were too fuzzy to really know what was allowed, but the new ones seem fair and clear.\r\n\r\nWendy, one question - is there some time limit within which teams must post pre-trained models they have used? Or could they in theory post their model one minute before the competition finishes?",
    "114952": "Well, the 4 hours I've invested (so far) labeling images is probably going to be less important now. Hmph.\r\n\r\nThere's been enough image processing contests on kaggle that perhaps these discussions can happen up-front with the sponsors rather than post launch? (The questions are always the same.)",
    "114960": "[quote=Wendy Kan;114940]\r\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\r\n[/quote]\r\n\r\nNot sure I understand this rule. So if I use pre-trained net, I must write about it on forums? After how much time? What if, as Bartek said, I paid $2k for additional data?",
    "114968": "[quote=Roman Ring;114960]\r\n\r\n[quote=Wendy Kan;114940]\r\n - External data (including pre-trained model) that is publicly available and free to use, is allowed, but must be posted in the forums for all to see and use.\r\n[/quote]\r\n\r\nNot sure I understand this rule. So if I use pre-trained net, I must write about it on forums? After how much time? What if, as Bartek said, I paid $2k for additional data?\r\n\r\n[/quote]\r\n\r\nI think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)",
    "114982": "[quote=Neil Slater;114968]\r\n\r\nI think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)\r\n[/quote]\r\n\r\nIt's not clear at all, there's multitude of ways rules can be bent with the way they're worded right now. \r\n\r\nFor example, I can collect the additional data, but wait until the very last second before submitting to LB.",
    "114983": "[quote=Roman Ring;114982]\r\n\r\n[quote=Neil Slater;114968]\r\n\r\nI think it is clear from the discussion above and new rules that if you paid $2k for additional data, you would have to make it available online to all competitors (and in fact become open data), and post a link in the forum thread to say where it was. Thank you for your generous public contribution to data science :-)\r\n[/quote]\r\n\r\nIt's not clear at all, there's multitude of ways rules can be bent with the way they're worded right now. \r\n\r\nFor example, I can collect the additional data, but wait until the very last second before submitting to LB.\r\n\r\n[/quote]\r\n\r\nYou won't get iron-clad rulings on timings, because it will be case-by-case what is reasonable. Clearly there needs to be some lead time, because making use of new data sources or models in this competition takes time and effort. \r\n\r\nWhat you will get is if you skirt too close to the edge and behave badly enough towards other competitors, Kaggle (and/or the competition sponsors) will exercise its right to disqualify your entry. \r\n\r\nIf you really want, you can take advantage of there not being a strict set of rules, and try playing a game where what you do is in a grey area and the offence to other competitors of you being disqualified \"unfairly\" would soften Kaggle's stance and merely cause a whole lot of grumpiness in the forums. In that case you could score a financial victory at the expense of goodwill from some other Kaggle users. I don't see that is worth it, but some people like the \"win at any cost\" approach, so I guess that route is open if you feel you want to gamble a bit of social capital to gain an edge.",
    "114985": "In my opinion there should be a deadline for informing about using any external-data or pretrained model. Maybe something 1 month before the end of competition.  It would be good for preventing the situation of waiting until the very last second before submitting to LB using any external data.",
    "114987": "In previous competitions, external data [had to be posted][1] \"within a reasonable timeframe of when you start to use it in your models, and definitely before the new entrants deadline.\" And further, \"Use common sense (don't post it 2 seconds before midnight UTC) and we'll apply the same common sense on the enforcement side.\"\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/rossmann-store-sales/forums/t/16905/clarification-on-future-external-data/95523#post95523",
    "115011": "Yes, as a rule of thumb for posting external data/model use, that [post][1] is a good resource. \r\n\r\n\"Data/model should be posted within a reasonable timeframe of **when you start to use it** in your models, and definitely **before the new entrants deadline**. Use common sense (don't post it 2 seconds before midnight UTC) and we'll apply the same common sense on the enforcement side.\"\r\n\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/rossmann-store-sales/forums/t/16905/clarification-on-future-external-data/95523#post95523",
    "115035": "Need some clarification on the definition of external data. I can clearly see discrepancies in the labels. Especially among the images tagged under driving safely and talking to passenger.  Cleaning the dataset would lead to different image-to-label mappings. Should this be considered as external data?\r\n\r\nIs it a good assumption to consider any info/labels used other than what was provided be considered as external data?",
    "115043": ">_<\r\n\r\nIt's a pretty safe bet that anything you do to the _training_ data is not considered external data. You can crop, label, sub label, add noise, transform, etc., etc., etc., whatever you want.\r\n\r\n(Note: You can't manually alter or manually label the test data.)\r\n\r\nRemoving bad images or re-labeling them (again, on the training set only) is no different that, e.g., removing outliers from numerical training data. So it's fine.\r\n\r\nWhat would be considered external data?  Pretty much anything not provided by the competition, or that isn't generated from the content given in the competition. (This doesn't apply to \"common knowledge\". For example, transforming a date feature into work-day/weekend is common knowledge.) Downloading images from the internet is external data. Using a library of facial emotions would be external data. Etc.",
    "115056": "What extent can we use the test data?  I can't be bothered to count directly but seem's to be more drivers than in the test set and if we were to create a model and use test images for CV and just not say so, seem's devious but I imagine it'd be a good idea.",
    "115059": "[quote=inversion;114952]\r\n\r\nWell, the 4 hours I've invested (so far) labeling images is probably going to be less important now. Hmph.\r\n\r\nThere's been enough image processing contests on kaggle that perhaps these discussions can happen up-front with the sponsors rather than post launch? (The questions are always the same.)\r\n\r\n[/quote]\r\nwhat do you labeled? I also spend many hours labeling keypoints of a driver , including the left right of many parts of arm and head",
    "115097": "[quote=hassiktir;115056]\r\n\r\nWhat extent can we use the test data?  I can't be bothered to count directly but seem's to be more drivers than in the test set and if we were to create a model and use test images for CV and just not say so, seem's devious but I imagine it'd be a good idea.\r\n\r\n[/quote]\r\n\r\nA semi-supervised run on the test data with auto-encoder or RBM is likely OK (it has been OK in previous competitions), and that might help a little.",
    "115787": "Michael, a participant of this competition, downloaded all the data and really spent $2k to label them. After that he sent the labeled data (~100k) to his buddy Lincoln.\r\n\r\nNow is the show time of Lincoln. He combined the ImageNet training data and the labeled competition data to train a powerful ‘object classification’ ConvNet, and even wrote a paper as the only author to arXiv with title such as ‘balabala convolutional neural network for object classification’. Along with this paper, the powerful pre-trained model has also been released, and was even posted on Caffe model zoo (now publicly available). God, it is not the so called ‘object classification pre-trained model’. It is a distracted driver detection model, but nobody knows that.\r\n\r\nMichael received the model and posted the link in the kaggle forum to announce he used it together with VGG19, ResNet, GoogleNet. However, who will really notice that. If someone notice it, who dare to use the model in a zero citation arXiv paper (who is Lincoln, never heard). Nevertheless, with the powerful over fitted model, Michael could get any performance as long as he wants, and all he should concern is not too obvious. Michael got the first place finally, but all the other kagglers think he is suspect. No evidence at all.",
    "117024": "[quote=godknowsall;115787]\r\n\r\nMichael [...] Lincoln.\r\n\r\n[/quote]\r\n\r\nInteresting choice of names.",
    "117025": "Two questions/requests:\r\n\r\n 1. To admins: Will some of the most frequently used pre-trained models be available in Kaggle scripts?\r\n 2. To other Kagglers: I have no prior experience whatsoever in computer vision, does someone have a link to a quick tutorial that might help me get started? Googling has not helped me.\r\n\r\nThanks.",
    "117387": "https://solderspot.wordpress.com/2014/10/18/using-opencv-for-simple-object-detection/\r\nif you can detect how much the steering is turning from frame to frame (toyota logo orientation) you might get a useful feature.",
    "119297": "innerproduct question 2.\r\n\r\nHave a look here [Kaggle Key Point detection deep learning tutoriall][1]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/facial-keypoints-detection/details/deep-learning-tutorial",
    "119332": "VGG16 and VGG19 available by CC BY-NC 4.0 http://creativecommons.org/licenses/by-nc/4.0/ \r\nAnd as said by admin are not allowed to be used. Is there any other good Pre-Trained models which have implementation with weights for Keras?",
    "119339": "[quote=ZFTurbo;119332]\r\n\r\nVGG16 and VGG19 available by CC BY-NC 4.0 http://creativecommons.org/licenses/by-nc/4.0/ \r\nAnd as said by admin are not allowed to be used. Is there any other good Pre-Trained models which have implementation with weights for Keras?\r\n\r\n[/quote]\r\n\r\nThe discussion is ongoing. VGG models should be fine in general, but something is a bit wrong with the licensing on published conversions to Keras.\r\n\r\nAs an aside, Creative Commons is designed for open sharing of text and media content, but is an awkward way to license code. I guess it may have been done in this case for attribution clause. I'm not sure if there is a good alternative that would achieve the same goal for the authors.",
    "658834": "Thank you"
  },
  "source": "meta"
}