{
  "id": 20759,
  "title": "External data thread",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/20759",
  "author_name": "",
  "post_date": "2016-05-06T07:31:02.073Z",
  "votes": null,
  "comment_count": 31,
  "views": 5128,
  "content": "<p>Since, external data is allowed, please share it here before you use them.. :)</p>",
  "messages": [
    {
      "id": "118932",
      "postDate": "05/06/2016 07:31:02",
      "content": "<p>Since, external data is allowed, please share it here before you use them.. :)</p>",
      "rawMarkdown": "Since, external data is allowed, please share it here before you use them.. :)",
      "votes": null
    },
    {
      "id": "118959",
      "postDate": "05/06/2016 11:04:49",
      "content": "<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>",
      "rawMarkdown": "I will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread",
      "votes": null
    },
    {
      "id": "118961",
      "postDate": "05/06/2016 11:29:44",
      "content": "<p>[quote=DataGeek;118959]</p>\n\n<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>\n\n<p>[/quote]</p>\n\n<p>I think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com</p>",
      "rawMarkdown": "[quote=DataGeek;118959]\r\n\r\nI will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\r\n\r\n[/quote]\r\n\r\nI think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com",
      "votes": null
    },
    {
      "id": "118963",
      "postDate": "05/06/2016 11:47:10",
      "content": "<p>[quote=Abhishek;118961]</p>\n\n<p>[quote=DataGeek;118959]</p>\n\n<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>\n\n<p>[/quote]</p>\n\n<p>I think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com</p>\n\n<p>[/quote]</p>\n\n<p>that's why I gave Kaggle link and didn't say &quot;I'll use the pretrained data accessible to everyone on www.google.com&quot;</p>",
      "rawMarkdown": "[quote=Abhishek;118961]\r\n\r\n[quote=DataGeek;118959]\r\n\r\nI will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\r\n\r\n[/quote]\r\n\r\nI think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com\r\n\r\n[/quote]\r\n\r\nthat's why I gave Kaggle link and didn't say \"I'll use the pretrained data accessible to everyone on www.google.com\"",
      "votes": null
    },
    {
      "id": "119038",
      "postDate": "05/06/2016 21:01:46",
      "content": "<p>[quote=DataGeek;118959]</p>\n\n<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>\n\n<p>[/quote]</p>\n\n<p>In my opinion this is not specific enough.</p>",
      "rawMarkdown": "[quote=DataGeek;118959]\r\n\r\nI will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\r\n\r\n[/quote]\r\n\r\nIn my opinion this is not specific enough.",
      "votes": null
    },
    {
      "id": "119064",
      "postDate": "05/07/2016 00:07:17",
      "content": "<p>Thinking that I will probably have a play with the WordNet Lexical DB\n<a href=\"http://wordnet.princeton.edu/\">http://wordnet.princeton.edu/</a> </p>",
      "rawMarkdown": "Thinking that I will probably have a play with the WordNet Lexical DB\r\nhttp://wordnet.princeton.edu/",
      "votes": null
    },
    {
      "id": "120953",
      "postDate": "05/22/2016 07:19:55",
      "content": "<p>We are using the russian stopwords from the NLTK corpus:</p>\n\n<p>In python:</p>\n\n<pre><code>from nltk.corpus import stopwords\nstop_words = stopwords.words('russian')\n</code></pre>",
      "rawMarkdown": "We are using the russian stopwords from the NLTK corpus:\r\n\r\nIn python:\r\n\r\n    from nltk.corpus import stopwords\r\n    stop_words = stopwords.words('russian')",
      "votes": null
    },
    {
      "id": "121852",
      "postDate": "05/30/2016 11:03:51",
      "content": "<p>Chris, there is an open <a href=\"http://globalwordnet.org/wordnets-in-the-world/\">Russian wordnet database</a>. Have you been able to use it?</p>\n\n<p>Neither <code>wn</code>, nor <code>wnb</code> (the UI), seem to find whatever words I throw at it. Not sure if I installed the database properly...</p>",
      "rawMarkdown": "Chris, there is an open [Russian wordnet database][1]. Have you been able to use it?\r\n\r\nNeither `wn`, nor `wnb` (the UI), seem to find whatever words I throw at it. Not sure if I installed the database properly...\r\n\r\n\r\n  [1]: http://globalwordnet.org/wordnets-in-the-world/",
      "votes": null
    },
    {
      "id": "122783",
      "postDate": "06/07/2016 05:24:38",
      "content": "<p>[quote=Abhishek;118961]</p>\n\n<p>[quote=DataGeek;118959]</p>\n\n<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>\n\n<p>[/quote]</p>\n\n<p>I think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "[quote=Abhishek;118961]\r\n\r\n[quote=DataGeek;118959]\r\n\r\nI will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\r\n\r\n[/quote]\r\n\r\nI think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "122784",
      "postDate": "06/07/2016 05:26:04",
      "content": "<p>hi i am a student</p>",
      "rawMarkdown": "hi i am a student",
      "votes": null
    },
    {
      "id": "122891",
      "postDate": "06/08/2016 03:08:12",
      "content": "<p>I will be using Lucene's default list of Russian stop words. (<a href=\"http://lucene.apache.org/\">http://lucene.apache.org</a>)</p>",
      "rawMarkdown": "I will be using Lucene's default list of Russian stop words. (http://lucene.apache.org)",
      "votes": null
    },
    {
      "id": "123564",
      "postDate": "06/12/2016 18:35:34",
      "content": "<p>We are using pretrained models from this page:</p>\n\n<p><a href=\"http://ling.go.mail.ru/misc/dialogue_2015.html\">http://ling.go.mail.ru/misc/dialogue_2015.html</a></p>",
      "rawMarkdown": "We are using pretrained models from this page:\r\n\r\nhttp://ling.go.mail.ru/misc/dialogue_2015.html",
      "votes": null
    },
    {
      "id": "123663",
      "postDate": "06/13/2016 11:08:26",
      "content": "<p>Pretrained models from: <br>\n<a href=\"https://github.com/dmlc/mxnet-model-gallery/\">https://github.com/dmlc/mxnet-model-gallery/</a></p>",
      "rawMarkdown": "Pretrained models from:  \r\nhttps://github.com/dmlc/mxnet-model-gallery/",
      "votes": null
    },
    {
      "id": "124176",
      "postDate": "06/16/2016 01:26:52",
      "content": "<p>Pretrained models from: <br>\n<a href=\"https://github.com/facebook/fb.resnet.torch\">https://github.com/facebook/fb.resnet.torch</a></p>",
      "rawMarkdown": "Pretrained models from:  \r\nhttps://github.com/facebook/fb.resnet.torch",
      "votes": null
    },
    {
      "id": "124200",
      "postDate": "06/16/2016 07:14:57",
      "content": "<p>We are using <code>russian.pickle</code> from the NLTK train_punkt tokenizers:\n<a href=\"https://github.com/mhq/train_punkt\">https://github.com/mhq/train_punkt</a></p>",
      "rawMarkdown": "We are using `russian.pickle` from the NLTK train_punkt tokenizers:\r\nhttps://github.com/mhq/train_punkt",
      "votes": null
    },
    {
      "id": "124349",
      "postDate": "06/17/2016 12:08:46",
      "content": "<p>[quote=anokas;124200]</p>\n\n<p>We are using <code>russian.pickle</code> from the NLTK train_punkt tokenizers:\n<a href=\"https://github.com/mhq/train_punkt\">https://github.com/mhq/train_punkt</a></p>\n\n<p>[/quote]</p>\n\n<p>I thought that gave me an error when I had tried it. I tried it again. cPickle.UnpicklingError: invalid load key, '. Did you get any such error ?</p>\n\n<p>Regds\nDeb</p>",
      "rawMarkdown": "[quote=anokas;124200]\r\n\r\nWe are using `russian.pickle` from the NLTK train_punkt tokenizers:\r\nhttps://github.com/mhq/train_punkt\r\n\r\n[/quote]\r\n\r\nI thought that gave me an error when I had tried it. I tried it again. cPickle.UnpicklingError: invalid load key, '. Did you get any such error ?\r\n\r\nRegds\r\nDeb",
      "votes": null
    },
    {
      "id": "124350",
      "postDate": "06/17/2016 12:09:44",
      "content": "<p>[quote=Gerard Toonstra;123564]</p>\n\n<p>We are using pretrained models from this page:\n<a href=\"http://ling.go.mail.ru/misc/dialogue_2015.html\">http://ling.go.mail.ru/misc/dialogue_2015.html</a></p>\n\n<p>[/quote]\nUs too</p>",
      "rawMarkdown": "[quote=Gerard Toonstra;123564]\r\n\r\nWe are using pretrained models from this page:\r\nhttp://ling.go.mail.ru/misc/dialogue_2015.html\r\n\r\n[/quote]\r\nUs too",
      "votes": null
    },
    {
      "id": "124387",
      "postDate": "06/17/2016 23:54:17",
      "content": "<p>Trained models from OpenCV: <br>\n<a href=\"https://github.com/Itseez/opencv/tree/master/data\">https://github.com/Itseez/opencv/tree/master/data</a> <br>\nfrom dlib: <br>\n<a href=\"https://github.com/davisking/dlib/blob/master/dlib/image_processing/frontal_face_detector.h\">https://github.com/davisking/dlib/blob/master/dlib/image_processing/frontal_face_detector.h</a> <br>\nOpenFace: <br>\n<a href=\"https://github.com/cmusatyalab/openface\">https://github.com/cmusatyalab/openface</a></p>",
      "rawMarkdown": "Trained models from OpenCV:  \r\nhttps://github.com/Itseez/opencv/tree/master/data  \r\nfrom dlib:  \r\nhttps://github.com/davisking/dlib/blob/master/dlib/image_processing/frontal_face_detector.h  \r\nOpenFace:  \r\nhttps://github.com/cmusatyalab/openface",
      "votes": null
    },
    {
      "id": "124391",
      "postDate": "06/18/2016 00:23:22",
      "content": "<p>Morphological analyzer for Russian language <a href=\"https://github.com/kmike/pymorphy2\">pymorphy2</a> which uses <a href=\"http://opencorpora.org/\">OpenCorpora</a> dictionary. </p>",
      "rawMarkdown": "Morphological analyzer for Russian language [pymorphy2](https://github.com/kmike/pymorphy2) which uses [OpenCorpora](http://opencorpora.org/) dictionary.",
      "votes": null
    },
    {
      "id": "124803",
      "postDate": "06/22/2016 13:31:11",
      "content": "<ul>\n<li><a href=\"https://www.openstreetmap.org/\">https://www.openstreetmap.org</a> (via nominatim)</li>\n<li><a href=\"https://www.languagetool.org/\">https://www.languagetool.org/</a> spellchecker</li>\n</ul>",
      "rawMarkdown": "https://www.openstreetmap.org (via nominatim)\r\n- https://www.languagetool.org/ spellchecker",
      "votes": null
    },
    {
      "id": "125583",
      "postDate": "06/30/2016 15:56:31",
      "content": "<ul>\n<li><a href=\"http://download.tensorflow.org/models/image/imagenet/inception-2015-12-05.tgz\">http://download.tensorflow.org/models/image/imagenet/inception-2015-12-05.tgz</a></li>\n</ul>",
      "rawMarkdown": "http://download.tensorflow.org/models/image/imagenet/inception-2015-12-05.tgz",
      "votes": null
    },
    {
      "id": "125627",
      "postDate": "06/30/2016 23:48:01",
      "content": "<p><a href=\"http://www.ranks.nl/stopwords/russian\">http://www.ranks.nl/stopwords/russian</a></p>",
      "rawMarkdown": "http://www.ranks.nl/stopwords/russian",
      "votes": null
    },
    {
      "id": "125671",
      "postDate": "07/01/2016 14:16:38",
      "content": "<p>Pretrained models from:\n<a href=\"https://gist.github.com/baraldilorenzo\">https://gist.github.com/baraldilorenzo</a></p>",
      "rawMarkdown": "Pretrained models from:\r\nhttps://gist.github.com/baraldilorenzo",
      "votes": null
    },
    {
      "id": "125818",
      "postDate": "07/03/2016 00:03:18",
      "content": "<p><a href=\"http://www.tageo.com/index-e-rs-cities-RU.htm\">Russia City &amp; Town Population</a></p>",
      "rawMarkdown": "[Russia City & Town Population][1]\r\n\r\n\r\n  [1]: http://www.tageo.com/index-e-rs-cities-RU.htm",
      "votes": null
    },
    {
      "id": "125868",
      "postDate": "07/03/2016 20:12:38",
      "content": "<p>I am going to use cifar10 dataset. <a href=\"https://www.cs.toronto.edu/~kriz/cifar.html\">https://www.cs.toronto.edu/~kriz/cifar.html</a></p>",
      "rawMarkdown": "I am going to use cifar10 dataset. https://www.cs.toronto.edu/~kriz/cifar.html",
      "votes": null
    },
    {
      "id": "125879",
      "postDate": "07/04/2016 00:23:24",
      "content": "<p>We are also looking at OpenCorpora dataset</p>",
      "rawMarkdown": "We are also looking at OpenCorpora dataset",
      "votes": null
    },
    {
      "id": "125901",
      "postDate": "07/04/2016 08:54:16",
      "content": "<p>Data set from previous competition \n<a href=\"https://www.kaggle.com/c/avito-prohibited-content\">https://www.kaggle.com/c/avito-prohibited-content</a></p>",
      "rawMarkdown": "Data set from previous competition \r\nhttps://www.kaggle.com/c/avito-prohibited-content",
      "votes": null
    },
    {
      "id": "125902",
      "postDate": "07/04/2016 08:55:45",
      "content": "<p>[quote=sh1ng;125901]</p>\n\n<p>Data set from previous competition \n<a href=\"https://www.kaggle.com/c/avito-prohibited-content\">https://www.kaggle.com/c/avito-prohibited-content</a></p>\n\n<p>[/quote]</p>\n\n<p>Isn't it against the rules of that competition to use the data outside of the competition? </p>",
      "rawMarkdown": "[quote=sh1ng;125901]\r\n\r\nData set from previous competition \r\nhttps://www.kaggle.com/c/avito-prohibited-content\r\n\r\n[/quote]\r\n\r\nIsn't it against the rules of that competition to use the data outside of the competition?",
      "votes": null
    },
    {
      "id": "125903",
      "postDate": "07/04/2016 09:04:57",
      "content": "<p>[quote=ololo;125902]</p>\n\n<p>[quote=sh1ng;125901]</p>\n\n<p>Data set from previous competition \n<a href=\"https://www.kaggle.com/c/avito-prohibited-content\">https://www.kaggle.com/c/avito-prohibited-content</a></p>\n\n<p>[/quote]</p>\n\n<p>Isn't it against the rules of that competition to use the data outside of the competition? </p>\n\n<p>[/quote]\nI think it's ok:\n[quote]\nExternal data (for example, language models and computer vision pre-trained models) that is publicly available and free to use is allowed. Any source of external data not already posted must be posted in the official competition forum before the First Submission Deadline.\n[/quote]</p>",
      "rawMarkdown": "[quote=ololo;125902]\r\n\r\n[quote=sh1ng;125901]\r\n\r\nData set from previous competition \r\nhttps://www.kaggle.com/c/avito-prohibited-content\r\n\r\n[/quote]\r\n\r\nIsn't it against the rules of that competition to use the data outside of the competition? \r\n\r\n[/quote]\r\nI think it's ok:\r\n[quote]\r\nExternal data (for example, language models and computer vision pre-trained models) that is publicly available and free to use is allowed. Any source of external data not already posted must be posted in the official competition forum before the First Submission Deadline.\r\n[/quote]",
      "votes": null
    },
    {
      "id": "125904",
      "postDate": "07/04/2016 09:07:12",
      "content": "<p>[quote=Alexander Ponomarchuk;125903]\nI think it's ok\n[/quote]</p>\n\n<p>From <a href=\"https://www.kaggle.com/c/avito-prohibited-content/rules\">https://www.kaggle.com/c/avito-prohibited-content/rules</a>:</p>\n\n<blockquote>\n  <p>Unless otherwise permitted by the terms of the Competition Website, Participants <strong>must use the Data solely for the purpose and duration of the Competition</strong>, including but not limited to reading and learning from the Data, analyzing the Data, modifying the Data and generally preparing your Submission and any underlying models and participating in forum discussions on the Website. </p>\n</blockquote>",
      "rawMarkdown": "[quote=Alexander Ponomarchuk;125903]\r\nI think it's ok\r\n[/quote]\r\n\r\nFrom https://www.kaggle.com/c/avito-prohibited-content/rules:\r\n\r\n> Unless otherwise permitted by the terms of the Competition Website, Participants **must use the Data solely for the purpose and duration of the Competition**, including but not limited to reading and learning from the Data, analyzing the Data, modifying the Data and generally preparing your Submission and any underlying models and participating in forum discussions on the Website.",
      "votes": null
    },
    {
      "id": "126594",
      "postDate": "07/10/2016 12:17:10",
      "content": "<p>We are using Inception v3 pre-trained model from <a href=\"https://github.com/Lasagne/Recipes/blob/master/modelzoo/inception_v3.py\">https://github.com/Lasagne/Recipes/blob/master/modelzoo/inception_v3.py</a></p>\n\n<p>UPD: sorry, didn't know about deadline for external data</p>",
      "rawMarkdown": "We are using Inception v3 pre-trained model from https://github.com/Lasagne/Recipes/blob/master/modelzoo/inception_v3.py\r\n\r\nUPD: sorry, didn't know about deadline for external data",
      "votes": null
    },
    {
      "id": "126596",
      "postDate": "07/10/2016 12:58:03",
      "content": "<blockquote>\n  <p>Any source of external data not already posted must be posted in the official competition forum <strong>before</strong> the First Submission Deadline.</p>\n</blockquote>",
      "rawMarkdown": "> Any source of external data not already posted must be posted in the official competition forum **before** the First Submission Deadline.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118959,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "05/06/2016 11:04:49",
      "content": "<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118961,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "05/06/2016 11:29:44",
      "content": "<p>[quote=DataGeek;118959]</p>\n\n<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>\n\n<p>[/quote]</p>\n\n<p>I think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118963,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "05/06/2016 11:47:10",
      "content": "<p>[quote=Abhishek;118961]</p>\n\n<p>[quote=DataGeek;118959]</p>\n\n<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>\n\n<p>[/quote]</p>\n\n<p>I think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com</p>\n\n<p>[/quote]</p>\n\n<p>that's why I gave Kaggle link and didn't say &quot;I'll use the pretrained data accessible to everyone on www.google.com&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119038,
      "author_name": "innerproduct",
      "author_url": "",
      "post_date": "05/06/2016 21:01:46",
      "content": "<p>[quote=DataGeek;118959]</p>\n\n<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>\n\n<p>[/quote]</p>\n\n<p>In my opinion this is not specific enough.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119064,
      "author_name": "cauldnz",
      "author_url": "",
      "post_date": "05/07/2016 00:07:17",
      "content": "<p>Thinking that I will probably have a play with the WordNet Lexical DB\n<a href=\"http://wordnet.princeton.edu/\">http://wordnet.princeton.edu/</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120953,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "05/22/2016 07:19:55",
      "content": "<p>We are using the russian stopwords from the NLTK corpus:</p>\n\n<p>In python:</p>\n\n<pre><code>from nltk.corpus import stopwords\nstop_words = stopwords.words('russian')\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121852,
      "author_name": "rpmcruz",
      "author_url": "",
      "post_date": "05/30/2016 11:03:51",
      "content": "<p>Chris, there is an open <a href=\"http://globalwordnet.org/wordnets-in-the-world/\">Russian wordnet database</a>. Have you been able to use it?</p>\n\n<p>Neither <code>wn</code>, nor <code>wnb</code> (the UI), seem to find whatever words I throw at it. Not sure if I installed the database properly...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122783,
      "author_name": "",
      "author_url": "",
      "post_date": "06/07/2016 05:24:38",
      "content": "<p>[quote=Abhishek;118961]</p>\n\n<p>[quote=DataGeek;118959]</p>\n\n<p>I will be using pre-trained models from this thread <br>\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread</a></p>\n\n<p>[/quote]</p>\n\n<p>I think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122784,
      "author_name": "",
      "author_url": "",
      "post_date": "06/07/2016 05:26:04",
      "content": "<p>hi i am a student</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122891,
      "author_name": "stickshift",
      "author_url": "",
      "post_date": "06/08/2016 03:08:12",
      "content": "<p>I will be using Lucene's default list of Russian stop words. (<a href=\"http://lucene.apache.org/\">http://lucene.apache.org</a>)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123564,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "06/12/2016 18:35:34",
      "content": "<p>We are using pretrained models from this page:</p>\n\n<p><a href=\"http://ling.go.mail.ru/misc/dialogue_2015.html\">http://ling.go.mail.ru/misc/dialogue_2015.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123663,
      "author_name": "u1234x1234",
      "author_url": "",
      "post_date": "06/13/2016 11:08:26",
      "content": "<p>Pretrained models from: <br>\n<a href=\"https://github.com/dmlc/mxnet-model-gallery/\">https://github.com/dmlc/mxnet-model-gallery/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124176,
      "author_name": "u1234x1234",
      "author_url": "",
      "post_date": "06/16/2016 01:26:52",
      "content": "<p>Pretrained models from: <br>\n<a href=\"https://github.com/facebook/fb.resnet.torch\">https://github.com/facebook/fb.resnet.torch</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124200,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "06/16/2016 07:14:57",
      "content": "<p>We are using <code>russian.pickle</code> from the NLTK train_punkt tokenizers:\n<a href=\"https://github.com/mhq/train_punkt\">https://github.com/mhq/train_punkt</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124349,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "06/17/2016 12:08:46",
      "content": "<p>[quote=anokas;124200]</p>\n\n<p>We are using <code>russian.pickle</code> from the NLTK train_punkt tokenizers:\n<a href=\"https://github.com/mhq/train_punkt\">https://github.com/mhq/train_punkt</a></p>\n\n<p>[/quote]</p>\n\n<p>I thought that gave me an error when I had tried it. I tried it again. cPickle.UnpicklingError: invalid load key, '. Did you get any such error ?</p>\n\n<p>Regds\nDeb</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124350,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "06/17/2016 12:09:44",
      "content": "<p>[quote=Gerard Toonstra;123564]</p>\n\n<p>We are using pretrained models from this page:\n<a href=\"http://ling.go.mail.ru/misc/dialogue_2015.html\">http://ling.go.mail.ru/misc/dialogue_2015.html</a></p>\n\n<p>[/quote]\nUs too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124387,
      "author_name": "u1234x1234",
      "author_url": "",
      "post_date": "06/17/2016 23:54:17",
      "content": "<p>Trained models from OpenCV: <br>\n<a href=\"https://github.com/Itseez/opencv/tree/master/data\">https://github.com/Itseez/opencv/tree/master/data</a> <br>\nfrom dlib: <br>\n<a href=\"https://github.com/davisking/dlib/blob/master/dlib/image_processing/frontal_face_detector.h\">https://github.com/davisking/dlib/blob/master/dlib/image_processing/frontal_face_detector.h</a> <br>\nOpenFace: <br>\n<a href=\"https://github.com/cmusatyalab/openface\">https://github.com/cmusatyalab/openface</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124391,
      "author_name": "u1234x1234",
      "author_url": "",
      "post_date": "06/18/2016 00:23:22",
      "content": "<p>Morphological analyzer for Russian language <a href=\"https://github.com/kmike/pymorphy2\">pymorphy2</a> which uses <a href=\"http://opencorpora.org/\">OpenCorpora</a> dictionary. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124803,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "06/22/2016 13:31:11",
      "content": "<ul>\n<li><a href=\"https://www.openstreetmap.org/\">https://www.openstreetmap.org</a> (via nominatim)</li>\n<li><a href=\"https://www.languagetool.org/\">https://www.languagetool.org/</a> spellchecker</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125583,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "06/30/2016 15:56:31",
      "content": "<ul>\n<li><a href=\"http://download.tensorflow.org/models/image/imagenet/inception-2015-12-05.tgz\">http://download.tensorflow.org/models/image/imagenet/inception-2015-12-05.tgz</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125627,
      "author_name": "stickshift",
      "author_url": "",
      "post_date": "06/30/2016 23:48:01",
      "content": "<p><a href=\"http://www.ranks.nl/stopwords/russian\">http://www.ranks.nl/stopwords/russian</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125671,
      "author_name": "dsalex",
      "author_url": "",
      "post_date": "07/01/2016 14:16:38",
      "content": "<p>Pretrained models from:\n<a href=\"https://gist.github.com/baraldilorenzo\">https://gist.github.com/baraldilorenzo</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125818,
      "author_name": "jpison",
      "author_url": "",
      "post_date": "07/03/2016 00:03:18",
      "content": "<p><a href=\"http://www.tageo.com/index-e-rs-cities-RU.htm\">Russia City &amp; Town Population</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125868,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "07/03/2016 20:12:38",
      "content": "<p>I am going to use cifar10 dataset. <a href=\"https://www.cs.toronto.edu/~kriz/cifar.html\">https://www.cs.toronto.edu/~kriz/cifar.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125879,
      "author_name": "stickshift",
      "author_url": "",
      "post_date": "07/04/2016 00:23:24",
      "content": "<p>We are also looking at OpenCorpora dataset</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125901,
      "author_name": "sh1ngg",
      "author_url": "",
      "post_date": "07/04/2016 08:54:16",
      "content": "<p>Data set from previous competition \n<a href=\"https://www.kaggle.com/c/avito-prohibited-content\">https://www.kaggle.com/c/avito-prohibited-content</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125902,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/04/2016 08:55:45",
      "content": "<p>[quote=sh1ng;125901]</p>\n\n<p>Data set from previous competition \n<a href=\"https://www.kaggle.com/c/avito-prohibited-content\">https://www.kaggle.com/c/avito-prohibited-content</a></p>\n\n<p>[/quote]</p>\n\n<p>Isn't it against the rules of that competition to use the data outside of the competition? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125903,
      "author_name": "dsalex",
      "author_url": "",
      "post_date": "07/04/2016 09:04:57",
      "content": "<p>[quote=ololo;125902]</p>\n\n<p>[quote=sh1ng;125901]</p>\n\n<p>Data set from previous competition \n<a href=\"https://www.kaggle.com/c/avito-prohibited-content\">https://www.kaggle.com/c/avito-prohibited-content</a></p>\n\n<p>[/quote]</p>\n\n<p>Isn't it against the rules of that competition to use the data outside of the competition? </p>\n\n<p>[/quote]\nI think it's ok:\n[quote]\nExternal data (for example, language models and computer vision pre-trained models) that is publicly available and free to use is allowed. Any source of external data not already posted must be posted in the official competition forum before the First Submission Deadline.\n[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125904,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/04/2016 09:07:12",
      "content": "<p>[quote=Alexander Ponomarchuk;125903]\nI think it's ok\n[/quote]</p>\n\n<p>From <a href=\"https://www.kaggle.com/c/avito-prohibited-content/rules\">https://www.kaggle.com/c/avito-prohibited-content/rules</a>:</p>\n\n<blockquote>\n  <p>Unless otherwise permitted by the terms of the Competition Website, Participants <strong>must use the Data solely for the purpose and duration of the Competition</strong>, including but not limited to reading and learning from the Data, analyzing the Data, modifying the Data and generally preparing your Submission and any underlying models and participating in forum discussions on the Website. </p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126594,
      "author_name": "evgenyeltyshev",
      "author_url": "",
      "post_date": "07/10/2016 12:17:10",
      "content": "<p>We are using Inception v3 pre-trained model from <a href=\"https://github.com/Lasagne/Recipes/blob/master/modelzoo/inception_v3.py\">https://github.com/Lasagne/Recipes/blob/master/modelzoo/inception_v3.py</a></p>\n\n<p>UPD: sorry, didn't know about deadline for external data</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126596,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "07/10/2016 12:58:03",
      "content": "<blockquote>\n  <p>Any source of external data not already posted must be posted in the official competition forum <strong>before</strong> the First Submission Deadline.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118932": "Since, external data is allowed, please share it here before you use them.. :)",
    "118959": "I will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread",
    "118961": "[quote=DataGeek;118959]\r\n\r\nI will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\r\n\r\n[/quote]\r\n\r\nI think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com",
    "118963": "[quote=Abhishek;118961]\r\n\r\n[quote=DataGeek;118959]\r\n\r\nI will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\r\n\r\n[/quote]\r\n\r\nI think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com\r\n\r\n[/quote]\r\n\r\nthat's why I gave Kaggle link and didn't say \"I'll use the pretrained data accessible to everyone on www.google.com\"",
    "119038": "[quote=DataGeek;118959]\r\n\r\nI will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\r\n\r\n[/quote]\r\n\r\nIn my opinion this is not specific enough.",
    "119064": "Thinking that I will probably have a play with the WordNet Lexical DB\r\nhttp://wordnet.princeton.edu/",
    "120953": "We are using the russian stopwords from the NLTK corpus:\r\n\r\nIn python:\r\n\r\n    from nltk.corpus import stopwords\r\n    stop_words = stopwords.words('russian')",
    "121852": "Chris, there is an open [Russian wordnet database][1]. Have you been able to use it?\r\n\r\nNeither `wn`, nor `wnb` (the UI), seem to find whatever words I throw at it. Not sure if I installed the database properly...\r\n\r\n\r\n  [1]: http://globalwordnet.org/wordnets-in-the-world/",
    "122783": "[quote=Abhishek;118961]\r\n\r\n[quote=DataGeek;118959]\r\n\r\nI will be using pre-trained models from this thread  \r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread\r\n\r\n[/quote]\r\n\r\nI think that is too general. One could say, I'll use the pretrained data accessible to everyone on www.google.com\r\n\r\n[/quote]",
    "122784": "hi i am a student",
    "122891": "I will be using Lucene's default list of Russian stop words. (http://lucene.apache.org)",
    "123564": "We are using pretrained models from this page:\r\n\r\nhttp://ling.go.mail.ru/misc/dialogue_2015.html",
    "123663": "Pretrained models from:  \r\nhttps://github.com/dmlc/mxnet-model-gallery/",
    "124176": "Pretrained models from:  \r\nhttps://github.com/facebook/fb.resnet.torch",
    "124200": "We are using `russian.pickle` from the NLTK train_punkt tokenizers:\r\nhttps://github.com/mhq/train_punkt",
    "124349": "[quote=anokas;124200]\r\n\r\nWe are using `russian.pickle` from the NLTK train_punkt tokenizers:\r\nhttps://github.com/mhq/train_punkt\r\n\r\n[/quote]\r\n\r\nI thought that gave me an error when I had tried it. I tried it again. cPickle.UnpicklingError: invalid load key, '. Did you get any such error ?\r\n\r\nRegds\r\nDeb",
    "124350": "[quote=Gerard Toonstra;123564]\r\n\r\nWe are using pretrained models from this page:\r\nhttp://ling.go.mail.ru/misc/dialogue_2015.html\r\n\r\n[/quote]\r\nUs too",
    "124387": "Trained models from OpenCV:  \r\nhttps://github.com/Itseez/opencv/tree/master/data  \r\nfrom dlib:  \r\nhttps://github.com/davisking/dlib/blob/master/dlib/image_processing/frontal_face_detector.h  \r\nOpenFace:  \r\nhttps://github.com/cmusatyalab/openface",
    "124391": "Morphological analyzer for Russian language [pymorphy2](https://github.com/kmike/pymorphy2) which uses [OpenCorpora](http://opencorpora.org/) dictionary.",
    "124803": "https://www.openstreetmap.org (via nominatim)\r\n- https://www.languagetool.org/ spellchecker",
    "125583": "http://download.tensorflow.org/models/image/imagenet/inception-2015-12-05.tgz",
    "125627": "http://www.ranks.nl/stopwords/russian",
    "125671": "Pretrained models from:\r\nhttps://gist.github.com/baraldilorenzo",
    "125818": "[Russia City & Town Population][1]\r\n\r\n\r\n  [1]: http://www.tageo.com/index-e-rs-cities-RU.htm",
    "125868": "I am going to use cifar10 dataset. https://www.cs.toronto.edu/~kriz/cifar.html",
    "125879": "We are also looking at OpenCorpora dataset",
    "125901": "Data set from previous competition \r\nhttps://www.kaggle.com/c/avito-prohibited-content",
    "125902": "[quote=sh1ng;125901]\r\n\r\nData set from previous competition \r\nhttps://www.kaggle.com/c/avito-prohibited-content\r\n\r\n[/quote]\r\n\r\nIsn't it against the rules of that competition to use the data outside of the competition?",
    "125903": "[quote=ololo;125902]\r\n\r\n[quote=sh1ng;125901]\r\n\r\nData set from previous competition \r\nhttps://www.kaggle.com/c/avito-prohibited-content\r\n\r\n[/quote]\r\n\r\nIsn't it against the rules of that competition to use the data outside of the competition? \r\n\r\n[/quote]\r\nI think it's ok:\r\n[quote]\r\nExternal data (for example, language models and computer vision pre-trained models) that is publicly available and free to use is allowed. Any source of external data not already posted must be posted in the official competition forum before the First Submission Deadline.\r\n[/quote]",
    "125904": "[quote=Alexander Ponomarchuk;125903]\r\nI think it's ok\r\n[/quote]\r\n\r\nFrom https://www.kaggle.com/c/avito-prohibited-content/rules:\r\n\r\n> Unless otherwise permitted by the terms of the Competition Website, Participants **must use the Data solely for the purpose and duration of the Competition**, including but not limited to reading and learning from the Data, analyzing the Data, modifying the Data and generally preparing your Submission and any underlying models and participating in forum discussions on the Website.",
    "126594": "We are using Inception v3 pre-trained model from https://github.com/Lasagne/Recipes/blob/master/modelzoo/inception_v3.py\r\n\r\nUPD: sorry, didn't know about deadline for external data",
    "126596": "> Any source of external data not already posted must be posted in the official competition forum **before** the First Submission Deadline."
  },
  "source": "meta"
}