{
  "id": 55897,
  "title": "Official External Data Thread",
  "url": "/competitions/avito-demand-prediction/discussion/55897",
  "author_name": "Sohier Dane",
  "post_date": "2018-05-02T22:29:14.300000",
  "votes": 11,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Post your (freely available) external data links here.</p>",
  "messages": [
    {
      "id": 326798,
      "postDate": "2018-05-10T11:03:19.453Z",
      "content": "<hr>\n\n<p>Text Models</p>\n\n<p><a href=\"https://s3-us-west-1.amazonaws.com/fasttext-vectors/wiki.ru.vec\">Fasttext trained on wikipedia (300d)</a></p>\n\n<p><a href=\"https://s3-us-west-1.amazonaws.com/fasttext-vectors/word-vectors-v2/cc.ru.300.vec.gz\">Fasttext trained on CommonCrawl (300d)</a></p>\n\n<p><a href=\"https://github.com/named-entity/nltk4russian\">nltk4russian</a></p>\n\n<p><a href=\"http://www.redhenlab.org/home/the-cognitive-core-research-topics-in-red-hen/the-barnyard/russian-nlp\">Redhenlab</a></p>\n\n<p><a href=\"https://github.com/nlpub/russe-evaluation/tree/master/russe/measures/word2vec\">Russian w2v</a></p>\n\n<hr>\n\n<p>Image Models</p>\n\n<p>All from <a href=\"https://keras.io/applications/\">keras.applications</a></p>\n\n<p>neural-image-assessment <a href=\"https://github.com/titu1994/neural-image-assessment\">git repo</a></p>\n\n<p>ILGnet <a href=\"https://github.com/BestiVictory/ILGnet\">git repo</a></p>",
      "rawMarkdown": "----------\n\nText Models\n\n[Fasttext trained on wikipedia (300d)][1]\n\n[Fasttext trained on CommonCrawl (300d)][2]\n\n[nltk4russian][3]\n\n[Redhenlab][4]\n\n[Russian w2v][5]\n\n----------\nImage Models\n\nAll from [keras.applications][6]\n\nneural-image-assessment [git repo][7]\n\nILGnet [git repo][8]\n\n\n  [1]: https://s3-us-west-1.amazonaws.com/fasttext-vectors/wiki.ru.vec\n  [2]: https://s3-us-west-1.amazonaws.com/fasttext-vectors/word-vectors-v2/cc.ru.300.vec.gz\n  [3]: https://github.com/named-entity/nltk4russian\n  [4]: http://www.redhenlab.org/home/the-cognitive-core-research-topics-in-red-hen/the-barnyard/russian-nlp\n  [5]: https://github.com/nlpub/russe-evaluation/tree/master/russe/measures/word2vec\n  [6]: https://keras.io/applications/\n  [7]: https://github.com/titu1994/neural-image-assessment\n  [8]: https://github.com/BestiVictory/ILGnet",
      "votes": 11
    },
    {
      "id": 322418,
      "postDate": "2018-05-02T22:29:14.300Z",
      "content": "<p>Post your (freely available) external data links here.</p>",
      "rawMarkdown": "Post your (freely available) external data links here.",
      "votes": 11
    },
    {
      "id": 345145,
      "postDate": "2018-06-19T09:14:06.110Z",
      "content": "<p>tfidf with 1-4 ngram ('3131473e84a9' and '75ebe6b373ec' were dropped):\n<a href=\"https://www.dropbox.com/s/hrbyfvmlsfmikwl/test_tfidf_sparse_1_4_clean_data_v1.npz\">https://www.dropbox.com/s/hrbyfvmlsfmikwl/test_tfidf_sparse_1_4_clean_data_v1.npz</a>\n<a href=\"https://www.dropbox.com/s/3xwnixrjexiz7qd/train_tfidf_sparse_1_4_clean_data_v1.npz\">https://www.dropbox.com/s/3xwnixrjexiz7qd/train_tfidf_sparse_1_4_clean_data_v1.npz</a></p>\n\n<p>PIL&amp;CV2 features:\n<a href=\"https://www.dropbox.com/s/0mfgb90m4iggn6i/train_img_features_v1.csv.gz\">https://www.dropbox.com/s/0mfgb90m4iggn6i/train_img_features_v1.csv.gz</a>\n<a href=\"https://www.dropbox.com/s/vr1fv94k2fukd89/test_img_features_v1.csv.gz\">https://www.dropbox.com/s/vr1fv94k2fukd89/test_img_features_v1.csv.gz</a></p>",
      "rawMarkdown": "tfidf with 1-4 ngram ('3131473e84a9' and '75ebe6b373ec' were dropped):\nhttps://www.dropbox.com/s/hrbyfvmlsfmikwl/test_tfidf_sparse_1_4_clean_data_v1.npz\nhttps://www.dropbox.com/s/3xwnixrjexiz7qd/train_tfidf_sparse_1_4_clean_data_v1.npz\n\nPIL&amp;CV2 features:\nhttps://www.dropbox.com/s/0mfgb90m4iggn6i/train_img_features_v1.csv.gz\nhttps://www.dropbox.com/s/vr1fv94k2fukd89/test_img_features_v1.csv.gz",
      "votes": 7,
      "replies": [
        {
          "id": 345470,
          "postDate": "2018-06-20T00:01:34.453Z",
          "content": "<p>Thank you!</p>\n\n<blockquote>\n  <p>tfidf with 1-4 ngram ('3131473e84a9' and '75ebe6b373ec' were dropped)</p>\n</blockquote>\n\n<p>Seems like hundred more item_id's were dropped.</p>",
          "rawMarkdown": "Thank you!\n&gt; tfidf with 1-4 ngram ('3131473e84a9' and '75ebe6b373ec' were dropped)\n\nSeems like hundred more item_id's were dropped.",
          "votes": 1
        },
        {
          "id": 345757,
          "postDate": "2018-06-20T11:59:44.237Z",
          "content": "<p>train_data[pd.to_datetime(train_data.activation_date) &lt;= pd.to_datetime('2017-03-28')]</p>",
          "rawMarkdown": "train_data[pd.to_datetime(train_data.activation_date) &lt;= pd.to_datetime('2017-03-28')]",
          "votes": 2
        }
      ]
    },
    {
      "id": 331017,
      "postDate": "2018-05-20T07:18:50.883Z",
      "content": "<p>List of Russian cities on Wikipedia, with census information:\n<a href=\"https://en.wikipedia.org/wiki/List_of_cities_and_towns_in_Russia_by_population\">https://en.wikipedia.org/wiki/List_of_cities_and_towns_in_Russia_by_population</a></p>\n\n<p>See also this kernel: <a href=\"https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia\">https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia</a></p>",
      "rawMarkdown": "List of Russian cities on Wikipedia, with census information:\nhttps://en.wikipedia.org/wiki/List_of_cities_and_towns_in_Russia_by_population\n\nSee also this kernel: https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia",
      "votes": 7,
      "replies": [
        {
          "id": 331371,
          "postDate": "2018-05-21T05:27:58.203Z",
          "content": "<p>Nice work !! May be a naive question how does this data get approved to be used in the competition oe when do we get to know if this data can be used. Please suggest !!</p>",
          "rawMarkdown": "Nice work !! May be a naive question how does this data get approved to be used in the competition oe when do we get to know if this data can be used. Please suggest !!"
        },
        {
          "id": 331404,
          "postDate": "2018-05-21T07:04:29.760Z",
          "content": "<p>creating a public kernel like <a href=\"https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\">region-and-city-details</a> with the collected data attached should be sufficient. However, @Sohier Dane please correct me if a am wrong</p>",
          "rawMarkdown": "creating a public kernel like [region-and-city-details][1] with the collected data attached should be sufficient. However, @Sohier Dane please correct me if a am wrong\n\n\n  [1]: https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters",
          "votes": 1
        },
        {
          "id": 331419,
          "postDate": "2018-05-21T07:36:37.637Z",
          "content": "<p>I posted here precisely to have feedback from the organizers. If they don't reply, I will take it as approval</p>",
          "rawMarkdown": "I posted here precisely to have feedback from the organizers. If they don't reply, I will take it as approval"
        },
        {
          "id": 331628,
          "postDate": "2018-05-21T15:51:42.647Z",
          "content": "<p>@Dieter is correct. External data is valid as long as it is available to all competitors and is listed in this thread. </p>\n\n<p>Making data available to other competitors via a kernel like <a href=\"https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\">region-and-city-details</a> is great.</p>",
          "rawMarkdown": "@Dieter is correct. External data is valid as long as it is available to all competitors and is listed in this thread. \n\nMaking data available to other competitors via a kernel like [region-and-city-details](https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters) is great.",
          "votes": 3
        },
        {
          "id": 331761,
          "postDate": "2018-05-21T20:22:58.917Z",
          "content": "<p>Thanks @Sohier Dane for the clarification. I added the file under \"Data\" in my script:\n<a href=\"https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia/data\">https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia/data</a></p>",
          "rawMarkdown": "Thanks @Sohier Dane for the clarification. I added the file under \"Data\" in my script:\nhttps://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia/data",
          "votes": 1
        }
      ]
    },
    {
      "id": 345940,
      "postDate": "2018-06-20T19:05:30.647Z",
      "content": "<p>Distributed training framework for Tensorflow and Pytorch \n<a href=\"https://github.com/uber/horovod\">https://github.com/uber/horovod</a></p>",
      "rawMarkdown": "Distributed training framework for Tensorflow and Pytorch \nhttps://github.com/uber/horovod",
      "votes": 1
    },
    {
      "id": 345790,
      "postDate": "2018-06-20T13:07:51.207Z",
      "content": "<p>Wages in regions:\n<a href=\"http://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/wages/\">http://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/wages/</a>\n<a href=\"http://www.gks.ru/free_doc/new_site/population/trud/sr-zarplata/t2.xlsx\">http://www.gks.ru/free_doc/new_site/population/trud/sr-zarplata/t2.xlsx</a></p>",
      "rawMarkdown": "Wages in regions:\nhttp://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/wages/\nhttp://www.gks.ru/free_doc/new_site/population/trud/sr-zarplata/t2.xlsx",
      "votes": 1
    },
    {
      "id": 345174,
      "postDate": "2018-06-19T10:47:29.953Z",
      "content": "<p>Fast.ai models and weights (<a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>, <a href=\"http://files.fast.ai/models/weights.tgz\">http://files.fast.ai/models/weights.tgz</a>).\nPytorch models and weights.</p>",
      "rawMarkdown": "Fast.ai models and weights (https://github.com/fastai/fastai, http://files.fast.ai/models/weights.tgz).\nPytorch models and weights.",
      "votes": 1
    },
    {
      "id": 345149,
      "postDate": "2018-06-19T09:22:18.867Z",
      "content": "<p>Hello!\nWe want to use NIMA pretrained models (<a href=\"https://github.com/titu1994/neural-image-assessment\">https://github.com/titu1994/neural-image-assessment</a>).</p>",
      "rawMarkdown": "Hello!\nWe want to use NIMA pretrained models (https://github.com/titu1994/neural-image-assessment).",
      "votes": 1
    },
    {
      "id": 344679,
      "postDate": "2018-06-18T14:05:56.137Z",
      "content": "<p><a href=\"https://pytorch.org/docs/stable/torchvision/models.html\">https://pytorch.org/docs/stable/torchvision/models.html</a></p>",
      "rawMarkdown": "https://pytorch.org/docs/stable/torchvision/models.html",
      "votes": 1
    },
    {
      "id": 322917,
      "postDate": "2018-05-03T23:11:14.487Z",
      "content": "<p>I don't know about other competitors, but I think hourly site traffic data from Avito may be a helpful feature for our model.  Alexa Rank improved my model, and something more detailed may be even more helpful.</p>",
      "rawMarkdown": "I don't know about other competitors, but I think hourly site traffic data from Avito may be a helpful feature for our model.  Alexa Rank improved my model, and something more detailed may be even more helpful.",
      "votes": 1,
      "replies": [
        {
          "id": 324737,
          "postDate": "2018-05-07T23:05:52.730Z",
          "content": "<p>How did you get Alexa data?</p>",
          "rawMarkdown": "How did you get Alexa data?"
        },
        {
          "id": 326892,
          "postDate": "2018-05-10T13:21:45.653Z",
          "content": "<p>Worst comes to worst, just gotta manually recreate it. Here's the Alexa ranks for avito.ru</p>\n\n<p><img src=\"https://traffic.alexa.com/graph?u=avito.ru\" alt=\"avito alexa rank\"></p>\n\n<p>When I get desperate enough / exhaust my data modeling, I'll go ahead and convert it to a csv and share in a kernel if no one else has done it by that point.</p>",
          "rawMarkdown": "Worst comes to worst, just gotta manually recreate it. Here's the Alexa ranks for avito.ru\n\n![avito alexa rank][1]\n\nWhen I get desperate enough / exhaust my data modeling, I'll go ahead and convert it to a csv and share in a kernel if no one else has done it by that point.\n\n  [1]: https://traffic.alexa.com/graph?u=avito.ru",
          "votes": 4
        },
        {
          "id": 328906,
          "postDate": "2018-05-15T10:49:51.287Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 331631,
          "postDate": "2018-05-21T15:54:30.427Z",
          "content": "<p><a href=\"/matthewa313\">@matthewa313</a> Is the Alexa data freely available? If it’s not freely available to everyone in the competition, and for the sponsors to use in their models after the competition, it can’t be used in this competition.</p>",
          "rawMarkdown": "@matthewa313 Is the Alexa data freely available? If it’s not freely available to everyone in the competition, and for the sponsors to use in their models after the competition, it can’t be used in this competition."
        },
        {
          "id": 331677,
          "postDate": "2018-05-21T17:13:39.670Z",
          "content": "<p><a href=\"/matthewa313\">@matthewa313</a> Based on the information provided by <a href=\"/authman\">@authman</a> in <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57138\">this thread</a> Alexa data is not sufficiently open for use in this competition. The two key facts here are that only a limited amount of the data is available for free and that all of the data requires a login.</p>",
          "rawMarkdown": "@matthewa313 Based on the information provided by @authman in [this thread](https://www.kaggle.com/c/avito-demand-prediction/discussion/57138) Alexa data is not sufficiently open for use in this competition. The two key facts here are that only a limited amount of the data is available for free and that all of the data requires a login.",
          "votes": 1
        }
      ]
    },
    {
      "id": 345094,
      "postDate": "2018-06-19T07:37:56.427Z",
      "content": "<p>Currency exchange rates: <a href=\"http://cbr.ru/eng/\">http://cbr.ru/eng/</a>\nExample: <a href=\"http://cbr.ru/eng/currency_base/daily/?date_req=15.03.2017\">http://cbr.ru/eng/currency_base/daily/?date_req=15.03.2017</a></p>\n\n<p>Weather data: <a href=\"https://www.gismeteo.ru/\">https://www.gismeteo.ru/</a>\nExample: <a href=\"https://www.gismeteo.ru/diary/4720/2017/3/\">https://www.gismeteo.ru/diary/4720/2017/3/</a></p>\n\n<p>ResNet152 model for Keras: \n<a href=\"https://github.com/broadinstitute/keras-resnet/blob/master/keras_resnet/models/_2d.py\">https://github.com/broadinstitute/keras-resnet/blob/master/keras_resnet/models/_2d.py</a>\nWeights: <a href=\"https://github.com/fizyr/keras-models/releases/tag/v0.0.1\">https://github.com/fizyr/keras-models/releases/tag/v0.0.1</a></p>",
      "rawMarkdown": "Currency exchange rates: http://cbr.ru/eng/\nExample: http://cbr.ru/eng/currency_base/daily/?date_req=15.03.2017\n\nWeather data: https://www.gismeteo.ru/\nExample: https://www.gismeteo.ru/diary/4720/2017/3/\n\nResNet152 model for Keras: \nhttps://github.com/broadinstitute/keras-resnet/blob/master/keras_resnet/models/_2d.py\nWeights: https://github.com/fizyr/keras-models/releases/tag/v0.0.1\n",
      "votes": 2
    },
    {
      "id": 345066,
      "postDate": "2018-06-19T06:25:09.157Z",
      "content": "<p>Pretrained russian word vectors from <a href=\"https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md\">this page</a></p>",
      "rawMarkdown": "Pretrained russian word vectors from [this page][1]\n\n\n  [1]: https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md",
      "votes": 2
    },
    {
      "id": 343181,
      "postDate": "2018-06-14T20:54:49.737Z",
      "content": "<p>Scraping avito.ru ads</p>",
      "rawMarkdown": "Scraping avito.ru ads",
      "votes": -3,
      "replies": [
        {
          "id": 343358,
          "postDate": "2018-06-15T07:56:43.887Z",
          "content": "<p>not allowed, as long as you do not share the scraped data to everybody. See one post above</p>",
          "rawMarkdown": "not allowed, as long as you do not share the scraped data to everybody. See one post above\n",
          "votes": 1
        },
        {
          "id": 343361,
          "postDate": "2018-06-15T08:05:39.667Z",
          "content": "<p>The rules only require you to share the source of the data in the discussion, not the data itself:</p>\n\n<pre><code>The following provision supersedes General Rules Section 7.C. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required.\" If you are using a source of external data, **you must post the source to the official external data** forum thread no later than one week prior to the deadline.\n</code></pre>\n\n<p>Note that Im still exploring (haven't really got any data yet) but <a href=\"/sohier\">@sohier</a> said \"External data is valid as long as it is available to all competitors and is listed in this thread.\" which I read as a necessary condition whereas \"Making data available to other competitors via a kernel like region-and-city-details is great.\" I read  as nice to have but not required.</p>\n\n<p><a href=\"/sohier\">@sohier</a> please clarify.</p>",
          "rawMarkdown": "The rules only require you to share the source of the data in the discussion, not the data itself:\n\n    The following provision supersedes General Rules Section 7.C. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required.\" If you are using a source of external data, **you must post the source to the official external data** forum thread no later than one week prior to the deadline.\n\nNote that Im still exploring (haven't really got any data yet) but @sohier said \"External data is valid as long as it is available to all competitors and is listed in this thread.\" which I read as a necessary condition whereas \"Making data available to other competitors via a kernel like region-and-city-details is great.\" I read  as nice to have but not required.\n\n@sohier please clarify.",
          "votes": 1
        },
        {
          "id": 343387,
          "postDate": "2018-06-15T09:06:53.903Z",
          "content": "<p>From my understanding others need to be able to reproduce your dataset. With scraping, since the data is not really static, your exact data is not \"available to all competitors\". Also the \"right and authority to use such external data\" is a grey area for scraping.</p>",
          "rawMarkdown": "From my understanding others need to be able to reproduce your dataset. With scraping, since the data is not really static, your exact data is not \"available to all competitors\". Also the \"right and authority to use such external data\" is a grey area for scraping.",
          "votes": 1
        },
        {
          "id": 343392,
          "postDate": "2018-06-15T09:14:12.423Z",
          "content": "<p>They mentioned source of the data, not data itself.</p>\n\n<p>FWIW, in another competition top winner of IEEE Camera competitition <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49367\">https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49367</a> scraped 300+ Gb of photos from Yandex and it wasn't an issue for Kaggle.</p>",
          "rawMarkdown": "They mentioned source of the data, not data itself.\n\nFWIW, in another competition top winner of IEEE Camera competitition https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49367 scraped 300+ Gb of photos from Yandex and it wasn't an issue for Kaggle.",
          "votes": 1
        },
        {
          "id": 343565,
          "postDate": "2018-06-15T15:23:31.963Z",
          "content": "<p>To be clear: scraping data is prohibited by the terms and conditions of Avito's website . Any data acquired by scraping Avito is therefore not generally available and is against the competition rules.</p>\n\n<p>We can't put you guys into a situation where you need to violate the rules of other websites in order to stay competitive.</p>",
          "rawMarkdown": "To be clear: scraping data is prohibited by the terms and conditions of Avito's website . Any data acquired by scraping Avito is therefore not generally available and is against the competition rules.\n\nWe can't put you guys into a situation where you need to violate the rules of other websites in order to stay competitive.",
          "votes": 4
        },
        {
          "id": 343568,
          "postDate": "2018-06-15T15:30:26.437Z",
          "content": "<p>Thanks for the quick clarification.</p>",
          "rawMarkdown": "Thanks for the quick clarification."
        },
        {
          "id": 343632,
          "postDate": "2018-06-15T17:46:25Z",
          "content": "<p>Happy to help!</p>\n\n<p>To be clear, the IEEE competition had different (and unusual) rules regarding external data. </p>",
          "rawMarkdown": "Happy to help!\n\nTo be clear, the IEEE competition had different (and unusual) rules regarding external data. "
        }
      ]
    },
    {
      "id": 349022,
      "postDate": "2018-06-27T17:11:39.863Z",
      "content": "<p>On be half of my team, all the external data we use are already mentioned here.</p>",
      "rawMarkdown": "On be half of my team, all the external data we use are already mentioned here."
    },
    {
      "id": 346940,
      "postDate": "2018-06-22T20:51:05.683Z",
      "content": "<p>cryptocurrency data from <a href=\"https://www.kaggle.com/mczielinski/bitcoin-historical-data/data\">https://www.kaggle.com/mczielinski/bitcoin-historical-data/data</a>\nand\n<a href=\"https://www.kaggle.com/kingburrito666/ethereum-historical-data/data\">https://www.kaggle.com/kingburrito666/ethereum-historical-data/data</a></p>",
      "rawMarkdown": "cryptocurrency data from https://www.kaggle.com/mczielinski/bitcoin-historical-data/data\nand\nhttps://www.kaggle.com/kingburrito666/ethereum-historical-data/data"
    },
    {
      "id": 346796,
      "postDate": "2018-06-22T11:58:25.997Z",
      "content": "<p>I have a question. Is it allowed to use external data other participants mentioned on this discussion forum.\n<a href=\"https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\">https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters</a>\nfor example this kernel's data.</p>",
      "rawMarkdown": "I have a question. Is it allowed to use external data other participants mentioned on this discussion forum.\nhttps://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\nfor example this kernel's data.",
      "replies": [
        {
          "id": 346842,
          "postDate": "2018-06-22T14:43:23.567Z",
          "content": "<p>Correct, you don't need to repeat an item that's already been listed.</p>",
          "rawMarkdown": "Correct, you don't need to repeat an item that's already been listed."
        },
        {
          "id": 346984,
          "postDate": "2018-06-22T23:47:25.103Z",
          "content": "<p>Thank you very much.</p>",
          "rawMarkdown": "Thank you very much."
        }
      ]
    },
    {
      "id": 346530,
      "postDate": "2018-06-21T22:17:13.027Z",
      "content": "<p>All posted already~</p>",
      "rawMarkdown": "All posted already~"
    },
    {
      "id": 346048,
      "postDate": "2018-06-21T01:23:03.753Z",
      "content": "<p>On be half of my team, all the external data we use are already mentioned here :)</p>",
      "rawMarkdown": "On be half of my team, all the external data we use are already mentioned here :)"
    },
    {
      "id": 346008,
      "postDate": "2018-06-20T22:21:26.980Z",
      "content": "<p>On behalf of our team, I warrant that all external data used has either already been mentioned in this thread, is available in a publicly shared kernel here, or is contained in this following freely available public CSV file: <a href=\"https://s3.amazonaws.com/avito-demand-kaggle/region_macro.csv\">https://s3.amazonaws.com/avito-demand-kaggle/region_macro.csv</a></p>",
      "rawMarkdown": "On behalf of our team, I warrant that all external data used has either already been mentioned in this thread, is available in a publicly shared kernel here, or is contained in this following freely available public CSV file: https://s3.amazonaws.com/avito-demand-kaggle/region_macro.csv"
    },
    {
      "id": 345979,
      "postDate": "2018-06-20T21:05:20.837Z",
      "content": "<p><a href=\"https://keras.io/applications/\">https://keras.io/applications/</a></p>",
      "rawMarkdown": "https://keras.io/applications/"
    },
    {
      "id": 345794,
      "postDate": "2018-06-20T13:17:01.727Z",
      "content": "<p>The Demographic Yearbook of Russia (not sure I'll use it, but it may be worth of giving a try):\n<a href=\"http://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/publications/catalog/doc_1137674209312\">http://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/publications/catalog/doc_1137674209312</a></p>",
      "rawMarkdown": "The Demographic Yearbook of Russia (not sure I'll use it, but it may be worth of giving a try):\nhttp://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/publications/catalog/doc_1137674209312"
    },
    {
      "id": 345355,
      "postDate": "2018-06-19T18:25:37.347Z",
      "content": "<p>keras pretrained image models - resnet, densnet, vgg19 <br>\nimagenet synsets <br>\nfasttext russian pretrain vectors &amp; <a href=\"http://panchenko.me/data/dsl-backup/w2v-ru/\">http://panchenko.me/data/dsl-backup/w2v-ru/</a> <br>\nnltk &amp; pymorph <br>\nlat lon coords shared in kernels     </p>",
      "rawMarkdown": "keras pretrained image models - resnet, densnet, vgg19   \nimagenet synsets   \nfasttext russian pretrain vectors &amp; http://panchenko.me/data/dsl-backup/w2v-ru/    \nnltk &amp; pymorph    \nlat lon coords shared in kernels     "
    },
    {
      "id": 343869,
      "postDate": "2018-06-16T11:43:03.053Z",
      "content": "<p>Hi, we are making use of the Forbes top 2000 ranking. <a href=\"https://www.forbes.com/global2000/#6507469f335d\">https://www.forbes.com/global2000/#6507469f335d</a></p>",
      "rawMarkdown": "Hi, we are making use of the Forbes top 2000 ranking. https://www.forbes.com/global2000/#6507469f335d\n"
    }
  ],
  "comments": [
    {
      "id": 326798,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-05-10T11:03:19.453000",
      "content": "<hr>\n\n<p>Text Models</p>\n\n<p><a href=\"https://s3-us-west-1.amazonaws.com/fasttext-vectors/wiki.ru.vec\">Fasttext trained on wikipedia (300d)</a></p>\n\n<p><a href=\"https://s3-us-west-1.amazonaws.com/fasttext-vectors/word-vectors-v2/cc.ru.300.vec.gz\">Fasttext trained on CommonCrawl (300d)</a></p>\n\n<p><a href=\"https://github.com/named-entity/nltk4russian\">nltk4russian</a></p>\n\n<p><a href=\"http://www.redhenlab.org/home/the-cognitive-core-research-topics-in-red-hen/the-barnyard/russian-nlp\">Redhenlab</a></p>\n\n<p><a href=\"https://github.com/nlpub/russe-evaluation/tree/master/russe/measures/word2vec\">Russian w2v</a></p>\n\n<hr>\n\n<p>Image Models</p>\n\n<p>All from <a href=\"https://keras.io/applications/\">keras.applications</a></p>\n\n<p>neural-image-assessment <a href=\"https://github.com/titu1994/neural-image-assessment\">git repo</a></p>\n\n<p>ILGnet <a href=\"https://github.com/BestiVictory/ILGnet\">git repo</a></p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 345145,
      "author_name": "Taras Baranyuk",
      "author_url": "",
      "post_date": "2018-06-19T09:14:06.110000",
      "content": "<p>tfidf with 1-4 ngram ('3131473e84a9' and '75ebe6b373ec' were dropped):\n<a href=\"https://www.dropbox.com/s/hrbyfvmlsfmikwl/test_tfidf_sparse_1_4_clean_data_v1.npz\">https://www.dropbox.com/s/hrbyfvmlsfmikwl/test_tfidf_sparse_1_4_clean_data_v1.npz</a>\n<a href=\"https://www.dropbox.com/s/3xwnixrjexiz7qd/train_tfidf_sparse_1_4_clean_data_v1.npz\">https://www.dropbox.com/s/3xwnixrjexiz7qd/train_tfidf_sparse_1_4_clean_data_v1.npz</a></p>\n\n<p>PIL&amp;CV2 features:\n<a href=\"https://www.dropbox.com/s/0mfgb90m4iggn6i/train_img_features_v1.csv.gz\">https://www.dropbox.com/s/0mfgb90m4iggn6i/train_img_features_v1.csv.gz</a>\n<a href=\"https://www.dropbox.com/s/vr1fv94k2fukd89/test_img_features_v1.csv.gz\">https://www.dropbox.com/s/vr1fv94k2fukd89/test_img_features_v1.csv.gz</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 345470,
          "author_name": "Leonid Sinev",
          "author_url": "",
          "post_date": "2018-06-20T00:01:34.453000",
          "content": "<p>Thank you!</p>\n\n<blockquote>\n  <p>tfidf with 1-4 ngram ('3131473e84a9' and '75ebe6b373ec' were dropped)</p>\n</blockquote>\n\n<p>Seems like hundred more item_id's were dropped.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345757,
          "author_name": "Taras Baranyuk",
          "author_url": "",
          "post_date": "2018-06-20T11:59:44.237000",
          "content": "<p>train_data[pd.to_datetime(train_data.activation_date) &lt;= pd.to_datetime('2017-03-28')]</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 331017,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2018-05-20T07:18:50.883000",
      "content": "<p>List of Russian cities on Wikipedia, with census information:\n<a href=\"https://en.wikipedia.org/wiki/List_of_cities_and_towns_in_Russia_by_population\">https://en.wikipedia.org/wiki/List_of_cities_and_towns_in_Russia_by_population</a></p>\n\n<p>See also this kernel: <a href=\"https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia\">https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 331371,
          "author_name": "Vishy",
          "author_url": "",
          "post_date": "2018-05-21T05:27:58.203000",
          "content": "<p>Nice work !! May be a naive question how does this data get approved to be used in the competition oe when do we get to know if this data can be used. Please suggest !!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331404,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-21T07:04:29.760000",
          "content": "<p>creating a public kernel like <a href=\"https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\">region-and-city-details</a> with the collected data attached should be sufficient. However, @Sohier Dane please correct me if a am wrong</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331419,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2018-05-21T07:36:37.637000",
          "content": "<p>I posted here precisely to have feedback from the organizers. If they don't reply, I will take it as approval</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331628,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-05-21T15:51:42.647000",
          "content": "<p>@Dieter is correct. External data is valid as long as it is available to all competitors and is listed in this thread. </p>\n\n<p>Making data available to other competitors via a kernel like <a href=\"https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\">region-and-city-details</a> is great.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 331761,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2018-05-21T20:22:58.917000",
          "content": "<p>Thanks @Sohier Dane for the clarification. I added the file under \"Data\" in my script:\n<a href=\"https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia/data\">https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia/data</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 345940,
      "author_name": "Olga Ivanova",
      "author_url": "",
      "post_date": "2018-06-20T19:05:30.647000",
      "content": "<p>Distributed training framework for Tensorflow and Pytorch \n<a href=\"https://github.com/uber/horovod\">https://github.com/uber/horovod</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 345790,
      "author_name": "daniilmaltsev",
      "author_url": "",
      "post_date": "2018-06-20T13:07:51.207000",
      "content": "<p>Wages in regions:\n<a href=\"http://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/wages/\">http://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/wages/</a>\n<a href=\"http://www.gks.ru/free_doc/new_site/population/trud/sr-zarplata/t2.xlsx\">http://www.gks.ru/free_doc/new_site/population/trud/sr-zarplata/t2.xlsx</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 345174,
      "author_name": "Andrea Rapuzzi",
      "author_url": "",
      "post_date": "2018-06-19T10:47:29.953000",
      "content": "<p>Fast.ai models and weights (<a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>, <a href=\"http://files.fast.ai/models/weights.tgz\">http://files.fast.ai/models/weights.tgz</a>).\nPytorch models and weights.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 345149,
      "author_name": "Valentina Biryukova",
      "author_url": "",
      "post_date": "2018-06-19T09:22:18.867000",
      "content": "<p>Hello!\nWe want to use NIMA pretrained models (<a href=\"https://github.com/titu1994/neural-image-assessment\">https://github.com/titu1994/neural-image-assessment</a>).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 344679,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2018-06-18T14:05:56.137000",
      "content": "<p><a href=\"https://pytorch.org/docs/stable/torchvision/models.html\">https://pytorch.org/docs/stable/torchvision/models.html</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 322917,
      "author_name": "Matthew Anderson",
      "author_url": "",
      "post_date": "2018-05-03T23:11:14.487000",
      "content": "<p>I don't know about other competitors, but I think hourly site traffic data from Avito may be a helpful feature for our model.  Alexa Rank improved my model, and something more detailed may be even more helpful.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 324737,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-07T23:05:52.730000",
          "content": "<p>How did you get Alexa data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326892,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-10T13:21:45.653000",
          "content": "<p>Worst comes to worst, just gotta manually recreate it. Here's the Alexa ranks for avito.ru</p>\n\n<p><img src=\"https://traffic.alexa.com/graph?u=avito.ru\" alt=\"avito alexa rank\"></p>\n\n<p>When I get desperate enough / exhaust my data modeling, I'll go ahead and convert it to a csv and share in a kernel if no one else has done it by that point.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 328906,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-15T10:49:51.287000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331631,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-05-21T15:54:30.427000",
          "content": "<p><a href=\"/matthewa313\">@matthewa313</a> Is the Alexa data freely available? If it’s not freely available to everyone in the competition, and for the sponsors to use in their models after the competition, it can’t be used in this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331677,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-05-21T17:13:39.670000",
          "content": "<p><a href=\"/matthewa313\">@matthewa313</a> Based on the information provided by <a href=\"/authman\">@authman</a> in <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57138\">this thread</a> Alexa data is not sufficiently open for use in this competition. The two key facts here are that only a limited amount of the data is available for free and that all of the data requires a login.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 345094,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2018-06-19T07:37:56.427000",
      "content": "<p>Currency exchange rates: <a href=\"http://cbr.ru/eng/\">http://cbr.ru/eng/</a>\nExample: <a href=\"http://cbr.ru/eng/currency_base/daily/?date_req=15.03.2017\">http://cbr.ru/eng/currency_base/daily/?date_req=15.03.2017</a></p>\n\n<p>Weather data: <a href=\"https://www.gismeteo.ru/\">https://www.gismeteo.ru/</a>\nExample: <a href=\"https://www.gismeteo.ru/diary/4720/2017/3/\">https://www.gismeteo.ru/diary/4720/2017/3/</a></p>\n\n<p>ResNet152 model for Keras: \n<a href=\"https://github.com/broadinstitute/keras-resnet/blob/master/keras_resnet/models/_2d.py\">https://github.com/broadinstitute/keras-resnet/blob/master/keras_resnet/models/_2d.py</a>\nWeights: <a href=\"https://github.com/fizyr/keras-models/releases/tag/v0.0.1\">https://github.com/fizyr/keras-models/releases/tag/v0.0.1</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 345066,
      "author_name": "thousandvoices",
      "author_url": "",
      "post_date": "2018-06-19T06:25:09.157000",
      "content": "<p>Pretrained russian word vectors from <a href=\"https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md\">this page</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 343181,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2018-06-14T20:54:49.737000",
      "content": "<p>Scraping avito.ru ads</p>",
      "votes": -3,
      "replies": [
        {
          "id": 343358,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-15T07:56:43.887000",
          "content": "<p>not allowed, as long as you do not share the scraped data to everybody. See one post above</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343361,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2018-06-15T08:05:39.667000",
          "content": "<p>The rules only require you to share the source of the data in the discussion, not the data itself:</p>\n\n<pre><code>The following provision supersedes General Rules Section 7.C. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required.\" If you are using a source of external data, **you must post the source to the official external data** forum thread no later than one week prior to the deadline.\n</code></pre>\n\n<p>Note that Im still exploring (haven't really got any data yet) but <a href=\"/sohier\">@sohier</a> said \"External data is valid as long as it is available to all competitors and is listed in this thread.\" which I read as a necessary condition whereas \"Making data available to other competitors via a kernel like region-and-city-details is great.\" I read  as nice to have but not required.</p>\n\n<p><a href=\"/sohier\">@sohier</a> please clarify.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343387,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-15T09:06:53.903000",
          "content": "<p>From my understanding others need to be able to reproduce your dataset. With scraping, since the data is not really static, your exact data is not \"available to all competitors\". Also the \"right and authority to use such external data\" is a grey area for scraping.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343392,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2018-06-15T09:14:12.423000",
          "content": "<p>They mentioned source of the data, not data itself.</p>\n\n<p>FWIW, in another competition top winner of IEEE Camera competitition <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49367\">https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49367</a> scraped 300+ Gb of photos from Yandex and it wasn't an issue for Kaggle.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343565,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-06-15T15:23:31.963000",
          "content": "<p>To be clear: scraping data is prohibited by the terms and conditions of Avito's website . Any data acquired by scraping Avito is therefore not generally available and is against the competition rules.</p>\n\n<p>We can't put you guys into a situation where you need to violate the rules of other websites in order to stay competitive.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 343568,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2018-06-15T15:30:26.437000",
          "content": "<p>Thanks for the quick clarification.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343632,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-06-15T17:46:25",
          "content": "<p>Happy to help!</p>\n\n<p>To be clear, the IEEE competition had different (and unusual) rules regarding external data. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349022,
      "author_name": "F.J.Martinez-de-Pison",
      "author_url": "",
      "post_date": "2018-06-27T17:11:39.863000",
      "content": "<p>On be half of my team, all the external data we use are already mentioned here.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 346940,
      "author_name": "Leonid Sinev",
      "author_url": "",
      "post_date": "2018-06-22T20:51:05.683000",
      "content": "<p>cryptocurrency data from <a href=\"https://www.kaggle.com/mczielinski/bitcoin-historical-data/data\">https://www.kaggle.com/mczielinski/bitcoin-historical-data/data</a>\nand\n<a href=\"https://www.kaggle.com/kingburrito666/ethereum-historical-data/data\">https://www.kaggle.com/kingburrito666/ethereum-historical-data/data</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 346796,
      "author_name": "takuoko",
      "author_url": "",
      "post_date": "2018-06-22T11:58:25.997000",
      "content": "<p>I have a question. Is it allowed to use external data other participants mentioned on this discussion forum.\n<a href=\"https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\">https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters</a>\nfor example this kernel's data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 346842,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-06-22T14:43:23.567000",
          "content": "<p>Correct, you don't need to repeat an item that's already been listed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346984,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-06-22T23:47:25.103000",
          "content": "<p>Thank you very much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 346530,
      "author_name": "khyeh",
      "author_url": "",
      "post_date": "2018-06-21T22:17:13.027000",
      "content": "<p>All posted already~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 346048,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2018-06-21T01:23:03.753000",
      "content": "<p>On be half of my team, all the external data we use are already mentioned here :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 346008,
      "author_name": "Peter Hurford",
      "author_url": "",
      "post_date": "2018-06-20T22:21:26.980000",
      "content": "<p>On behalf of our team, I warrant that all external data used has either already been mentioned in this thread, is available in a publicly shared kernel here, or is contained in this following freely available public CSV file: <a href=\"https://s3.amazonaws.com/avito-demand-kaggle/region_macro.csv\">https://s3.amazonaws.com/avito-demand-kaggle/region_macro.csv</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 345979,
      "author_name": "Matt Motoki",
      "author_url": "",
      "post_date": "2018-06-20T21:05:20.837000",
      "content": "<p><a href=\"https://keras.io/applications/\">https://keras.io/applications/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 345794,
      "author_name": "daniilmaltsev",
      "author_url": "",
      "post_date": "2018-06-20T13:17:01.727000",
      "content": "<p>The Demographic Yearbook of Russia (not sure I'll use it, but it may be worth of giving a try):\n<a href=\"http://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/publications/catalog/doc_1137674209312\">http://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/publications/catalog/doc_1137674209312</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 345355,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2018-06-19T18:25:37.347000",
      "content": "<p>keras pretrained image models - resnet, densnet, vgg19 <br>\nimagenet synsets <br>\nfasttext russian pretrain vectors &amp; <a href=\"http://panchenko.me/data/dsl-backup/w2v-ru/\">http://panchenko.me/data/dsl-backup/w2v-ru/</a> <br>\nnltk &amp; pymorph <br>\nlat lon coords shared in kernels     </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 343869,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-16T11:43:03.053000",
      "content": "<p>Hi, we are making use of the Forbes top 2000 ranking. <a href=\"https://www.forbes.com/global2000/#6507469f335d\">https://www.forbes.com/global2000/#6507469f335d</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "326798": "----------\n\nText Models\n\n[Fasttext trained on wikipedia (300d)][1]\n\n[Fasttext trained on CommonCrawl (300d)][2]\n\n[nltk4russian][3]\n\n[Redhenlab][4]\n\n[Russian w2v][5]\n\n----------\nImage Models\n\nAll from [keras.applications][6]\n\nneural-image-assessment [git repo][7]\n\nILGnet [git repo][8]\n\n\n  [1]: https://s3-us-west-1.amazonaws.com/fasttext-vectors/wiki.ru.vec\n  [2]: https://s3-us-west-1.amazonaws.com/fasttext-vectors/word-vectors-v2/cc.ru.300.vec.gz\n  [3]: https://github.com/named-entity/nltk4russian\n  [4]: http://www.redhenlab.org/home/the-cognitive-core-research-topics-in-red-hen/the-barnyard/russian-nlp\n  [5]: https://github.com/nlpub/russe-evaluation/tree/master/russe/measures/word2vec\n  [6]: https://keras.io/applications/\n  [7]: https://github.com/titu1994/neural-image-assessment\n  [8]: https://github.com/BestiVictory/ILGnet",
    "322418": "Post your (freely available) external data links here.",
    "345145": "tfidf with 1-4 ngram ('3131473e84a9' and '75ebe6b373ec' were dropped):\nhttps://www.dropbox.com/s/hrbyfvmlsfmikwl/test_tfidf_sparse_1_4_clean_data_v1.npz\nhttps://www.dropbox.com/s/3xwnixrjexiz7qd/train_tfidf_sparse_1_4_clean_data_v1.npz\n\nPIL&amp;CV2 features:\nhttps://www.dropbox.com/s/0mfgb90m4iggn6i/train_img_features_v1.csv.gz\nhttps://www.dropbox.com/s/vr1fv94k2fukd89/test_img_features_v1.csv.gz",
    "331017": "List of Russian cities on Wikipedia, with census information:\nhttps://en.wikipedia.org/wiki/List_of_cities_and_towns_in_Russia_by_population\n\nSee also this kernel: https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia",
    "345940": "Distributed training framework for Tensorflow and Pytorch \nhttps://github.com/uber/horovod",
    "345790": "Wages in regions:\nhttp://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/wages/\nhttp://www.gks.ru/free_doc/new_site/population/trud/sr-zarplata/t2.xlsx",
    "345174": "Fast.ai models and weights (https://github.com/fastai/fastai, http://files.fast.ai/models/weights.tgz).\nPytorch models and weights.",
    "345149": "Hello!\nWe want to use NIMA pretrained models (https://github.com/titu1994/neural-image-assessment).",
    "344679": "https://pytorch.org/docs/stable/torchvision/models.html",
    "322917": "I don't know about other competitors, but I think hourly site traffic data from Avito may be a helpful feature for our model.  Alexa Rank improved my model, and something more detailed may be even more helpful.",
    "345094": "Currency exchange rates: http://cbr.ru/eng/\nExample: http://cbr.ru/eng/currency_base/daily/?date_req=15.03.2017\n\nWeather data: https://www.gismeteo.ru/\nExample: https://www.gismeteo.ru/diary/4720/2017/3/\n\nResNet152 model for Keras: \nhttps://github.com/broadinstitute/keras-resnet/blob/master/keras_resnet/models/_2d.py\nWeights: https://github.com/fizyr/keras-models/releases/tag/v0.0.1\n",
    "345066": "Pretrained russian word vectors from [this page][1]\n\n\n  [1]: https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md",
    "343181": "Scraping avito.ru ads",
    "349022": "On be half of my team, all the external data we use are already mentioned here.",
    "346940": "cryptocurrency data from https://www.kaggle.com/mczielinski/bitcoin-historical-data/data\nand\nhttps://www.kaggle.com/kingburrito666/ethereum-historical-data/data",
    "346796": "I have a question. Is it allowed to use external data other participants mentioned on this discussion forum.\nhttps://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\nfor example this kernel's data.",
    "346530": "All posted already~",
    "346048": "On be half of my team, all the external data we use are already mentioned here :)",
    "346008": "On behalf of our team, I warrant that all external data used has either already been mentioned in this thread, is available in a publicly shared kernel here, or is contained in this following freely available public CSV file: https://s3.amazonaws.com/avito-demand-kaggle/region_macro.csv",
    "345979": "https://keras.io/applications/",
    "345794": "The Demographic Yearbook of Russia (not sure I'll use it, but it may be worth of giving a try):\nhttp://www.gks.ru/wps/wcm/connect/rosstat_main/rosstat/ru/statistics/publications/catalog/doc_1137674209312",
    "345355": "keras pretrained image models - resnet, densnet, vgg19   \nimagenet synsets   \nfasttext russian pretrain vectors &amp; http://panchenko.me/data/dsl-backup/w2v-ru/    \nnltk &amp; pymorph    \nlat lon coords shared in kernels     ",
    "343869": "Hi, we are making use of the Forbes top 2000 ranking. https://www.forbes.com/global2000/#6507469f335d\n"
  }
}