{
  "id": 55503,
  "title": "External Data Thread",
  "url": "/competitions/avito-demand-prediction/discussion/55503",
  "author_name": "",
  "post_date": "2018-04-27T09:35:11.874439Z",
  "votes": 18,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Since there is not yet an official external data thread, but people might want to get an overview of pretrained models that can be used for this competition, I already start one. I personally go for a separation of models for the individual features so I also split my references here</p>\n\n<hr>\n\n<p>Text Models</p>\n\n<p><a href=\"https://s3-us-west-1.amazonaws.com/fasttext-vectors/wiki.ru.vec\">Fasttext trained on wikipedia (300d)</a></p>\n\n<p><a href=\"https://s3-us-west-1.amazonaws.com/fasttext-vectors/word-vectors-v2/cc.ru.300.vec.gz\">Fasttext trained on CommonCrawl (300d)</a></p>\n\n<p><a href=\"https://github.com/named-entity/nltk4russian\">nltk4russian</a></p>\n\n<p><a href=\"http://www.redhenlab.org/home/the-cognitive-core-research-topics-in-red-hen/the-barnyard/russian-nlp\">Redhenlab</a></p>\n\n<p><a href=\"https://github.com/nlpub/russe-evaluation/tree/master/russe/measures/word2vec\">Russian w2v</a></p>\n\n<hr>\n\n<p>Image Models</p>\n\n<p>All from <a href=\"https://keras.io/applications/\">keras.applications</a></p>",
  "messages": [
    {
      "id": "320016",
      "postDate": "04/27/2018 09:35:11",
      "content": "<p>Since there is not yet an official external data thread, but people might want to get an overview of pretrained models that can be used for this competition, I already start one. I personally go for a separation of models for the individual features so I also split my references here</p>\n\n<hr>\n\n<p>Text Models</p>\n\n<p><a href=\"https://s3-us-west-1.amazonaws.com/fasttext-vectors/wiki.ru.vec\">Fasttext trained on wikipedia (300d)</a></p>\n\n<p><a href=\"https://s3-us-west-1.amazonaws.com/fasttext-vectors/word-vectors-v2/cc.ru.300.vec.gz\">Fasttext trained on CommonCrawl (300d)</a></p>\n\n<p><a href=\"https://github.com/named-entity/nltk4russian\">nltk4russian</a></p>\n\n<p><a href=\"http://www.redhenlab.org/home/the-cognitive-core-research-topics-in-red-hen/the-barnyard/russian-nlp\">Redhenlab</a></p>\n\n<p><a href=\"https://github.com/nlpub/russe-evaluation/tree/master/russe/measures/word2vec\">Russian w2v</a></p>\n\n<hr>\n\n<p>Image Models</p>\n\n<p>All from <a href=\"https://keras.io/applications/\">keras.applications</a></p>",
      "rawMarkdown": "Since there is not yet an official external data thread, but people might want to get an overview of pretrained models that can be used for this competition, I already start one. I personally go for a separation of models for the individual features so I also split my references here\n\n----------\n\nText Models\n\n[Fasttext trained on wikipedia (300d)][1]\n\n[Fasttext trained on CommonCrawl (300d)][2]\n\n[nltk4russian][3]\n\n[Redhenlab][4]\n\n[Russian w2v][5]\n\n----------\nImage Models\n\nAll from [keras.applications][6]\n\n\n  [1]: https://s3-us-west-1.amazonaws.com/fasttext-vectors/wiki.ru.vec\n  [2]: https://s3-us-west-1.amazonaws.com/fasttext-vectors/word-vectors-v2/cc.ru.300.vec.gz\n  [3]: https://github.com/named-entity/nltk4russian\n  [4]: http://www.redhenlab.org/home/the-cognitive-core-research-topics-in-red-hen/the-barnyard/russian-nlp\n  [5]: https://github.com/nlpub/russe-evaluation/tree/master/russe/measures/word2vec\n  [6]: https://keras.io/applications/",
      "votes": null
    },
    {
      "id": "320128",
      "postDate": "04/27/2018 15:28:15",
      "content": "<p>These Russian Stopwords improved my model from no stop words: <a href=\"https://gist.github.com/menzenski/7047705\">https://gist.github.com/menzenski/7047705</a></p>\n\n<p>Past Avito.ru competitions have shown that making everything lowercase also helps tf-idf (although # of capital letters or uppercase:lowercase ratio may be a good idea).</p>",
      "rawMarkdown": "These Russian Stopwords improved my model from no stop words: https://gist.github.com/menzenski/7047705\n\nPast Avito.ru competitions have shown that making everything lowercase also helps tf-idf (although # of capital letters or uppercase:lowercase ratio may be a good idea).",
      "votes": null
    },
    {
      "id": "320427",
      "postDate": "04/28/2018 16:38:12",
      "content": "<p>Would you create a dataset in kaggle for fasttext pre-trained model files?</p>",
      "rawMarkdown": "Would you create a dataset in kaggle for fasttext pre-trained model files?",
      "votes": null
    },
    {
      "id": "326796",
      "postDate": "05/10/2018 11:02:06",
      "content": "<p>sure, you can get it here: <a href=\"https://www.kaggle.com/christofhenkel/fasttest-common-crawl-russian\">Fasttext Russian</a></p>",
      "rawMarkdown": "sure, you can get it here: [Fasttext Russian][1]\n\n\n  [1]: https://www.kaggle.com/christofhenkel/fasttest-common-crawl-russian",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 320128,
      "author_name": "matthewa313",
      "author_url": "",
      "post_date": "04/27/2018 15:28:15",
      "content": "<p>These Russian Stopwords improved my model from no stop words: <a href=\"https://gist.github.com/menzenski/7047705\">https://gist.github.com/menzenski/7047705</a></p>\n\n<p>Past Avito.ru competitions have shown that making everything lowercase also helps tf-idf (although # of capital letters or uppercase:lowercase ratio may be a good idea).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320427,
      "author_name": "classtag",
      "author_url": "",
      "post_date": "04/28/2018 16:38:12",
      "content": "<p>Would you create a dataset in kaggle for fasttext pre-trained model files?</p>",
      "votes": null,
      "replies": [
        {
          "id": 326796,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "05/10/2018 11:02:06",
          "content": "<p>sure, you can get it here: <a href=\"https://www.kaggle.com/christofhenkel/fasttest-common-crawl-russian\">Fasttext Russian</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "320016": "Since there is not yet an official external data thread, but people might want to get an overview of pretrained models that can be used for this competition, I already start one. I personally go for a separation of models for the individual features so I also split my references here\n\n----------\n\nText Models\n\n[Fasttext trained on wikipedia (300d)][1]\n\n[Fasttext trained on CommonCrawl (300d)][2]\n\n[nltk4russian][3]\n\n[Redhenlab][4]\n\n[Russian w2v][5]\n\n----------\nImage Models\n\nAll from [keras.applications][6]\n\n\n  [1]: https://s3-us-west-1.amazonaws.com/fasttext-vectors/wiki.ru.vec\n  [2]: https://s3-us-west-1.amazonaws.com/fasttext-vectors/word-vectors-v2/cc.ru.300.vec.gz\n  [3]: https://github.com/named-entity/nltk4russian\n  [4]: http://www.redhenlab.org/home/the-cognitive-core-research-topics-in-red-hen/the-barnyard/russian-nlp\n  [5]: https://github.com/nlpub/russe-evaluation/tree/master/russe/measures/word2vec\n  [6]: https://keras.io/applications/",
    "320128": "These Russian Stopwords improved my model from no stop words: https://gist.github.com/menzenski/7047705\n\nPast Avito.ru competitions have shown that making everything lowercase also helps tf-idf (although # of capital letters or uppercase:lowercase ratio may be a good idea).",
    "320427": "Would you create a dataset in kaggle for fasttext pre-trained model files?",
    "326796": "sure, you can get it here: [Fasttext Russian][1]\n\n\n  [1]: https://www.kaggle.com/christofhenkel/fasttest-common-crawl-russian"
  },
  "source": "meta"
}