{
  "id": 57307,
  "title": "Exceeding Memory Limitations",
  "url": "/competitions/avito-demand-prediction/discussion/57307",
  "author_name": "",
  "post_date": "2018-05-22T11:37:23.656799900Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello, is anyone else running into memory issues with Kaggle kernels when trying to apply TFIDF to the textual data? I have combined the \"Title\" and \"Description\" fields into a single \"txt\" feature. I then try and run TFIDF on this (just using code that many other kernels on here are using) but it keeps exceeding the 17.2GB RAM limit and then the kernel dies so I never get any results.</p>\n\n<p>Does anyone have any suggestions? Would be preferred if it could be run on Kaggle rather than having to turn to e.g. AWS.</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "332031",
      "postDate": "05/22/2018 11:37:23",
      "content": "<p>Hello, is anyone else running into memory issues with Kaggle kernels when trying to apply TFIDF to the textual data? I have combined the \"Title\" and \"Description\" fields into a single \"txt\" feature. I then try and run TFIDF on this (just using code that many other kernels on here are using) but it keeps exceeding the 17.2GB RAM limit and then the kernel dies so I never get any results.</p>\n\n<p>Does anyone have any suggestions? Would be preferred if it could be run on Kaggle rather than having to turn to e.g. AWS.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hello, is anyone else running into memory issues with Kaggle kernels when trying to apply TFIDF to the textual data? I have combined the \"Title\" and \"Description\" fields into a single \"txt\" feature. I then try and run TFIDF on this (just using code that many other kernels on here are using) but it keeps exceeding the 17.2GB RAM limit and then the kernel dies so I never get any results.\n\nDoes anyone have any suggestions? Would be preferred if it could be run on Kaggle rather than having to turn to e.g. AWS.\n\nThanks",
      "votes": null
    },
    {
      "id": "332089",
      "postDate": "05/22/2018 13:50:24",
      "content": "<p>If you are using only train.csv and test.csv that should be doable. Did you try using \"del X\" on variable X that you don't use anymore ? and then gc.collect() to ask for a garbage collection ?</p>\n\n<p>If you want to run all data (train_active + test_active + the periods), then I don't think 17GB would be enough. I have been using AWS and it's very cheap for spot instances.</p>",
      "rawMarkdown": "If you are using only train.csv and test.csv that should be doable. Did you try using \"del X\" on variable X that you don't use anymore ? and then gc.collect() to ask for a garbage collection ?\n\nIf you want to run all data (train_active + test_active + the periods), then I don't think 17GB would be enough. I have been using AWS and it's very cheap for spot instances.",
      "votes": null
    },
    {
      "id": "332191",
      "postDate": "05/22/2018 18:17:38",
      "content": "<p>You can use the gensim package to perform tf-idf (and other transformations) in a memory friendly manner.</p>",
      "rawMarkdown": "You can use the gensim package to perform tf-idf (and other transformations) in a memory friendly manner.",
      "votes": null
    },
    {
      "id": "332528",
      "postDate": "05/23/2018 10:27:40",
      "content": "<p>you can load the data to AWS using wget commands, right? </p>\n\n<p>how did you unzip the image files - I keep getting an error that each of them is part of a group of zipped files. do i need to unzip them all at once or something?</p>",
      "rawMarkdown": "you can load the data to AWS using wget commands, right? \n\nhow did you unzip the image files - I keep getting an error that each of them is part of a group of zipped files. do i need to unzip them all at once or something?",
      "votes": null
    },
    {
      "id": "332529",
      "postDate": "05/23/2018 10:28:38",
      "content": "<p>what does this package do differently? what's its downside? otherwise presumably everyone would use it...</p>",
      "rawMarkdown": "what does this package do differently? what's its downside? otherwise presumably everyone would use it...",
      "votes": null
    },
    {
      "id": "332567",
      "postDate": "05/23/2018 12:15:10",
      "content": "<p>Not sure you mean by differently, to what exactly? It is widely used, for example, see a popular kernel in this competition, <a href=\"https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\">here</a>. And <a href=\"https://www.kaggle.com/c/word2vec-nlp-tutorial#part-2-word-vectors\">this</a> previous competition. Both of these use the gensim word2vec model but tutorials are on the package site for tf-idf. That said, having tried the tf-idf route I'm not convinced by it due to the huge resulting dimensionality (you can cap, obviously, but I'm also not convinced about this given the declensions in the Russian language).</p>\n\n<p>As to its downside, I'm no authority but I guess the main thing for most people is that it's not sklearn and so needs a little extra work to integrate in the workflow. </p>",
      "rawMarkdown": "Not sure you mean by differently, to what exactly? It is widely used, for example, see a popular kernel in this competition, [here][1]. And [this][2] previous competition. Both of these use the gensim word2vec model but tutorials are on the package site for tf-idf. That said, having tried the tf-idf route I'm not convinced by it due to the huge resulting dimensionality (you can cap, obviously, but I'm also not convinced about this given the declensions in the Russian language).\n\nAs to its downside, I'm no authority but I guess the main thing for most people is that it's not sklearn and so needs a little extra work to integrate in the workflow. \n\n  [1]: https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\n  [2]: https://www.kaggle.com/c/word2vec-nlp-tutorial#part-2-word-vectors",
      "votes": null
    },
    {
      "id": "332634",
      "postDate": "05/23/2018 13:40:14",
      "content": "<p>You can't use wget but there is a very easy tool to download all your files :\n<a href=\"https://github.com/Kaggle/kaggle-api\">https://github.com/Kaggle/kaggle-api</a></p>",
      "rawMarkdown": "You can't use wget but there is a very easy tool to download all your files :\nhttps://github.com/Kaggle/kaggle-api",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 332089,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "05/22/2018 13:50:24",
      "content": "<p>If you are using only train.csv and test.csv that should be doable. Did you try using \"del X\" on variable X that you don't use anymore ? and then gc.collect() to ask for a garbage collection ?</p>\n\n<p>If you want to run all data (train_active + test_active + the periods), then I don't think 17GB would be enough. I have been using AWS and it's very cheap for spot instances.</p>",
      "votes": null,
      "replies": [
        {
          "id": 332528,
          "author_name": "derrington",
          "author_url": "",
          "post_date": "05/23/2018 10:27:40",
          "content": "<p>you can load the data to AWS using wget commands, right? </p>\n\n<p>how did you unzip the image files - I keep getting an error that each of them is part of a group of zipped files. do i need to unzip them all at once or something?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332634,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "05/23/2018 13:40:14",
          "content": "<p>You can't use wget but there is a very easy tool to download all your files :\n<a href=\"https://github.com/Kaggle/kaggle-api\">https://github.com/Kaggle/kaggle-api</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332191,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "05/22/2018 18:17:38",
      "content": "<p>You can use the gensim package to perform tf-idf (and other transformations) in a memory friendly manner.</p>",
      "votes": null,
      "replies": [
        {
          "id": 332529,
          "author_name": "derrington",
          "author_url": "",
          "post_date": "05/23/2018 10:28:38",
          "content": "<p>what does this package do differently? what's its downside? otherwise presumably everyone would use it...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332567,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "05/23/2018 12:15:10",
          "content": "<p>Not sure you mean by differently, to what exactly? It is widely used, for example, see a popular kernel in this competition, <a href=\"https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\">here</a>. And <a href=\"https://www.kaggle.com/c/word2vec-nlp-tutorial#part-2-word-vectors\">this</a> previous competition. Both of these use the gensim word2vec model but tutorials are on the package site for tf-idf. That said, having tried the tf-idf route I'm not convinced by it due to the huge resulting dimensionality (you can cap, obviously, but I'm also not convinced about this given the declensions in the Russian language).</p>\n\n<p>As to its downside, I'm no authority but I guess the main thing for most people is that it's not sklearn and so needs a little extra work to integrate in the workflow. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "332031": "Hello, is anyone else running into memory issues with Kaggle kernels when trying to apply TFIDF to the textual data? I have combined the \"Title\" and \"Description\" fields into a single \"txt\" feature. I then try and run TFIDF on this (just using code that many other kernels on here are using) but it keeps exceeding the 17.2GB RAM limit and then the kernel dies so I never get any results.\n\nDoes anyone have any suggestions? Would be preferred if it could be run on Kaggle rather than having to turn to e.g. AWS.\n\nThanks",
    "332089": "If you are using only train.csv and test.csv that should be doable. Did you try using \"del X\" on variable X that you don't use anymore ? and then gc.collect() to ask for a garbage collection ?\n\nIf you want to run all data (train_active + test_active + the periods), then I don't think 17GB would be enough. I have been using AWS and it's very cheap for spot instances.",
    "332191": "You can use the gensim package to perform tf-idf (and other transformations) in a memory friendly manner.",
    "332528": "you can load the data to AWS using wget commands, right? \n\nhow did you unzip the image files - I keep getting an error that each of them is part of a group of zipped files. do i need to unzip them all at once or something?",
    "332529": "what does this package do differently? what's its downside? otherwise presumably everyone would use it...",
    "332567": "Not sure you mean by differently, to what exactly? It is widely used, for example, see a popular kernel in this competition, [here][1]. And [this][2] previous competition. Both of these use the gensim word2vec model but tutorials are on the package site for tf-idf. That said, having tried the tf-idf route I'm not convinced by it due to the huge resulting dimensionality (you can cap, obviously, but I'm also not convinced about this given the declensions in the Russian language).\n\nAs to its downside, I'm no authority but I guess the main thing for most people is that it's not sklearn and so needs a little extra work to integrate in the workflow. \n\n  [1]: https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\n  [2]: https://www.kaggle.com/c/word2vec-nlp-tutorial#part-2-word-vectors",
    "332634": "You can't use wget but there is a very easy tool to download all your files :\nhttps://github.com/Kaggle/kaggle-api"
  },
  "source": "meta"
}