{
  "id": 204941,
  "title": "Loading Features in Submission",
  "url": "/competitions/riiid-test-answer-prediction/discussion/204941",
  "author_name": "Nono",
  "post_date": "2020-12-17T15:22:40.733000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi all,<br>\nI have a dictionary which represents one feature, where the keys are strings representing 'userid_contentid'. The number of keys is &gt;80 millions coming from the training set.<br>\nWhen I load the dictionary in Kaggle, I search for a specific 'userid_contentid' key in the dictonary. Then I assign the corresponding value to a Numpy array  (for each chunk in the Submission file). Using Numpy and in-memory dictionary, my submission takes only 1 hour, the drawback is the Ram Usage as the dictionary memory size is 8GB. As I am adding more features, this method of loading a dictionary is not sustainable due to RAM limitations.</p>\n<p>I tried to use Dask instead of loading the dictionary in memory. <br>\nI was able to overcome the RAM limitation, however it takes 25 seconds  to filter a Dask Array (while it takes 5e-05 seconds to access a key in a dictionary). I am not sure if I am doing something wrong but Dask seems very slow and the submission might take more than the 9hr time limit.</p>\n<p>I am hearing people talking about having 40+ number of features which is very impressive.<br>\nI must be using the wrong method in terms of how to save/load/handle features.</p>\n<p>Any suggestions please? </p>\n<p>Many thanks</p>",
  "messages": [
    {
      "id": 1116946,
      "postDate": "2020-12-17T15:22:40.733Z",
      "content": "<p>Hi all,<br>\nI have a dictionary which represents one feature, where the keys are strings representing 'userid_contentid'. The number of keys is &gt;80 millions coming from the training set.<br>\nWhen I load the dictionary in Kaggle, I search for a specific 'userid_contentid' key in the dictonary. Then I assign the corresponding value to a Numpy array  (for each chunk in the Submission file). Using Numpy and in-memory dictionary, my submission takes only 1 hour, the drawback is the Ram Usage as the dictionary memory size is 8GB. As I am adding more features, this method of loading a dictionary is not sustainable due to RAM limitations.</p>\n<p>I tried to use Dask instead of loading the dictionary in memory. <br>\nI was able to overcome the RAM limitation, however it takes 25 seconds  to filter a Dask Array (while it takes 5e-05 seconds to access a key in a dictionary). I am not sure if I am doing something wrong but Dask seems very slow and the submission might take more than the 9hr time limit.</p>\n<p>I am hearing people talking about having 40+ number of features which is very impressive.<br>\nI must be using the wrong method in terms of how to save/load/handle features.</p>\n<p>Any suggestions please? </p>\n<p>Many thanks</p>",
      "rawMarkdown": "Hi all,\nI have a dictionary which represents one feature, where the keys are strings representing 'userid_contentid'. The number of keys is >80 millions coming from the training set.\nWhen I load the dictionary in Kaggle, I search for a specific 'userid_contentid' key in the dictonary. Then I assign the corresponding value to a Numpy array  (for each chunk in the Submission file). Using Numpy and in-memory dictionary, my submission takes only 1 hour, the drawback is the Ram Usage as the dictionary memory size is 8GB. As I am adding more features, this method of loading a dictionary is not sustainable due to RAM limitations.\n\nI tried to use Dask instead of loading the dictionary in memory. \nI was able to overcome the RAM limitation, however it takes 25 seconds  to filter a Dask Array (while it takes 5e-05 seconds to access a key in a dictionary). I am not sure if I am doing something wrong but Dask seems very slow and the submission might take more than the 9hr time limit.\n\nI am hearing people talking about having 40+ number of features which is very impressive.\nI must be using the wrong method in terms of how to save/load/handle features.\n\nAny suggestions please? \n\nMany thanks\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1116946": "Hi all,\nI have a dictionary which represents one feature, where the keys are strings representing 'userid_contentid'. The number of keys is >80 millions coming from the training set.\nWhen I load the dictionary in Kaggle, I search for a specific 'userid_contentid' key in the dictonary. Then I assign the corresponding value to a Numpy array  (for each chunk in the Submission file). Using Numpy and in-memory dictionary, my submission takes only 1 hour, the drawback is the Ram Usage as the dictionary memory size is 8GB. As I am adding more features, this method of loading a dictionary is not sustainable due to RAM limitations.\n\nI tried to use Dask instead of loading the dictionary in memory. \nI was able to overcome the RAM limitation, however it takes 25 seconds  to filter a Dask Array (while it takes 5e-05 seconds to access a key in a dictionary). I am not sure if I am doing something wrong but Dask seems very slow and the submission might take more than the 9hr time limit.\n\nI am hearing people talking about having 40+ number of features which is very impressive.\nI must be using the wrong method in terms of how to save/load/handle features.\n\nAny suggestions please? \n\nMany thanks\n"
  }
}