{
  "id": 201664,
  "title": "How do I handle a large numpy dataset with tensorflow.Dataset?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201664",
  "author_name": "nadare",
  "post_date": "2020-12-06T02:48:02.649000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I define a dataset and train it as follows.</p>\n<pre><code>train_data = tf.data.Dataset.from_tensor_slices((user_ids.astype(np.int32),\n                                                 content_ids.astype(np.int32),\n                                                 timestamps.astype(np.int64),\n                                                 features.astype(np.float32)))\n\npredictions = np.zeros(len(train_data))\nwith tqdm(total=len(train_data.batch(batch_size))) as pbar_train:\n    for i, data in enumerate(train_data.batch(batch_size).map(create_feature)):\n        pred = train_step(data)\n        predictions[i*batch_size:(i+1)*batch_size] = pred.numpy()\n        progress_text = \"Loss: {:.5f}, AUC: {:.5f}\".format(train_loss.result(), train_auc.result())\n        pbar_train.set_postfix_str(progress_text)\n        pbar_train.update(1)\n</code></pre>\n<p>However, this method will gradually consume memory because tensorflow will track the entire generator.<br>\nAny good ideas for working with numpy datasets in memory?</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 1103551,
      "postDate": "2020-12-06T02:48:02.650Z",
      "content": "<p>I define a dataset and train it as follows.</p>\n<pre><code>train_data = tf.data.Dataset.from_tensor_slices((user_ids.astype(np.int32),\n                                                 content_ids.astype(np.int32),\n                                                 timestamps.astype(np.int64),\n                                                 features.astype(np.float32)))\n\npredictions = np.zeros(len(train_data))\nwith tqdm(total=len(train_data.batch(batch_size))) as pbar_train:\n    for i, data in enumerate(train_data.batch(batch_size).map(create_feature)):\n        pred = train_step(data)\n        predictions[i*batch_size:(i+1)*batch_size] = pred.numpy()\n        progress_text = \"Loss: {:.5f}, AUC: {:.5f}\".format(train_loss.result(), train_auc.result())\n        pbar_train.set_postfix_str(progress_text)\n        pbar_train.update(1)\n</code></pre>\n<p>However, this method will gradually consume memory because tensorflow will track the entire generator.<br>\nAny good ideas for working with numpy datasets in memory?</p>\n<p>Thank you.</p>",
      "rawMarkdown": "I define a dataset and train it as follows.\n```\ntrain_data = tf.data.Dataset.from_tensor_slices((user_ids.astype(np.int32),\n                                                 content_ids.astype(np.int32),\n                                                 timestamps.astype(np.int64),\n                                                 features.astype(np.float32)))\n\npredictions = np.zeros(len(train_data))\nwith tqdm(total=len(train_data.batch(batch_size))) as pbar_train:\n    for i, data in enumerate(train_data.batch(batch_size).map(create_feature)):\n        pred = train_step(data)\n        predictions[i*batch_size:(i+1)*batch_size] = pred.numpy()\n        progress_text = \"Loss: {:.5f}, AUC: {:.5f}\".format(train_loss.result(), train_auc.result())\n        pbar_train.set_postfix_str(progress_text)\n        pbar_train.update(1)\n```\n\nHowever, this method will gradually consume memory because tensorflow will track the entire generator.\nAny good ideas for working with numpy datasets in memory?\n\nThank you."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1103551": "I define a dataset and train it as follows.\n```\ntrain_data = tf.data.Dataset.from_tensor_slices((user_ids.astype(np.int32),\n                                                 content_ids.astype(np.int32),\n                                                 timestamps.astype(np.int64),\n                                                 features.astype(np.float32)))\n\npredictions = np.zeros(len(train_data))\nwith tqdm(total=len(train_data.batch(batch_size))) as pbar_train:\n    for i, data in enumerate(train_data.batch(batch_size).map(create_feature)):\n        pred = train_step(data)\n        predictions[i*batch_size:(i+1)*batch_size] = pred.numpy()\n        progress_text = \"Loss: {:.5f}, AUC: {:.5f}\".format(train_loss.result(), train_auc.result())\n        pbar_train.set_postfix_str(progress_text)\n        pbar_train.update(1)\n```\n\nHowever, this method will gradually consume memory because tensorflow will track the entire generator.\nAny good ideas for working with numpy datasets in memory?\n\nThank you."
  }
}