{
  "id": 324179,
  "title": "Machine Choice and Memory Reduction Technique?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/324179",
  "author_name": "",
  "post_date": "2022-05-10T12:59:18.515633700Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>One of the big challenges for me in this competition was memory consumption. When I transitioned from a rule-based approach to a machine-learning-based, retrieval-ranker approach, I soon realized that the 16GB RAM provided in Kaggle Notebook was far from sufficient.</p>\n<p>So, I'd like to ask <strong>what you did to cope with the memory problem</strong>. I think there are two directions: <strong>using external computing resources</strong> and <strong>reducing memory consumption</strong>.</p>\n<p>For example, here is what our team did. I'm not sure if this is the best practice, but BigQuery was great because it made me forget about memory limitation.</p>\n<h1>Machine choice</h1>\n<ul>\n<li>Kaggle Notebook to generate customer-article pairs</li>\n<li>BigQuery to merge pairs with features<ul>\n<li>Created tables for static and dynamic features (i.e., data mart, feature store, whatever you call) </li>\n<li>Generated customer-article interaction features on the fly</li></ul></li>\n<li>Vertex AI Workbench (<code>e2-highmem-16</code>; 128GB RAM) to train the ranker<ul>\n<li>Had 60 million rows (pairs) and 33 columns (features) for two weeks of train set</li></ul></li>\n</ul>\n<h1>Memory reduction</h1>\n<ul>\n<li>Cast <code>customer_id</code> to int32</li>\n<li>Treated <code>article_id</code> as integer (not category) in LGBMRanker</li>\n</ul>",
  "messages": [
    {
      "id": "1783524",
      "postDate": "05/10/2022 12:59:18",
      "content": "<p>One of the big challenges for me in this competition was memory consumption. When I transitioned from a rule-based approach to a machine-learning-based, retrieval-ranker approach, I soon realized that the 16GB RAM provided in Kaggle Notebook was far from sufficient.</p>\n<p>So, I'd like to ask <strong>what you did to cope with the memory problem</strong>. I think there are two directions: <strong>using external computing resources</strong> and <strong>reducing memory consumption</strong>.</p>\n<p>For example, here is what our team did. I'm not sure if this is the best practice, but BigQuery was great because it made me forget about memory limitation.</p>\n<h1>Machine choice</h1>\n<ul>\n<li>Kaggle Notebook to generate customer-article pairs</li>\n<li>BigQuery to merge pairs with features<ul>\n<li>Created tables for static and dynamic features (i.e., data mart, feature store, whatever you call) </li>\n<li>Generated customer-article interaction features on the fly</li></ul></li>\n<li>Vertex AI Workbench (<code>e2-highmem-16</code>; 128GB RAM) to train the ranker<ul>\n<li>Had 60 million rows (pairs) and 33 columns (features) for two weeks of train set</li></ul></li>\n</ul>\n<h1>Memory reduction</h1>\n<ul>\n<li>Cast <code>customer_id</code> to int32</li>\n<li>Treated <code>article_id</code> as integer (not category) in LGBMRanker</li>\n</ul>",
      "rawMarkdown": "One of the big challenges for me in this competition was memory consumption. When I transitioned from a rule-based approach to a machine-learning-based, retrieval-ranker approach, I soon realized that the 16GB RAM provided in Kaggle Notebook was far from sufficient.\n\nSo, I'd like to ask **what you did to cope with the memory problem**. I think there are two directions: **using external computing resources** and **reducing memory consumption**.\n\nFor example, here is what our team did. I'm not sure if this is the best practice, but BigQuery was great because it made me forget about memory limitation.\n\n# Machine choice\n- Kaggle Notebook to generate customer-article pairs\n- BigQuery to merge pairs with features\n  - Created tables for static and dynamic features (i.e., data mart, feature store, whatever you call) \n  - Generated customer-article interaction features on the fly\n- Vertex AI Workbench (`e2-highmem-16`; 128GB RAM) to train the ranker\n  - Had 60 million rows (pairs) and 33 columns (features) for two weeks of train set\n\n# Memory reduction\n- Cast `customer_id` to int32\n- Treated `article_id` as integer (not category) in LGBMRanker",
      "votes": null
    },
    {
      "id": "1784108",
      "postDate": "05/11/2022 00:01:48",
      "content": "<p>I chose BigQuery too. Besides, I used Vertex Training to easily manage machine resource (finally used 512 GB RAM).<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324293\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324293</a></p>",
      "rawMarkdown": "I chose BigQuery too. Besides, I used Vertex Training to easily manage machine resource (finally used 512 GB RAM).\nhttps://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324293",
      "votes": null
    },
    {
      "id": "1784117",
      "postDate": "05/11/2022 00:15:11",
      "content": "<p>Vertex Training sounds nice! I'll give it a try next time!</p>",
      "rawMarkdown": "Vertex Training sounds nice! I'll give it a try next time!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1784108,
      "author_name": "keisuke07",
      "author_url": "",
      "post_date": "05/11/2022 00:01:48",
      "content": "<p>I chose BigQuery too. Besides, I used Vertex Training to easily manage machine resource (finally used 512 GB RAM).<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324293\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324293</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1784117,
          "author_name": "shionhonda",
          "author_url": "",
          "post_date": "05/11/2022 00:15:11",
          "content": "<p>Vertex Training sounds nice! I'll give it a try next time!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1783524": "One of the big challenges for me in this competition was memory consumption. When I transitioned from a rule-based approach to a machine-learning-based, retrieval-ranker approach, I soon realized that the 16GB RAM provided in Kaggle Notebook was far from sufficient.\n\nSo, I'd like to ask **what you did to cope with the memory problem**. I think there are two directions: **using external computing resources** and **reducing memory consumption**.\n\nFor example, here is what our team did. I'm not sure if this is the best practice, but BigQuery was great because it made me forget about memory limitation.\n\n# Machine choice\n- Kaggle Notebook to generate customer-article pairs\n- BigQuery to merge pairs with features\n  - Created tables for static and dynamic features (i.e., data mart, feature store, whatever you call) \n  - Generated customer-article interaction features on the fly\n- Vertex AI Workbench (`e2-highmem-16`; 128GB RAM) to train the ranker\n  - Had 60 million rows (pairs) and 33 columns (features) for two weeks of train set\n\n# Memory reduction\n- Cast `customer_id` to int32\n- Treated `article_id` as integer (not category) in LGBMRanker",
    "1784108": "I chose BigQuery too. Besides, I used Vertex Training to easily manage machine resource (finally used 512 GB RAM).\nhttps://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324293",
    "1784117": "Vertex Training sounds nice! I'll give it a try next time!"
  },
  "source": "meta"
}