{
  "id": 307409,
  "title": "Is there anyone like me who wanna get the transaction table indicating the index of customer_id?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/307409",
  "author_name": "",
  "post_date": "2022-02-14T05:44:30.817070900Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Example is below.</p>\n<p><br>\ncustomer_id1    ---&gt;      3<br>\ncustomer_id3    ---&gt;     10<br>\netc…</p>\n<p>I am going to use tensorflow recommender model.<br>\nBut because of memory and <strong>searching time cost</strong>, I can't use recommender model.<br>\nSo I let one kernel do transforming the customer_id in transaction table to the index of customer table.<br>\nIs there any one who got problem like me? I will share the output after the work.</p>",
  "messages": [
    {
      "id": "1689256",
      "postDate": "02/14/2022 05:44:30",
      "content": "<p>Example is below.</p>\n<p><br>\ncustomer_id1    ---&gt;      3<br>\ncustomer_id3    ---&gt;     10<br>\netc…</p>\n<p>I am going to use tensorflow recommender model.<br>\nBut because of memory and <strong>searching time cost</strong>, I can't use recommender model.<br>\nSo I let one kernel do transforming the customer_id in transaction table to the index of customer table.<br>\nIs there any one who got problem like me? I will share the output after the work.</p>",
      "rawMarkdown": "Example is below.\n\n<Transaction Table>\ncustomer_id1    --->      3\ncustomer_id3    --->     10\netc...\n\nI am going to use tensorflow recommender model.\nBut because of memory and **searching time cost**, I can't use recommender model.\nSo I let one kernel do transforming the customer_id in transaction table to the index of customer table.\nIs there any one who got problem like me? I will share the output after the work.",
      "votes": null
    },
    {
      "id": "1689353",
      "postDate": "02/14/2022 07:06:39",
      "content": "<p>I did something like that with data (not with tf)<br>\n<code>customers['CID'] = customers.index+1</code><br>\n<code>d = transactions.set_index(['customer_id']).join(customers[['customer_id','CID']].set_index(['customer_id']))</code><br>\n<code>d.reset_index(drop=True, inplace=True)</code></p>",
      "rawMarkdown": "I did something like that with data (not with tf)\n`customers['CID'] = customers.index+1 `\n`d = transactions.set_index(['customer_id']).join(customers[['customer_id','CID']].set_index(['customer_id'])) `\n`d.reset_index(drop=True, inplace=True)`",
      "votes": null
    },
    {
      "id": "1698720",
      "postDate": "02/20/2022 15:54:53",
      "content": "<p>Here's some simple code to do it:</p>\n<pre><code>id_to_index_dict = dict(zip(customers[\"customer_id\"], customers.index))\nindex_to_id_dict = dict(zip(customers.index, customers[\"customer_id\"]))\n\ntransactions[\"customer_id\"] = transactions[\"customer_id\"].map(id_to_index_dict)\n\n# when you want to make submission\nsub[\"customer_id\"] = sub[\"customer_id\"].map(index_to_id_dict)\n</code></pre>",
      "rawMarkdown": "Here's some simple code to do it:\n\n```\nid_to_index_dict = dict(zip(customers[\"customer_id\"], customers.index))\nindex_to_id_dict = dict(zip(customers.index, customers[\"customer_id\"]))\n\ntransactions[\"customer_id\"] = transactions[\"customer_id\"].map(id_to_index_dict)\n\n# when you want to make submission\nsub[\"customer_id\"] = sub[\"customer_id\"].map(index_to_id_dict)\n```",
      "votes": null
    },
    {
      "id": "1707495",
      "postDate": "02/28/2022 14:07:33",
      "content": "<p>I found that pandas <strong>(also RAPIDS cudf)</strong> 'get_loc' method algorithm is very fast for searching index.<br>\nIf you want fast searching algorithm, <strong>using the pandas 'get_loc' method will be a good choice !</strong></p>",
      "rawMarkdown": "I found that pandas **(also RAPIDS cudf)** 'get_loc' method algorithm is very fast for searching index.\nIf you want fast searching algorithm, **using the pandas 'get_loc' method will be a good choice !**",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1689353,
      "author_name": "aekaap",
      "author_url": "",
      "post_date": "02/14/2022 07:06:39",
      "content": "<p>I did something like that with data (not with tf)<br>\n<code>customers['CID'] = customers.index+1</code><br>\n<code>d = transactions.set_index(['customer_id']).join(customers[['customer_id','CID']].set_index(['customer_id']))</code><br>\n<code>d.reset_index(drop=True, inplace=True)</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1698720,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "02/20/2022 15:54:53",
      "content": "<p>Here's some simple code to do it:</p>\n<pre><code>id_to_index_dict = dict(zip(customers[\"customer_id\"], customers.index))\nindex_to_id_dict = dict(zip(customers.index, customers[\"customer_id\"]))\n\ntransactions[\"customer_id\"] = transactions[\"customer_id\"].map(id_to_index_dict)\n\n# when you want to make submission\nsub[\"customer_id\"] = sub[\"customer_id\"].map(index_to_id_dict)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1707495,
      "author_name": "cafelatte1",
      "author_url": "",
      "post_date": "02/28/2022 14:07:33",
      "content": "<p>I found that pandas <strong>(also RAPIDS cudf)</strong> 'get_loc' method algorithm is very fast for searching index.<br>\nIf you want fast searching algorithm, <strong>using the pandas 'get_loc' method will be a good choice !</strong></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1689256": "Example is below.\n\n<Transaction Table>\ncustomer_id1    --->      3\ncustomer_id3    --->     10\netc...\n\nI am going to use tensorflow recommender model.\nBut because of memory and **searching time cost**, I can't use recommender model.\nSo I let one kernel do transforming the customer_id in transaction table to the index of customer table.\nIs there any one who got problem like me? I will share the output after the work.",
    "1689353": "I did something like that with data (not with tf)\n`customers['CID'] = customers.index+1 `\n`d = transactions.set_index(['customer_id']).join(customers[['customer_id','CID']].set_index(['customer_id'])) `\n`d.reset_index(drop=True, inplace=True)`",
    "1698720": "Here's some simple code to do it:\n\n```\nid_to_index_dict = dict(zip(customers[\"customer_id\"], customers.index))\nindex_to_id_dict = dict(zip(customers.index, customers[\"customer_id\"]))\n\ntransactions[\"customer_id\"] = transactions[\"customer_id\"].map(id_to_index_dict)\n\n# when you want to make submission\nsub[\"customer_id\"] = sub[\"customer_id\"].map(index_to_id_dict)\n```",
    "1707495": "I found that pandas **(also RAPIDS cudf)** 'get_loc' method algorithm is very fast for searching index.\nIf you want fast searching algorithm, **using the pandas 'get_loc' method will be a good choice !**"
  },
  "source": "meta"
}