{
  "id": 21455,
  "title": "R syntax for three-way join on itemPairsTrain.itemID_{1,2} == itemInfoTrain.itemID ?",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/21455",
  "author_name": "",
  "post_date": "2016-06-06T03:08:51.187Z",
  "votes": null,
  "comment_count": 2,
  "views": 328,
  "content": "<p>Sorry to ask an R syntax question here, but what's a simple (and performant) way to do a three-way merge/join/reshape between all these (and drop the other columns):</p>\n\n<ul>\n<li>itemPairsTrain.itemID_1, itemInfoTrain.itemID to join itemInfoTrain.categoryID\nas categoryID_1</li>\n<li>itemPairsTrain.itemID_2, itemInfoTrain.itemID to join itemInfoTrain.categoryID as categoryID_2</li>\n</ul>\n\n<p>(I did check SO and the dplyr and reshape doc)</p>",
  "messages": [
    {
      "id": "122644",
      "postDate": "06/06/2016 03:08:51",
      "content": "<p>Sorry to ask an R syntax question here, but what's a simple (and performant) way to do a three-way merge/join/reshape between all these (and drop the other columns):</p>\n\n<ul>\n<li>itemPairsTrain.itemID_1, itemInfoTrain.itemID to join itemInfoTrain.categoryID\nas categoryID_1</li>\n<li>itemPairsTrain.itemID_2, itemInfoTrain.itemID to join itemInfoTrain.categoryID as categoryID_2</li>\n</ul>\n\n<p>(I did check SO and the dplyr and reshape doc)</p>",
      "rawMarkdown": "Sorry to ask an R syntax question here, but what's a simple (and performant) way to do a three-way merge/join/reshape between all these (and drop the other columns):\r\n\r\n - itemPairsTrain.itemID_1, itemInfoTrain.itemID to join itemInfoTrain.categoryID\r\n   as categoryID_1\r\n - itemPairsTrain.itemID_2, itemInfoTrain.itemID to join itemInfoTrain.categoryID as categoryID_2\r\n\r\n(I did check SO and the dplyr and reshape doc)",
      "votes": null
    },
    {
      "id": "122646",
      "postDate": "06/06/2016 03:45:59",
      "content": "<p>Here's a dplyr way, using nested inner_joins:</p>\n\n<pre><code>itemPairsTrain.cats &lt;- inner_join(\n\n  inner_join(itemPairsTrain,\n             itemInfoTrain %&gt;% select(itemID_1=itemID, categoryID_1=categoryID),\n             by = 'itemID_1'),\n\n  itemInfoTrain %&gt;% select(itemID_2=itemID, categoryID_2=categoryID), by='itemID_2'\n\n)\n</code></pre>",
      "rawMarkdown": "Here's a dplyr way, using nested inner_joins:\r\n\r\n    itemPairsTrain.cats <- inner_join(\r\n\r\n      inner_join(itemPairsTrain,\r\n                 itemInfoTrain %>% select(itemID_1=itemID, categoryID_1=categoryID),\r\n                 by = 'itemID_1'),\r\n      \r\n      itemInfoTrain %>% select(itemID_2=itemID, categoryID_2=categoryID), by='itemID_2'\r\n\r\n    )",
      "votes": null
    },
    {
      "id": "122647",
      "postDate": "06/06/2016 03:54:24",
      "content": "<p>(Anyway this is all irrelevant since it turns out categoryID_1 == categoryID_2 throughout all of train and test.)</p>",
      "rawMarkdown": "(Anyway this is all irrelevant since it turns out categoryID_1 == categoryID_2 throughout all of train and test.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122646,
      "author_name": "smcinerney",
      "author_url": "",
      "post_date": "06/06/2016 03:45:59",
      "content": "<p>Here's a dplyr way, using nested inner_joins:</p>\n\n<pre><code>itemPairsTrain.cats &lt;- inner_join(\n\n  inner_join(itemPairsTrain,\n             itemInfoTrain %&gt;% select(itemID_1=itemID, categoryID_1=categoryID),\n             by = 'itemID_1'),\n\n  itemInfoTrain %&gt;% select(itemID_2=itemID, categoryID_2=categoryID), by='itemID_2'\n\n)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122647,
      "author_name": "smcinerney",
      "author_url": "",
      "post_date": "06/06/2016 03:54:24",
      "content": "<p>(Anyway this is all irrelevant since it turns out categoryID_1 == categoryID_2 throughout all of train and test.)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122644": "Sorry to ask an R syntax question here, but what's a simple (and performant) way to do a three-way merge/join/reshape between all these (and drop the other columns):\r\n\r\n - itemPairsTrain.itemID_1, itemInfoTrain.itemID to join itemInfoTrain.categoryID\r\n   as categoryID_1\r\n - itemPairsTrain.itemID_2, itemInfoTrain.itemID to join itemInfoTrain.categoryID as categoryID_2\r\n\r\n(I did check SO and the dplyr and reshape doc)",
    "122646": "Here's a dplyr way, using nested inner_joins:\r\n\r\n    itemPairsTrain.cats <- inner_join(\r\n\r\n      inner_join(itemPairsTrain,\r\n                 itemInfoTrain %>% select(itemID_1=itemID, categoryID_1=categoryID),\r\n                 by = 'itemID_1'),\r\n      \r\n      itemInfoTrain %>% select(itemID_2=itemID, categoryID_2=categoryID), by='itemID_2'\r\n\r\n    )",
    "122647": "(Anyway this is all irrelevant since it turns out categoryID_1 == categoryID_2 throughout all of train and test.)"
  },
  "source": "meta"
}