{
  "id": 24183,
  "title": "R Need MUCH faster way to create output",
  "url": "/competitions/outbrain-click-prediction/discussion/24183",
  "author_name": "",
  "post_date": "2016-10-07T22:24:38.633Z",
  "votes": null,
  "comment_count": 3,
  "views": 564,
  "content": "<p>I need help generating the output format. I create te2 which is a data.frame of display_id and ad_id sorted by ad click probability. I've tried every trick I know to speed up this lookup: I create a matrix, create an index, name the index rows. This takes incredibly long to run. I think that this needs an Rcpp solution, but I don't know c++ so I'm stuck.</p>\n\n<pre><code>te3 &lt;- as.matrix(te2[, c(1,2)])\n\nte2$idx &lt;- 1:nrow(te2)\n\nte3_idx &lt;- te2 %&gt;%\n  group_by(display_id) %&gt;%\n  summarize(start = min(idx),\n            end = max(idx)) %&gt;%\n  as.matrix()\n\nrow.names(te3_idx) &lt;- te3_idx[,1]\n\nfast_top_12 &lt;- function(id, dat = te3, index = te3_idx){\n  idx_st &lt;- index[id,2]\n  idx_en &lt;- index[id,3]\n  cnt &lt;- min(12, max(idx_en - idx_st + 1, 1))\n  return(paste(dat[idx_st:idx_en, 2], collapse = &quot; &quot;))\n}\n\ntest_list &lt;- as.character(unique(te2$display_id))\ntest_top12 &lt;- sapply(test_list, fast_top_12)\n</code></pre>",
  "messages": [
    {
      "id": "138282",
      "postDate": "10/07/2016 22:24:38",
      "content": "<p>I need help generating the output format. I create te2 which is a data.frame of display_id and ad_id sorted by ad click probability. I've tried every trick I know to speed up this lookup: I create a matrix, create an index, name the index rows. This takes incredibly long to run. I think that this needs an Rcpp solution, but I don't know c++ so I'm stuck.</p>\n\n<pre><code>te3 &lt;- as.matrix(te2[, c(1,2)])\n\nte2$idx &lt;- 1:nrow(te2)\n\nte3_idx &lt;- te2 %&gt;%\n  group_by(display_id) %&gt;%\n  summarize(start = min(idx),\n            end = max(idx)) %&gt;%\n  as.matrix()\n\nrow.names(te3_idx) &lt;- te3_idx[,1]\n\nfast_top_12 &lt;- function(id, dat = te3, index = te3_idx){\n  idx_st &lt;- index[id,2]\n  idx_en &lt;- index[id,3]\n  cnt &lt;- min(12, max(idx_en - idx_st + 1, 1))\n  return(paste(dat[idx_st:idx_en, 2], collapse = &quot; &quot;))\n}\n\ntest_list &lt;- as.character(unique(te2$display_id))\ntest_top12 &lt;- sapply(test_list, fast_top_12)\n</code></pre>",
      "rawMarkdown": "I need help generating the output format. I create te2 which is a data.frame of display_id and ad_id sorted by ad click probability. I've tried every trick I know to speed up this lookup: I create a matrix, create an index, name the index rows. This takes incredibly long to run. I think that this needs an Rcpp solution, but I don't know c++ so I'm stuck.\r\n\r\n\r\n\r\n    te3 <- as.matrix(te2[, c(1,2)])\r\n    \r\n    te2$idx <- 1:nrow(te2)\r\n    \r\n    te3_idx <- te2 %>%\r\n      group_by(display_id) %>%\r\n      summarize(start = min(idx),\r\n                end = max(idx)) %>%\r\n      as.matrix()\r\n    \r\n    row.names(te3_idx) <- te3_idx[,1]\r\n    \r\n    fast_top_12 <- function(id, dat = te3, index = te3_idx){\r\n      idx_st <- index[id,2]\r\n      idx_en <- index[id,3]\r\n      cnt <- min(12, max(idx_en - idx_st + 1, 1))\r\n      return(paste(dat[idx_st:idx_en, 2], collapse = \" \"))\r\n    }\r\n    \r\n    test_list <- as.character(unique(te2$display_id))\r\n    test_top12 <- sapply(test_list, fast_top_12)",
      "votes": null
    },
    {
      "id": "138807",
      "postDate": "10/10/2016 23:23:12",
      "content": "<p>FYI - I opened a thread with some suggestions on creating a submission file fast from R. actually, I do it outside R :)\nI still work on a faster R version. Will publish soon.</p>",
      "rawMarkdown": "FYI - I opened a thread with some suggestions on creating a submission file fast from R. actually, I do it outside R :)\r\nI still work on a faster R version. Will publish soon.",
      "votes": null
    },
    {
      "id": "138999",
      "postDate": "10/12/2016 01:20:42",
      "content": "<p>Thanks Amit! I look forward to your solution.</p>\n\n<p>idle_speculation posted a solution in <a href=\"https://www.kaggle.com/speculation/outbrain-click-prediction/btb-in-r/code\">this kernel</a>. I have no idea how the Reduce() function works so I have some reading to do.</p>\n\n<p><strong>idle_speculation's solution</strong></p>\n\n<pre><code>SUB=Reduce(\n    function(x,y){\n        M=merge(x,y,by='display_id',all.x=T)\n        M$ad_id.y[is.na(M$ad_id.y)]=&quot;&quot;\n        M[,list(display_id,ad_id=paste(ad_id.x,ad_id.y))]\n    }\n    ,lapply(1:max_length,function(k){\n        P[ad_rank==k,list(display_id,ad_id)]\n    })\n)\n</code></pre>",
      "rawMarkdown": "Thanks Amit! I look forward to your solution.\r\n\r\nidle_speculation posted a solution in [this kernel][1]. I have no idea how the Reduce() function works so I have some reading to do.\r\n\r\n\r\n**idle_speculation's solution**\r\n\r\n    SUB=Reduce(\r\n    \tfunction(x,y){\r\n    \t\tM=merge(x,y,by='display_id',all.x=T)\r\n    \t\tM$ad_id.y[is.na(M$ad_id.y)]=\"\"\r\n    \t\tM[,list(display_id,ad_id=paste(ad_id.x,ad_id.y))]\r\n    \t}\r\n    \t,lapply(1:max_length,function(k){\r\n    \t\tP[ad_rank==k,list(display_id,ad_id)]\r\n    \t})\r\n    )\r\n\r\n\r\n\r\n  [1]: https://www.kaggle.com/speculation/outbrain-click-prediction/btb-in-r/code",
      "votes": null
    },
    {
      "id": "139008",
      "postDate": "10/12/2016 02:52:32",
      "content": "<p>Take a look at this <a href=\"https://www.kaggle.com/jpuigde/outbrain-click-prediction/simple-r/discussion\">script</a> instead.  Jpuigde obtains the same result in quicker and more intuitive way on lines 32-34.</p>",
      "rawMarkdown": "Take a look at this [script][1] instead.  Jpuigde obtains the same result in quicker and more intuitive way on lines 32-34.\r\n\r\n\r\n  [1]: https://www.kaggle.com/jpuigde/outbrain-click-prediction/simple-r/discussion",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 138807,
      "author_name": "amitgal",
      "author_url": "",
      "post_date": "10/10/2016 23:23:12",
      "content": "<p>FYI - I opened a thread with some suggestions on creating a submission file fast from R. actually, I do it outside R :)\nI still work on a faster R version. Will publish soon.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138999,
      "author_name": "jeffhebert",
      "author_url": "",
      "post_date": "10/12/2016 01:20:42",
      "content": "<p>Thanks Amit! I look forward to your solution.</p>\n\n<p>idle_speculation posted a solution in <a href=\"https://www.kaggle.com/speculation/outbrain-click-prediction/btb-in-r/code\">this kernel</a>. I have no idea how the Reduce() function works so I have some reading to do.</p>\n\n<p><strong>idle_speculation's solution</strong></p>\n\n<pre><code>SUB=Reduce(\n    function(x,y){\n        M=merge(x,y,by='display_id',all.x=T)\n        M$ad_id.y[is.na(M$ad_id.y)]=&quot;&quot;\n        M[,list(display_id,ad_id=paste(ad_id.x,ad_id.y))]\n    }\n    ,lapply(1:max_length,function(k){\n        P[ad_rank==k,list(display_id,ad_id)]\n    })\n)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139008,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "10/12/2016 02:52:32",
      "content": "<p>Take a look at this <a href=\"https://www.kaggle.com/jpuigde/outbrain-click-prediction/simple-r/discussion\">script</a> instead.  Jpuigde obtains the same result in quicker and more intuitive way on lines 32-34.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "138282": "I need help generating the output format. I create te2 which is a data.frame of display_id and ad_id sorted by ad click probability. I've tried every trick I know to speed up this lookup: I create a matrix, create an index, name the index rows. This takes incredibly long to run. I think that this needs an Rcpp solution, but I don't know c++ so I'm stuck.\r\n\r\n\r\n\r\n    te3 <- as.matrix(te2[, c(1,2)])\r\n    \r\n    te2$idx <- 1:nrow(te2)\r\n    \r\n    te3_idx <- te2 %>%\r\n      group_by(display_id) %>%\r\n      summarize(start = min(idx),\r\n                end = max(idx)) %>%\r\n      as.matrix()\r\n    \r\n    row.names(te3_idx) <- te3_idx[,1]\r\n    \r\n    fast_top_12 <- function(id, dat = te3, index = te3_idx){\r\n      idx_st <- index[id,2]\r\n      idx_en <- index[id,3]\r\n      cnt <- min(12, max(idx_en - idx_st + 1, 1))\r\n      return(paste(dat[idx_st:idx_en, 2], collapse = \" \"))\r\n    }\r\n    \r\n    test_list <- as.character(unique(te2$display_id))\r\n    test_top12 <- sapply(test_list, fast_top_12)",
    "138807": "FYI - I opened a thread with some suggestions on creating a submission file fast from R. actually, I do it outside R :)\r\nI still work on a faster R version. Will publish soon.",
    "138999": "Thanks Amit! I look forward to your solution.\r\n\r\nidle_speculation posted a solution in [this kernel][1]. I have no idea how the Reduce() function works so I have some reading to do.\r\n\r\n\r\n**idle_speculation's solution**\r\n\r\n    SUB=Reduce(\r\n    \tfunction(x,y){\r\n    \t\tM=merge(x,y,by='display_id',all.x=T)\r\n    \t\tM$ad_id.y[is.na(M$ad_id.y)]=\"\"\r\n    \t\tM[,list(display_id,ad_id=paste(ad_id.x,ad_id.y))]\r\n    \t}\r\n    \t,lapply(1:max_length,function(k){\r\n    \t\tP[ad_rank==k,list(display_id,ad_id)]\r\n    \t})\r\n    )\r\n\r\n\r\n\r\n  [1]: https://www.kaggle.com/speculation/outbrain-click-prediction/btb-in-r/code",
    "139008": "Take a look at this [script][1] instead.  Jpuigde obtains the same result in quicker and more intuitive way on lines 32-34.\r\n\r\n\r\n  [1]: https://www.kaggle.com/jpuigde/outbrain-click-prediction/simple-r/discussion"
  },
  "source": "meta"
}