{
  "id": 53792,
  "title": "R multicore usage with multidplyr",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53792",
  "author_name": "",
  "post_date": "2018-04-05T08:28:06.129566600Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Just came across the great package multidplyr (<a href=\"https://github.com/hadley/multidplyr\">https://github.com/hadley/multidplyr</a> ), which allows to distribute groupings over several cores and perform calculations like counts, summaries etc in parallel using dplyr syntax. Comes in very handy for this data set. For me using 12 cores it speeds up computations around 8x ! </p>",
  "messages": [
    {
      "id": "309400",
      "postDate": "04/05/2018 08:28:06",
      "content": "<p>Just came across the great package multidplyr (<a href=\"https://github.com/hadley/multidplyr\">https://github.com/hadley/multidplyr</a> ), which allows to distribute groupings over several cores and perform calculations like counts, summaries etc in parallel using dplyr syntax. Comes in very handy for this data set. For me using 12 cores it speeds up computations around 8x ! </p>",
      "rawMarkdown": "Just came across the great package multidplyr (https://github.com/hadley/multidplyr ), which allows to distribute groupings over several cores and perform calculations like counts, summaries etc in parallel using dplyr syntax. Comes in very handy for this data set. For me using 12 cores it speeds up computations around 8x !",
      "votes": null
    },
    {
      "id": "309411",
      "postDate": "04/05/2018 09:14:43",
      "content": "<p>Thank you for sharing @Malte. </p>",
      "rawMarkdown": "Thank you for sharing @Malte.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 309411,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "04/05/2018 09:14:43",
      "content": "<p>Thank you for sharing @Malte. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "309400": "Just came across the great package multidplyr (https://github.com/hadley/multidplyr ), which allows to distribute groupings over several cores and perform calculations like counts, summaries etc in parallel using dplyr syntax. Comes in very handy for this data set. For me using 12 cores it speeds up computations around 8x !",
    "309411": "Thank you for sharing @Malte."
  },
  "source": "meta"
}