{
  "id": 56280,
  "title": "Technology stack",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56280",
  "author_name": "",
  "post_date": "2018-05-08T08:20:08.779972600Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>It seems in this competition the choice of technologies, both hardware and software, was probably as important as the ML techniques. Hence, I open this thread so that everybody can explain what was different in their technology stack that allowed them to iterate faster or use more data or whatever other competitive advantage they may get from it.</p>\n\n<p>On my side, very naive, I just went with R using multidplyr (that is not released as official version, as has some issues that sure will be fixed for the final version), and used Amazon m10xlarge (40 CPUS and 160 GB).  </p>\n\n<p>As per other threads I have read, <a href=\"/tkm2261\">@tkm2261</a>: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56250\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56250</a> has used Google Cloud Platform with Big Query.</p>\n\n<p>Please add to this thread regarding this part of your solution.</p>",
  "messages": [
    {
      "id": "325209",
      "postDate": "05/08/2018 08:20:08",
      "content": "<p>It seems in this competition the choice of technologies, both hardware and software, was probably as important as the ML techniques. Hence, I open this thread so that everybody can explain what was different in their technology stack that allowed them to iterate faster or use more data or whatever other competitive advantage they may get from it.</p>\n\n<p>On my side, very naive, I just went with R using multidplyr (that is not released as official version, as has some issues that sure will be fixed for the final version), and used Amazon m10xlarge (40 CPUS and 160 GB).  </p>\n\n<p>As per other threads I have read, <a href=\"/tkm2261\">@tkm2261</a>: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56250\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56250</a> has used Google Cloud Platform with Big Query.</p>\n\n<p>Please add to this thread regarding this part of your solution.</p>",
      "rawMarkdown": "It seems in this competition the choice of technologies, both hardware and software, was probably as important as the ML techniques. Hence, I open this thread so that everybody can explain what was different in their technology stack that allowed them to iterate faster or use more data or whatever other competitive advantage they may get from it.\n\nOn my side, very naive, I just went with R using multidplyr (that is not released as official version, as has some issues that sure will be fixed for the final version), and used Amazon m10xlarge (40 CPUS and 160 GB).  \n\nAs per other threads I have read, @tkm2261: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56250 has used Google Cloud Platform with Big Query.\n\nPlease add to this thread regarding this part of your solution.",
      "votes": null
    },
    {
      "id": "325220",
      "postDate": "05/08/2018 08:33:38",
      "content": "<p>It is my first competition and my first attempt at machine learning. </p>\n\n<p>I've managed to create an NN that generated an LB of ~0.98 with less than 200k params.</p>\n\n<p>I trained it on an Intel quad-core i7 4400k, 32 GB RAM DDR3, nVidia 1070 (gigabyte g1). An epoch would take me about 1h to run and needed 6 epochs to get to that score. Obviously I've iterated over multiple NN model versions. I've only used features available at the time of each click - meaning I haven't relayed on aggregations of all available data but only of data available up-to the moment in time the row was generated. </p>\n\n<p>In the following days I would generate a Dataset with all the fancy features people used to get over 0.98 result and I will check if my idea of NN would be able to train as fast as previous iterations producing similar results as the top ones relaying on LightGBM. I've never ran into memory issues, my GPU used only about 3 GB for my 300k batch, I consider it to train pretty fast to an okish AUC.</p>\n\n<p>Forgot to mention :-) \npython, pytorch, pandas, sklearn (for AUC)</p>",
      "rawMarkdown": "It is my first competition and my first attempt at machine learning. \n\nI've managed to create an NN that generated an LB of ~0.98 with less than 200k params.\n\nI trained it on an Intel quad-core i7 4400k, 32 GB RAM DDR3, nVidia 1070 (gigabyte g1). An epoch would take me about 1h to run and needed 6 epochs to get to that score. Obviously I've iterated over multiple NN model versions. I've only used features available at the time of each click - meaning I haven't relayed on aggregations of all available data but only of data available up-to the moment in time the row was generated. \n\nIn the following days I would generate a Dataset with all the fancy features people used to get over 0.98 result and I will check if my idea of NN would be able to train as fast as previous iterations producing similar results as the top ones relaying on LightGBM. I've never ran into memory issues, my GPU used only about 3 GB for my 300k batch, I consider it to train pretty fast to an okish AUC.\n\nForgot to mention :-) \npython, pytorch, pandas, sklearn (for AUC)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325220,
      "author_name": "profetul",
      "author_url": "",
      "post_date": "05/08/2018 08:33:38",
      "content": "<p>It is my first competition and my first attempt at machine learning. </p>\n\n<p>I've managed to create an NN that generated an LB of ~0.98 with less than 200k params.</p>\n\n<p>I trained it on an Intel quad-core i7 4400k, 32 GB RAM DDR3, nVidia 1070 (gigabyte g1). An epoch would take me about 1h to run and needed 6 epochs to get to that score. Obviously I've iterated over multiple NN model versions. I've only used features available at the time of each click - meaning I haven't relayed on aggregations of all available data but only of data available up-to the moment in time the row was generated. </p>\n\n<p>In the following days I would generate a Dataset with all the fancy features people used to get over 0.98 result and I will check if my idea of NN would be able to train as fast as previous iterations producing similar results as the top ones relaying on LightGBM. I've never ran into memory issues, my GPU used only about 3 GB for my 300k batch, I consider it to train pretty fast to an okish AUC.</p>\n\n<p>Forgot to mention :-) \npython, pytorch, pandas, sklearn (for AUC)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325209": "It seems in this competition the choice of technologies, both hardware and software, was probably as important as the ML techniques. Hence, I open this thread so that everybody can explain what was different in their technology stack that allowed them to iterate faster or use more data or whatever other competitive advantage they may get from it.\n\nOn my side, very naive, I just went with R using multidplyr (that is not released as official version, as has some issues that sure will be fixed for the final version), and used Amazon m10xlarge (40 CPUS and 160 GB).  \n\nAs per other threads I have read, @tkm2261: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56250 has used Google Cloud Platform with Big Query.\n\nPlease add to this thread regarding this part of your solution.",
    "325220": "It is my first competition and my first attempt at machine learning. \n\nI've managed to create an NN that generated an LB of ~0.98 with less than 200k params.\n\nI trained it on an Intel quad-core i7 4400k, 32 GB RAM DDR3, nVidia 1070 (gigabyte g1). An epoch would take me about 1h to run and needed 6 epochs to get to that score. Obviously I've iterated over multiple NN model versions. I've only used features available at the time of each click - meaning I haven't relayed on aggregations of all available data but only of data available up-to the moment in time the row was generated. \n\nIn the following days I would generate a Dataset with all the fancy features people used to get over 0.98 result and I will check if my idea of NN would be able to train as fast as previous iterations producing similar results as the top ones relaying on LightGBM. I've never ran into memory issues, my GPU used only about 3 GB for my 300k batch, I consider it to train pretty fast to an okish AUC.\n\nForgot to mention :-) \npython, pytorch, pandas, sklearn (for AUC)"
  },
  "source": "meta"
}