{
  "id": 12423,
  "title": "What tool do you guys use for this problem?",
  "url": "/competitions/malware-classification/discussion/12423",
  "author_name": "",
  "post_date": "2015-02-06T01:25:48.597Z",
  "votes": 2,
  "comment_count": 24,
  "views": 10599,
  "content": "<p>I have never processed data of such a huge size. I am wondering what tool everyone use to build model out of the data?</p>\n<p>I saw&nbsp;Microsoft Azure Machine Learning in the&nbsp;Acknowledgements of this competition. Does any use that? If so, is it good.</p>\n<p>I am familiar with R and Azure has support to run R code, but not sure how that scales to large data.</p>",
  "messages": [
    {
      "id": "63671",
      "postDate": "02/06/2015 01:25:48",
      "content": "<p>I have never processed data of such a huge size. I am wondering what tool everyone use to build model out of the data?</p>\n<p>I saw&nbsp;Microsoft Azure Machine Learning in the&nbsp;Acknowledgements of this competition. Does any use that? If so, is it good.</p>\n<p>I am familiar with R and Azure has support to run R code, but not sure how that scales to large data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63679",
      "postDate": "02/06/2015 06:49:15",
      "content": "<p>&quot;Data is not information, information is not knowledge.&quot;</p>\n<p>The huge size is not a problem, its all about how you transform that (kind of)raw data into really useful knowledge, once you do that, it will not be as big.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63689",
      "postDate": "02/06/2015 11:46:28",
      "content": "<p>NxGTR makes a good point, the first step is some data normalisation, and data reduction, to extract features that may be useful.</p>\n<p>My first attack is using FFT to transform each .bytes file's hex&nbsp;values into a frequency spectrum, so that at least the files produce a set of comparable features. &nbsp;Once I have the features, I will try some traditional learning algorithms, and feature selection, to build a predictive model. &nbsp;I have no idea whether these features will contain the information we need, until it's finished of course!</p>\n<p>If you want to try the FFT approach,&nbsp;here's an R snippet.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63691",
      "postDate": "02/06/2015 11:50:18",
      "content": "<p>after some processing my training and test dataset size is approx 50mb. Im still trying to figure out how 0.02 logloss is even possible... :D</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63692",
      "postDate": "02/06/2015 11:52:18",
      "content": "<p>Maybe one way to get really low log loss is to convert the hex back to binary and throw the files at a virus checker...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63696",
      "postDate": "02/06/2015 13:36:24",
      "content": "<p>Yes, it is possible :), and can even be improved.</p>\n<p>Did not expect it to happen so early, but this is how sports is.</p>\n<p>Good luck to all!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63697",
      "postDate": "02/06/2015 13:40:31",
      "content": "<p>[quote=jieqchen;63671]</p>\n<p>I have never processed data of such a huge size. I am wondering what tool everyone use to build model out of the data?</p>\n<p>I saw&nbsp;Microsoft Azure Machine Learning in the&nbsp;Acknowledgements of this competition. Does any use that? If so, is it good.</p>\n<p>I am familiar with R and Azure has support to run R code, but not sure how that scales to large data.</p>\n<p>[/quote]</p>\n<p>You can use R within the Azure Machine Learning platform.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63712",
      "postDate": "02/06/2015 17:56:39",
      "content": "<p>[quote=WWW BIG - Cup Committee;63696]</p>\n<p>Yes, it is possible :), and can even be improved.</p>\n<p>Did not expect it to happen so early, but this is how sports is.</p>\n<p>Good luck to all!</p>\n<p>[/quote]</p>\n<p>Now I believe you! :D</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63826",
      "postDate": "02/09/2015 20:17:22",
      "content": "<p>For this competition I use the good old unix tools. wc, grep, uniq are really your friends.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63901",
      "postDate": "02/10/2015 09:17:10",
      "content": "<p>So just to be clear - Is it allowed to use a virus checker to label the test data?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63911",
      "postDate": "02/10/2015 11:08:25",
      "content": "<p>[quote=Jay Moore;63689]</p>\n<p>NxGTR makes a good point, the first step is some data normalisation, and data reduction, to extract features that may be useful.</p>\n<p>My first attack is using FFT to transform each .bytes file's hex&nbsp;values into a frequency spectrum, so that at least the files produce a set of comparable features. &nbsp;Once I have the features, I will try some traditional learning algorithms, and feature selection, to build a predictive model. &nbsp;I have no idea whether these features will contain the information we need, until it's finished of course!</p>\n<p>If you want to try the FFT approach,&nbsp;here's an R snippet.</p>\n<p>[/quote]</p>\n<p>Hey Jay, thank you for the solution. I was wondering if we use your FFT code, we will be downloading all the data we need quickly?</p>\n<p>&nbsp;Correct me if i were wrong.&nbsp;</p>\n\n<p>thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63914",
      "postDate": "02/10/2015 12:09:59",
      "content": "<p>@Nubee To use my code you would first need to download the data, and unzip it, sorry!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63916",
      "postDate": "02/10/2015 12:19:58",
      "content": "<p>@jay: any success with FFT features yet?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63917",
      "postDate": "02/10/2015 12:38:46",
      "content": "<p>@Jay Moore - many thanks for the R code. &nbsp;I've never Kaggled before or heard of FFT so your code has given me something to look at, investigate and learn! &nbsp;Much appreciated.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63919",
      "postDate": "02/10/2015 13:42:19",
      "content": "<p>I now have an idea of how FFT works but want to understand the results produced by Jay Moore. The first row of hex in the file 0A32eTdBKayjCWhZqDOQ.bytes is:</p>\n<p>&nbsp;00401000 56 8D 44 24 08 50 8B F1 E8 1C 1B 00 00 C7 06 08</p>\n<p>and the first 3 numbers in the matrix for this file are:</p>\n<p>127849556 122725.111116317 357349.360212411</p>\n<p>What are these numbers ... are they the coefficients to the FFT formula?<br>Could you plug these numbers into an&nbsp;oscilloscope to see the signal?<br>Can/would you plot these in R?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "63932",
      "postDate": "02/10/2015 16:34:53",
      "content": "<p>@Iain, the code treats the binary values for the whole file as though it were a timeseries amplitude signal, and uses FFT to convert this into the corresponding frequency spectrum. &nbsp;</p>\n<p>It represents the spectrum as a series of complex numbers, each representing the signal in&nbsp;a particular frequency window. &nbsp;The code splits each complex number into its amplitude and phase components, and uses the first 1000 of each as features. &nbsp;</p>\n<p>The three large numbers you showed I guess are the amplitude components for the first three (lowest) frequency windows. If you plotted the first 1000 values in R it would show you the frequency spectrum of the binary values in the file, as a spectral analysis scope might.</p>\n<p>FFT is a reversible process, so if you had enough of these windows, you could in theory regenerate your original signal. &nbsp;I thought it would be an easy way of turning files of different sizes into a common basis.</p>\n<p>@Abishek, it looks better than random and gave me a LB score something around 0.9 when I chose the 9 best features, so might offer something in combination with other approaches. &nbsp;</p>\n<p>Most of the signal is in a small number of the amplitude components, but when an amplitude component is a good feature, then the corresponding phase&nbsp;component is also better than other phase&nbsp;components.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64075",
      "postDate": "02/11/2015 23:05:30",
      "content": "<p>Hi Guys,</p>\n<p>I am fairly new to this contest so don't have much knowledge about intricacies of such competition. But, I have a question, why cannot we run simple Machine learning algorithm in a Hadoop infrastructure with possibly cassandra as repository. For hadoop, 17 Gigs of data should not possess a big challenge. I have already worked on almost 3 gigs of data and ran naive bayes classification under hadoop infrastructure(single node infrastructure), it just took 2-3 mins.</p>\n<p>How about using weka API's in Hadoop/Mapreduce to approach this problem?</p>\n<p>-Vishal</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64078",
      "postDate": "02/11/2015 23:38:13",
      "content": "<p>out of curiosity: and what would you use as features into this simple algorithm? as far as i can see, none are given, so a feature extraction step is necessary anyway... Hadoop / MapReduce and all the other good stuff are at the next step of the pipeline.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64322",
      "postDate": "02/16/2015 07:37:00",
      "content": "<p>@Konrad: Thanks for pointing that out. I did not check the data set properly and fairly new to Machine learning. Thanks Again.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64330",
      "postDate": "02/16/2015 09:43:54",
      "content": "<p>[quote=jieqchen;63671]</p>\n<p>I have never processed data of such a huge size. I am wondering what tool everyone use to build model out of the data?</p>\n<p>I saw&nbsp;Microsoft Azure Machine Learning in the&nbsp;Acknowledgements of this competition. Does any use that? If so, is it good.</p>\n<p>I am familiar with R and Azure has support to run R code, but not sure how that scales to large data.</p>\n<p>[/quote]</p>\n<p>I use Azure ML studio for work and it runs fine with our few-millions row(10s of gigs) datasets...however, for use in this competition, I&nbsp;wouldnt&nbsp;really want to use AML as the integration with R packages is still being built by MS and getting in packages which are not pre-built is kinda painful vis-a-vis using RStudio on Amazon EC2. I would wait AML to come out of beta to use them for personal use... (BTW, the web API for AML makes putting my models in production remarkably easy)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64336",
      "postDate": "02/16/2015 11:39:54",
      "content": "<p>Hi There,</p>\n<p>Just wanted to check whether if anyone tried to extract features with N-gram (for byte sequences in .bytes files) or Static Instruction Sequences with Aprior Algorithm (for operand sequences from .asm files) with a Neural Network Classifier.</p>\n<p>Thanks in advance !!</p>\n<p>Rahul&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64338",
      "postDate": "02/16/2015 12:11:41",
      "content": "<p>[quote=Rahul Anand;64336]</p>\n<p>Hi There,</p>\n<p>Just wanted to check whether if anyone had a chance to tried extract features with N-gram (for byte sequences in .bytes files) or Static Instruction Sequences with Aprior Algorithm (for operand sequences from .asm files) with a Neural Network Classifier.</p>\n<p>Thanks in advance !!</p>\n<p>Rahul&nbsp;</p>\n<p>[/quote]</p>\n<p>I'm planning to use N-gram. And I think I should know more about&nbsp;assembly language, otherwise I'm not able to use .asm files...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70061",
      "postDate": "04/08/2015 23:02:50",
      "content": "<p>Take the question one step further.</p>\n<p>Not only did I want to use R either from HDInsight or Machine Learning on Azure.&nbsp; I was expecting to be to find the data set already loaded there.</p>\n<p>I was hoping to find these under the experiment templates.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70070",
      "postDate": "04/09/2015 02:00:54",
      "content": "<p>I used Python via Pypy and R. I've always used Python via Pypy (plus standard with sklearn) and R for basically every single Kaggle. Also I usually use multiple AWS instances at spot pricing. Setup is fast but you risk losing data/sub/code if you don't backup after every change given a random burst.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71258",
      "postDate": "04/11/2015 01:09:26",
      "content": "<p>I am using&nbsp;python to extract the data, and R to do pre-processing and fitting.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 63679,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "02/06/2015 06:49:15",
      "content": "<p>&quot;Data is not information, information is not knowledge.&quot;</p>\n<p>The huge size is not a problem, its all about how you transform that (kind of)raw data into really useful knowledge, once you do that, it will not be as big.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63689,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/06/2015 11:46:28",
      "content": "<p>NxGTR makes a good point, the first step is some data normalisation, and data reduction, to extract features that may be useful.</p>\n<p>My first attack is using FFT to transform each .bytes file's hex&nbsp;values into a frequency spectrum, so that at least the files produce a set of comparable features. &nbsp;Once I have the features, I will try some traditional learning algorithms, and feature selection, to build a predictive model. &nbsp;I have no idea whether these features will contain the information we need, until it's finished of course!</p>\n<p>If you want to try the FFT approach,&nbsp;here's an R snippet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63691,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/06/2015 11:50:18",
      "content": "<p>after some processing my training and test dataset size is approx 50mb. Im still trying to figure out how 0.02 logloss is even possible... :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63692,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/06/2015 11:52:18",
      "content": "<p>Maybe one way to get really low log loss is to convert the hex back to binary and throw the files at a virus checker...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63696,
      "author_name": "wwwbigcupcommittee",
      "author_url": "",
      "post_date": "02/06/2015 13:36:24",
      "content": "<p>Yes, it is possible :), and can even be improved.</p>\n<p>Did not expect it to happen so early, but this is how sports is.</p>\n<p>Good luck to all!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63697,
      "author_name": "wwwbigcupcommittee",
      "author_url": "",
      "post_date": "02/06/2015 13:40:31",
      "content": "<p>[quote=jieqchen;63671]</p>\n<p>I have never processed data of such a huge size. I am wondering what tool everyone use to build model out of the data?</p>\n<p>I saw&nbsp;Microsoft Azure Machine Learning in the&nbsp;Acknowledgements of this competition. Does any use that? If so, is it good.</p>\n<p>I am familiar with R and Azure has support to run R code, but not sure how that scales to large data.</p>\n<p>[/quote]</p>\n<p>You can use R within the Azure Machine Learning platform.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63712,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/06/2015 17:56:39",
      "content": "<p>[quote=WWW BIG - Cup Committee;63696]</p>\n<p>Yes, it is possible :), and can even be improved.</p>\n<p>Did not expect it to happen so early, but this is how sports is.</p>\n<p>Good luck to all!</p>\n<p>[/quote]</p>\n<p>Now I believe you! :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63826,
      "author_name": "leobuettiker",
      "author_url": "",
      "post_date": "02/09/2015 20:17:22",
      "content": "<p>For this competition I use the good old unix tools. wc, grep, uniq are really your friends.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63901,
      "author_name": "izikgo",
      "author_url": "",
      "post_date": "02/10/2015 09:17:10",
      "content": "<p>So just to be clear - Is it allowed to use a virus checker to label the test data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63911,
      "author_name": "nubeee",
      "author_url": "",
      "post_date": "02/10/2015 11:08:25",
      "content": "<p>[quote=Jay Moore;63689]</p>\n<p>NxGTR makes a good point, the first step is some data normalisation, and data reduction, to extract features that may be useful.</p>\n<p>My first attack is using FFT to transform each .bytes file's hex&nbsp;values into a frequency spectrum, so that at least the files produce a set of comparable features. &nbsp;Once I have the features, I will try some traditional learning algorithms, and feature selection, to build a predictive model. &nbsp;I have no idea whether these features will contain the information we need, until it's finished of course!</p>\n<p>If you want to try the FFT approach,&nbsp;here's an R snippet.</p>\n<p>[/quote]</p>\n<p>Hey Jay, thank you for the solution. I was wondering if we use your FFT code, we will be downloading all the data we need quickly?</p>\n<p>&nbsp;Correct me if i were wrong.&nbsp;</p>\n\n<p>thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63914,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/10/2015 12:09:59",
      "content": "<p>@Nubee To use my code you would first need to download the data, and unzip it, sorry!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63916,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/10/2015 12:19:58",
      "content": "<p>@jay: any success with FFT features yet?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63917,
      "author_name": "iaindobson",
      "author_url": "",
      "post_date": "02/10/2015 12:38:46",
      "content": "<p>@Jay Moore - many thanks for the R code. &nbsp;I've never Kaggled before or heard of FFT so your code has given me something to look at, investigate and learn! &nbsp;Much appreciated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63919,
      "author_name": "iaindobson",
      "author_url": "",
      "post_date": "02/10/2015 13:42:19",
      "content": "<p>I now have an idea of how FFT works but want to understand the results produced by Jay Moore. The first row of hex in the file 0A32eTdBKayjCWhZqDOQ.bytes is:</p>\n<p>&nbsp;00401000 56 8D 44 24 08 50 8B F1 E8 1C 1B 00 00 C7 06 08</p>\n<p>and the first 3 numbers in the matrix for this file are:</p>\n<p>127849556 122725.111116317 357349.360212411</p>\n<p>What are these numbers ... are they the coefficients to the FFT formula?<br>Could you plug these numbers into an&nbsp;oscilloscope to see the signal?<br>Can/would you plot these in R?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 63932,
      "author_name": "jaymoore",
      "author_url": "",
      "post_date": "02/10/2015 16:34:53",
      "content": "<p>@Iain, the code treats the binary values for the whole file as though it were a timeseries amplitude signal, and uses FFT to convert this into the corresponding frequency spectrum. &nbsp;</p>\n<p>It represents the spectrum as a series of complex numbers, each representing the signal in&nbsp;a particular frequency window. &nbsp;The code splits each complex number into its amplitude and phase components, and uses the first 1000 of each as features. &nbsp;</p>\n<p>The three large numbers you showed I guess are the amplitude components for the first three (lowest) frequency windows. If you plotted the first 1000 values in R it would show you the frequency spectrum of the binary values in the file, as a spectral analysis scope might.</p>\n<p>FFT is a reversible process, so if you had enough of these windows, you could in theory regenerate your original signal. &nbsp;I thought it would be an easy way of turning files of different sizes into a common basis.</p>\n<p>@Abishek, it looks better than random and gave me a LB score something around 0.9 when I chose the 9 best features, so might offer something in combination with other approaches. &nbsp;</p>\n<p>Most of the signal is in a small number of the amplitude components, but when an amplitude component is a good feature, then the corresponding phase&nbsp;component is also better than other phase&nbsp;components.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64075,
      "author_name": "vishalruhela",
      "author_url": "",
      "post_date": "02/11/2015 23:05:30",
      "content": "<p>Hi Guys,</p>\n<p>I am fairly new to this contest so don't have much knowledge about intricacies of such competition. But, I have a question, why cannot we run simple Machine learning algorithm in a Hadoop infrastructure with possibly cassandra as repository. For hadoop, 17 Gigs of data should not possess a big challenge. I have already worked on almost 3 gigs of data and ran naive bayes classification under hadoop infrastructure(single node infrastructure), it just took 2-3 mins.</p>\n<p>How about using weka API's in Hadoop/Mapreduce to approach this problem?</p>\n<p>-Vishal</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64078,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "02/11/2015 23:38:13",
      "content": "<p>out of curiosity: and what would you use as features into this simple algorithm? as far as i can see, none are given, so a feature extraction step is necessary anyway... Hadoop / MapReduce and all the other good stuff are at the next step of the pipeline.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64322,
      "author_name": "vishalruhela",
      "author_url": "",
      "post_date": "02/16/2015 07:37:00",
      "content": "<p>@Konrad: Thanks for pointing that out. I did not check the data set properly and fairly new to Machine learning. Thanks Again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64330,
      "author_name": "udatta",
      "author_url": "",
      "post_date": "02/16/2015 09:43:54",
      "content": "<p>[quote=jieqchen;63671]</p>\n<p>I have never processed data of such a huge size. I am wondering what tool everyone use to build model out of the data?</p>\n<p>I saw&nbsp;Microsoft Azure Machine Learning in the&nbsp;Acknowledgements of this competition. Does any use that? If so, is it good.</p>\n<p>I am familiar with R and Azure has support to run R code, but not sure how that scales to large data.</p>\n<p>[/quote]</p>\n<p>I use Azure ML studio for work and it runs fine with our few-millions row(10s of gigs) datasets...however, for use in this competition, I&nbsp;wouldnt&nbsp;really want to use AML as the integration with R packages is still being built by MS and getting in packages which are not pre-built is kinda painful vis-a-vis using RStudio on Amazon EC2. I would wait AML to come out of beta to use them for personal use... (BTW, the web API for AML makes putting my models in production remarkably easy)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64336,
      "author_name": "rahulaakkina",
      "author_url": "",
      "post_date": "02/16/2015 11:39:54",
      "content": "<p>Hi There,</p>\n<p>Just wanted to check whether if anyone tried to extract features with N-gram (for byte sequences in .bytes files) or Static Instruction Sequences with Aprior Algorithm (for operand sequences from .asm files) with a Neural Network Classifier.</p>\n<p>Thanks in advance !!</p>\n<p>Rahul&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64338,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "02/16/2015 12:11:41",
      "content": "<p>[quote=Rahul Anand;64336]</p>\n<p>Hi There,</p>\n<p>Just wanted to check whether if anyone had a chance to tried extract features with N-gram (for byte sequences in .bytes files) or Static Instruction Sequences with Aprior Algorithm (for operand sequences from .asm files) with a Neural Network Classifier.</p>\n<p>Thanks in advance !!</p>\n<p>Rahul&nbsp;</p>\n<p>[/quote]</p>\n<p>I'm planning to use N-gram. And I think I should know more about&nbsp;assembly language, otherwise I'm not able to use .asm files...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70061,
      "author_name": "tomburke",
      "author_url": "",
      "post_date": "04/08/2015 23:02:50",
      "content": "<p>Take the question one step further.</p>\n<p>Not only did I want to use R either from HDInsight or Machine Learning on Azure.&nbsp; I was expecting to be to find the data set already loaded there.</p>\n<p>I was hoping to find these under the experiment templates.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70070,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "04/09/2015 02:00:54",
      "content": "<p>I used Python via Pypy and R. I've always used Python via Pypy (plus standard with sklearn) and R for basically every single Kaggle. Also I usually use multiple AWS instances at spot pricing. Setup is fast but you risk losing data/sub/code if you don't backup after every change given a random burst.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71258,
      "author_name": "thekannman",
      "author_url": "",
      "post_date": "04/11/2015 01:09:26",
      "content": "<p>I am using&nbsp;python to extract the data, and R to do pre-processing and fitting.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "63671": "",
    "63679": "",
    "63689": "",
    "63691": "",
    "63692": "",
    "63696": "",
    "63697": "",
    "63712": "",
    "63826": "",
    "63901": "",
    "63911": "",
    "63914": "",
    "63916": "",
    "63917": "",
    "63919": "",
    "63932": "",
    "64075": "",
    "64078": "",
    "64322": "",
    "64330": "",
    "64336": "",
    "64338": "",
    "70061": "",
    "70070": "",
    "71258": ""
  },
  "source": "meta"
}