{
  "id": 12389,
  "title": "Welcome!",
  "url": "/competitions/malware-classification/discussion/12389",
  "author_name": "Will Cukierski",
  "post_date": "2015-02-03T19:11:44.237000",
  "votes": 13,
  "comment_count": 43,
  "views": 12601,
  "content": "<p>If you're the type to read user terms, you'll notice&nbsp;that ours say:</p>\n<p><em>KAGGLE DOES NOT WARRANT THAT THE WEBSITE AND RELATED SERVICES AND THE DATA PROVIDED THROUGH IT, INCLUDING THE ENTRIES AND MODELS, TO BE AVAILABLE, ACCURATE, USEFUL, OR FREE OF ERRORS, <strong>VIRUSES</strong>, OR OTHER HARMFUL COMPONENTS.</em></p>\n<p>We can finally&nbsp;make good on this promise and&nbsp;guarantee your fair share--dare we say, a lifetime's worth--of viruses from Kaggle. You may even need a new hard drive to store them all!</p>\n<p>All jokes&nbsp;aside, this competition presents a very large and interesting feature extraction step standing between you and&nbsp;a potentially difficult multi-class classification problem. Good luck and enjoy!</p>",
  "messages": [
    {
      "id": 63494,
      "postDate": "2015-02-03T19:11:44.237Z",
      "content": "<p>If you're the type to read user terms, you'll notice&nbsp;that ours say:</p>\n<p><em>KAGGLE DOES NOT WARRANT THAT THE WEBSITE AND RELATED SERVICES AND THE DATA PROVIDED THROUGH IT, INCLUDING THE ENTRIES AND MODELS, TO BE AVAILABLE, ACCURATE, USEFUL, OR FREE OF ERRORS, <strong>VIRUSES</strong>, OR OTHER HARMFUL COMPONENTS.</em></p>\n<p>We can finally&nbsp;make good on this promise and&nbsp;guarantee your fair share--dare we say, a lifetime's worth--of viruses from Kaggle. You may even need a new hard drive to store them all!</p>\n<p>All jokes&nbsp;aside, this competition presents a very large and interesting feature extraction step standing between you and&nbsp;a potentially difficult multi-class classification problem. Good luck and enjoy!</p>",
      "votes": 13
    },
    {
      "id": 63592,
      "postDate": "2015-02-04T16:33:38.633Z",
      "content": "<p>We will not be splitting the dataset and providing bytes-only or IDA-only versions. People blindly download all files available, so providing another copy of the same data just doubles our bandwidth load&nbsp;and confuses participants.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 442790,
          "postDate": "2018-12-20T14:07:52.527Z",
          "content": "<p>Hi, I need an expert help.\nI have malware .exe files, from that i am able to extract the opcodes, Addresses etc by using some tools or by using library's. but how to extract the class labels? and assign to each file. Here, in kaggle organizers have given the label data and assemble data to proceed the task. But in my case i have to prepare my own data from .exe files. Can anyone suggest how to do future engineering for exe files.</p>",
          "rawMarkdown": "Hi, I need an expert help.\nI have malware .exe files, from that i am able to extract the opcodes, Addresses etc by using some tools or by using library's. but how to extract the class labels? and assign to each file. Here, in kaggle organizers have given the label data and assemble data to proceed the task. But in my case i have to prepare my own data from .exe files. Can anyone suggest how to do future engineering for exe files."
        },
        {
          "id": 690701,
          "postDate": "2019-12-09T01:08:52.313Z",
          "content": "<p>you can generate malware label from <a href=\"https://www.virustotal.com/gui/home/upload\">virustotal</a>.\nI have same requirement to extract opcodes from my own file dataset, for academic purpose, if you still have research this problem, maybe we can discuss together.</p>",
          "rawMarkdown": "you can generate malware label from [virustotal](https://www.virustotal.com/gui/home/upload).\nI have same requirement to extract opcodes from my own file dataset, for academic purpose, if you still have research this problem, maybe we can discuss together."
        }
      ]
    },
    {
      "id": 63565,
      "postDate": "2015-02-04T13:26:41.090Z",
      "content": "<p>Also, is it possible to provide md5 of the datasets?</p>",
      "votes": 3
    },
    {
      "id": 63576,
      "postDate": "2015-02-04T13:54:25.447Z",
      "content": "<p>[quote=Ashish Mishra;63573]</p>\n<p>Data is Hex from using&nbsp;&nbsp;in IDA Pro disassembler. Do we have to parse it ourself or there is some python library for that? :-) &nbsp;</p>\n<p>[/quote]</p>\n<p>https://docs.python.org/2/library/re.html&nbsp;;)</p>",
      "votes": 4
    },
    {
      "id": 63509,
      "postDate": "2015-02-03T21:22:25.053Z",
      "content": "<p>Yes. Both the train set and test set are ~200GB each (for a total of ~400GB).</p>",
      "votes": 4
    },
    {
      "id": 64732,
      "postDate": "2015-02-22T09:27:35.643Z",
      "content": "<p>its already on data page.</p>",
      "votes": 1
    },
    {
      "id": 63879,
      "postDate": "2015-02-10T05:55:25.137Z",
      "content": "<p>Any one trying out a solution with employing HMM for extracting features from .asm files and then may be employing a back propagation network (ANN)&nbsp;</p>\n<p>http://scholarworks.sjsu.edu/cgi/viewcontent.cgi?article=1356&amp;context=etd_projects</p>",
      "votes": 1
    },
    {
      "id": 63610,
      "postDate": "2015-02-04T21:37:30.177Z",
      "content": "<p>I think we're talking about different things? A long time ago, Kaggle used to compress datasets in multiple formats so you could have zip, 7z, gz, etc., all for the same file. When we&nbsp;looked at the download logs, we'd see a ton of people&nbsp;would grab all the different flavors of the same file, even though they were grouped together in the UI and there was a warning that they were the same thing.</p>\n<p>Anyways, this is unrelated to the test set / submission ids, which is what Abhishek is asking about. It's clear if you look at the dataset or read the data description that there are two files for the same hash. If there is a mistake, let us know. I don't see one.</p>",
      "votes": 1
    },
    {
      "id": 63607,
      "postDate": "2015-02-04T20:43:28.580Z",
      "content": "<p>[quote=Abhishek;63606]</p>\n<p>We dont need predictions for all the files in the test set?</p>\n<p>[/quote]</p>\n<p>I count 21746 files in test set (two per malware).</p>\n<p>sampleSubmission has 10873 ids, 10873 * 2 = 21746</p>\n<p>What are you seeing?</p>",
      "votes": 1
    },
    {
      "id": 63569,
      "postDate": "2015-02-04T13:34:48.673Z",
      "content": "<p>The metadata(asm files) represents roughly about&nbsp;75% of the size.</p>",
      "votes": 1
    },
    {
      "id": 63567,
      "postDate": "2015-02-04T13:32:38.370Z",
      "content": "<p>Looking at the sample data most of the data is not really binary, it is annotated hex from the disassembler. &nbsp;I guess the point of the comp (from the sponsor's side) is to be able to identify patterns in this disassembled code, as they hope this is likely to expose features common among obfuscated versions of the same attack, amirite? &nbsp;</p>\n<p>In that sense, my thought is that the metadata is all buried in the .asm files, and what we need first are some decent .asm parsers to extract features from the hideous (human readable) text format of the .asm files.</p>",
      "votes": 1
    },
    {
      "id": 63562,
      "postDate": "2015-02-04T13:08:19.590Z",
      "content": "<p>As far as I understand, most of the data is binary content. What about releasing metadata-only train and test sets?</p>",
      "votes": 1
    },
    {
      "id": 63510,
      "postDate": "2015-02-03T21:23:32.530Z",
      "content": "<p>Thanks.</p>",
      "votes": 1
    },
    {
      "id": 64491,
      "postDate": "2015-02-18T13:31:18.190Z",
      "content": "<p>eric@glamdring:~/workspace/microsoft-malware-kaggle/mnt$ ls train/*.bytes | wc -l<br>10868<br>eric@glamdring:~/workspace/microsoft-malware-kaggle/mnt$ ls train/*.asm | wc -l<br>10868<br>eric@glamdring:~/workspace/microsoft-malware-kaggle/mnt$&nbsp;echo &quot;10868 * 2&quot; | bc<br>21736</p>\n<p>It looks like you're good to go.</p>",
      "votes": 2
    },
    {
      "id": 63571,
      "postDate": "2015-02-04T13:40:08.643Z",
      "content": "<p>[quote=Abhishek;63565]</p>\n<p>Also, is it possible to provide md5 of the datasets?</p>\n<p>[/quote]</p>\n<p>Will post these on the data page.</p>",
      "votes": 2
    },
    {
      "id": 64731,
      "postDate": "2015-02-22T09:10:54.507Z",
      "content": "<p>[quote=William Cukierski;63571]</p>\n<p>[quote=Abhishek;63565]</p>\n<p>Also, is it possible to provide md5 of the datasets?</p>\n<p>[/quote]</p>\n<p>Will post these on the data page.</p>\n<p>[/quote]</p>\n<p>Hi,</p>\n<p>It can be very useful to have those MD5. When do you plan to post them?</p>\n<p><br>Thanks</p>"
    },
    {
      "id": 1875858,
      "postDate": "2022-07-29T11:04:58.303Z",
      "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> <a href=\"https://www.kaggle.com/marianr\" target=\"_blank\">@marianr</a> Sorry to be reviving such an old thread, but could you perhaps let me know if the dataset is shareable (and under which license) outside of Kaggle now that the competition is over? I would like to add it to Hugging Face datasets (<a href=\"https://huggingface.co/datasets\" target=\"_blank\">https://huggingface.co/datasets</a>) which would also support gated access subject to agreeing to additional terms and being logged in with an email address.</p>",
      "rawMarkdown": "@wcukierski @marianr Sorry to be reviving such an old thread, but could you perhaps let me know if the dataset is shareable (and under which license) outside of Kaggle now that the competition is over? I would like to add it to Hugging Face datasets (https://huggingface.co/datasets) which would also support gated access subject to agreeing to additional terms and being logged in with an email address.",
      "replies": [
        {
          "id": 1876491,
          "postDate": "2022-07-29T19:50:30.200Z",
          "content": "<p><a href=\"https://www.kaggle.com/cakiki\" target=\"_blank\">@cakiki</a> Kaggle staff unfortunately have no special ownership in the datasets we use for competitions. I'd just have to refer you to the rules <a href=\"https://www.kaggle.com/competitions/malware-classification/rules\" target=\"_blank\">https://www.kaggle.com/competitions/malware-classification/rules</a>. Apologies I can't be of help!</p>",
          "rawMarkdown": "@cakiki Kaggle staff unfortunately have no special ownership in the datasets we use for competitions. I'd just have to refer you to the rules https://www.kaggle.com/competitions/malware-classification/rules. Apologies I can't be of help!"
        }
      ]
    },
    {
      "id": 783369,
      "postDate": "2020-03-23T08:56:01.663Z",
      "content": "<p>Hello Guys! \nCan anyone post a mini dataset of around 2 GB zip as its not feasible for students like me to have data connection to download 50GB dataset. Thanks in advance </p>",
      "rawMarkdown": "Hello Guys! \nCan anyone post a mini dataset of around 2 GB zip as its not feasible for students like me to have data connection to download 50GB dataset. Thanks in advance "
    },
    {
      "id": 367203,
      "postDate": "2018-08-07T09:51:03.323Z",
      "content": "<p>how to download data?</p>",
      "rawMarkdown": "how to download data?"
    },
    {
      "id": 194580,
      "postDate": "2017-06-21T04:58:23.677Z",
      "content": "<p>HI, \n    Why is the file I downloaded is empty? Is the tag data not showing? Can you provide feature dataset for me and label about test data to validate my approach? Thank you very much!</p>",
      "rawMarkdown": "HI, \n    Why is the file I downloaded is empty? Is the tag data not showing? Can you provide feature dataset for me and label about test data to validate my approach? Thank you very much!"
    },
    {
      "id": 96863,
      "postDate": "2015-10-20T20:34:57.650Z",
      "content": "<p>Hi Arun,\nWe are also unable to send you a message.\nPlease send your contact info to: malwarecompetition@outlook.com</p>",
      "rawMarkdown": "Hi Arun,\r\nWe are also unable to send you a message.\r\nPlease send your contact info to: malwarecompetition@outlook.com",
      "replies": [
        {
          "id": 305030,
          "postDate": "2018-03-28T11:39:17.610Z",
          "content": "<p>Hi, I'm trying to do a research project and would need the hashes of the malware samples in the dataset, is it possible for you to send it to me? I've already sent 2 emails to the address malwarecompetition@outlook.com but still have not received any replies.</p>",
          "rawMarkdown": "Hi, I'm trying to do a research project and would need the hashes of the malware samples in the dataset, is it possible for you to send it to me? I've already sent 2 emails to the address malwarecompetition@outlook.com but still have not received any replies."
        }
      ]
    },
    {
      "id": 96637,
      "postDate": "2015-10-19T10:48:32.817Z",
      "content": "<p>Admin -- Thanks for the reply. Kaggle tells me I don't have enough points to send a private message.   Could you send me a message instead?</p>",
      "rawMarkdown": "Admin -- Thanks for the reply. Kaggle tells me I don't have enough points to send a private message.   Could you send me a message instead?",
      "replies": [
        {
          "id": 1216633,
          "postDate": "2021-02-24T11:41:18.957Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 96632,
      "postDate": "2015-10-19T09:11:19.963Z",
      "content": "<p>Dear Arun,</p>\n\n<p>We did not upload the md5 to the public site.</p>\n\n<p>Please contact us through a private message including your mail. </p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Dear Arun,\r\n\r\nWe did not upload the md5 to the public site.\r\n\r\nPlease contact us through a private message including your mail. \r\n\r\nThanks!"
    },
    {
      "id": 96579,
      "postDate": "2015-10-18T19:07:53.917Z",
      "content": "<p>I learned about this competition while reviewing a paper that used this data. The paper also mentioned availability of md5. I see a promise of uploading md5 in the discussion. But I don't see the data on the website.</p>\n\n<p>Are the md5s available at some other location?</p>",
      "rawMarkdown": "I learned about this competition while reviewing a paper that used this data. The paper also mentioned availability of md5. I see a promise of uploading md5 in the discussion. But I don't see the data on the website.\r\n\r\nAre the md5s available at some other location?\r\n\r\n"
    },
    {
      "id": 64747,
      "postDate": "2015-02-22T22:32:16.330Z",
      "content": "<p>Still nothing on this???:&nbsp;<br>http://www.kaggle.com/c/malware-classification/forums/t/12537/organizers-validity-of-train-labels</p>"
    },
    {
      "id": 64733,
      "postDate": "2015-02-22T12:14:18.100Z",
      "content": "<p>[quote=Abhishek;64732]</p>\n<p>its already on data page.</p>\n<p>[/quote]</p>\n<p>I thought about the MD5 hash of each sample, not on the entire set&nbsp;:)</p>\n<p>Do you think it is possible to have them?</p>"
    },
    {
      "id": 64665,
      "postDate": "2015-02-20T20:04:33.157Z",
      "content": "<p>[quote=Marian Radu;63589]</p>\n<p>Yes, all are x86. You can disassemble the code directly&nbsp;from the bytes files but it is not straight forward(PE header is missing), nor will you get the same level of detail.</p>\n<p>[/quote]</p>\n<p>Given that disassembling code from&nbsp;the bytes files is hard without the corresponding headers, can you make the corresponding IDA generated .idb files available, so one&nbsp;can employ IDA to extract a much richer set&nbsp;of features to work with? &nbsp;I can understand that distributing the raw binaries is unsafe due to their malicious nature, but having to work with&nbsp;lossy,&nbsp;ASCII representations of those binaries&nbsp;seems like an unnecessary, artificial constraint. Making the corresponding .idb files would greatly alleviate that constraint.</p>"
    },
    {
      "id": 64627,
      "postDate": "2015-02-20T04:30:28.443Z",
      "content": "<p>Hello,</p>\n<p>Many files have GAP(??) in both bytes and asm. Is this deliberate? Do others have this too?</p>"
    },
    {
      "id": 64470,
      "postDate": "2015-02-18T08:17:12.307Z",
      "content": "<p>[quote=William Cukierski;63607]</p>\n<p>[quote=Abhishek;63606]</p>\n<p>We dont need predictions for all the files in the test set?</p>\n<p>[/quote]</p>\n<p>I count 21746 files in test set (two per malware).</p>\n<p>sampleSubmission has 10873 ids, 10873 * 2 = 21746</p>\n<p>What are you seeing?</p>\n<p>[/quote]</p>\n\n<p>There are 21736 files in the training set , right? Just want to confirm that the extraction is fully successful. Thanks.</p>"
    },
    {
      "id": 63617,
      "postDate": "2015-02-04T22:27:30.270Z",
      "content": "<p>[quote=William Cukierski;63610]</p>\n<p>I think we're talking about different things?</p>\n<p>[/quote]</p>\n<p>Yep. It all makes sense now. Thx.</p>"
    },
    {
      "id": 63615,
      "postDate": "2015-02-04T22:17:14.417Z",
      "content": "<p>as usual, the mistake is on my part. Extraction of files got terminated and i was confused why sample submission has less rows....pff!&nbsp;</p>\n<p>Thanks!</p>"
    },
    {
      "id": 63609,
      "postDate": "2015-02-04T21:27:44.333Z",
      "content": "<p>[quote=William Cukierski;63607]</p>\n<p>I count 21746 files in test set (two per malware).</p>\n<p>sampleSubmission has 10873 ids, 10873 * 2 = 21746</p>\n<p>What are you seeing?</p>\n<p>[/quote]</p>\n<p>Perhaps it was your comment &quot;people blindly downloading all files available&quot; that may have caused some head scratching. (Myself included.)</p>"
    },
    {
      "id": 63606,
      "postDate": "2015-02-04T20:34:01.197Z",
      "content": "<p>We dont need predictions for all the files in the test set?</p>"
    },
    {
      "id": 63590,
      "postDate": "2015-02-04T15:47:30.933Z",
      "content": "<p>[quote=Marian Radu;63589]</p>\n<p>Yes, all are x86. You can disassemble the code directly&nbsp;from the bytes files but it is not straight forward(PE header is missing), nor will you get the same level of detail.</p>\n<p>[/quote]</p>\n<p>Thanks, It will be helpful if you seperately upload bytes and IDA disassembled file as IDA file 75%of the whole data.&nbsp;</p>"
    },
    {
      "id": 63589,
      "postDate": "2015-02-04T15:09:42.180Z",
      "content": "<p>Yes, all are x86. You can disassemble the code directly&nbsp;from the bytes files but it is not straight forward(PE header is missing), nor will you get the same level of detail.</p>"
    },
    {
      "id": 63580,
      "postDate": "2015-02-04T14:01:29.927Z",
      "content": "<p>2 more questions:</p>\n<p>1. I dont think its mentioned but All are x86 PE file?</p>\n<p>2. Can it be disassembled from packed malware binary using one or many packers?</p>"
    },
    {
      "id": 63573,
      "postDate": "2015-02-04T13:44:50.417Z",
      "content": "<p>Data is Hex from using&nbsp;&nbsp;in IDA Pro disassembler. Do we have to parse it ourself or there is some python library for that? :-) &nbsp;</p>"
    },
    {
      "id": 63508,
      "postDate": "2015-02-03T21:14:02.747Z",
      "content": "<p>Before I even download the data, is this half a terabyte total between test and train?</p>"
    },
    {
      "id": 442788,
      "postDate": "2018-12-20T14:05:25.643Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 63592,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2015-02-04T16:33:38.633000",
      "content": "<p>We will not be splitting the dataset and providing bytes-only or IDA-only versions. People blindly download all files available, so providing another copy of the same data just doubles our bandwidth load&nbsp;and confuses participants.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 442790,
          "author_name": "vikram Reddy",
          "author_url": "",
          "post_date": "2018-12-20T14:07:52.527000",
          "content": "<p>Hi, I need an expert help.\nI have malware .exe files, from that i am able to extract the opcodes, Addresses etc by using some tools or by using library's. but how to extract the class labels? and assign to each file. Here, in kaggle organizers have given the label data and assemble data to proceed the task. But in my case i have to prepare my own data from .exe files. Can anyone suggest how to do future engineering for exe files.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 690701,
          "author_name": "MAO-HSUN",
          "author_url": "",
          "post_date": "2019-12-09T01:08:52.313000",
          "content": "<p>you can generate malware label from <a href=\"https://www.virustotal.com/gui/home/upload\">virustotal</a>.\nI have same requirement to extract opcodes from my own file dataset, for academic purpose, if you still have research this problem, maybe we can discuss together.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 63565,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2015-02-04T13:26:41.090000",
      "content": "<p>Also, is it possible to provide md5 of the datasets?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 63576,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2015-02-04T13:54:25.447000",
      "content": "<p>[quote=Ashish Mishra;63573]</p>\n<p>Data is Hex from using&nbsp;&nbsp;in IDA Pro disassembler. Do we have to parse it ourself or there is some python library for that? :-) &nbsp;</p>\n<p>[/quote]</p>\n<p>https://docs.python.org/2/library/re.html&nbsp;;)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 63509,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2015-02-03T21:22:25.053000",
      "content": "<p>Yes. Both the train set and test set are ~200GB each (for a total of ~400GB).</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 64732,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2015-02-22T09:27:35.643000",
      "content": "<p>its already on data page.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 63879,
      "author_name": "Rahul Anand",
      "author_url": "",
      "post_date": "2015-02-10T05:55:25.137000",
      "content": "<p>Any one trying out a solution with employing HMM for extracting features from .asm files and then may be employing a back propagation network (ANN)&nbsp;</p>\n<p>http://scholarworks.sjsu.edu/cgi/viewcontent.cgi?article=1356&amp;context=etd_projects</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 63610,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2015-02-04T21:37:30.177000",
      "content": "<p>I think we're talking about different things? A long time ago, Kaggle used to compress datasets in multiple formats so you could have zip, 7z, gz, etc., all for the same file. When we&nbsp;looked at the download logs, we'd see a ton of people&nbsp;would grab all the different flavors of the same file, even though they were grouped together in the UI and there was a warning that they were the same thing.</p>\n<p>Anyways, this is unrelated to the test set / submission ids, which is what Abhishek is asking about. It's clear if you look at the dataset or read the data description that there are two files for the same hash. If there is a mistake, let us know. I don't see one.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 63607,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2015-02-04T20:43:28.580000",
      "content": "<p>[quote=Abhishek;63606]</p>\n<p>We dont need predictions for all the files in the test set?</p>\n<p>[/quote]</p>\n<p>I count 21746 files in test set (two per malware).</p>\n<p>sampleSubmission has 10873 ids, 10873 * 2 = 21746</p>\n<p>What are you seeing?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 63569,
      "author_name": "Marian",
      "author_url": "",
      "post_date": "2015-02-04T13:34:48.673000",
      "content": "<p>The metadata(asm files) represents roughly about&nbsp;75% of the size.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 63567,
      "author_name": "Jay Moore",
      "author_url": "",
      "post_date": "2015-02-04T13:32:38.370000",
      "content": "<p>Looking at the sample data most of the data is not really binary, it is annotated hex from the disassembler. &nbsp;I guess the point of the comp (from the sponsor's side) is to be able to identify patterns in this disassembled code, as they hope this is likely to expose features common among obfuscated versions of the same attack, amirite? &nbsp;</p>\n<p>In that sense, my thought is that the metadata is all buried in the .asm files, and what we need first are some decent .asm parsers to extract features from the hideous (human readable) text format of the .asm files.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 63562,
      "author_name": "Foxtrot",
      "author_url": "",
      "post_date": "2015-02-04T13:08:19.590000",
      "content": "<p>As far as I understand, most of the data is binary content. What about releasing metadata-only train and test sets?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 63510,
      "author_name": "Giulio",
      "author_url": "",
      "post_date": "2015-02-03T21:23:32.530000",
      "content": "<p>Thanks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 64491,
      "author_name": "Eric Whyne",
      "author_url": "",
      "post_date": "2015-02-18T13:31:18.190000",
      "content": "<p>eric@glamdring:~/workspace/microsoft-malware-kaggle/mnt$ ls train/*.bytes | wc -l<br>10868<br>eric@glamdring:~/workspace/microsoft-malware-kaggle/mnt$ ls train/*.asm | wc -l<br>10868<br>eric@glamdring:~/workspace/microsoft-malware-kaggle/mnt$&nbsp;echo &quot;10868 * 2&quot; | bc<br>21736</p>\n<p>It looks like you're good to go.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 63571,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2015-02-04T13:40:08.643000",
      "content": "<p>[quote=Abhishek;63565]</p>\n<p>Also, is it possible to provide md5 of the datasets?</p>\n<p>[/quote]</p>\n<p>Will post these on the data page.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 64731,
      "author_name": "Dan Journo",
      "author_url": "",
      "post_date": "2015-02-22T09:10:54.507000",
      "content": "<p>[quote=William Cukierski;63571]</p>\n<p>[quote=Abhishek;63565]</p>\n<p>Also, is it possible to provide md5 of the datasets?</p>\n<p>[/quote]</p>\n<p>Will post these on the data page.</p>\n<p>[/quote]</p>\n<p>Hi,</p>\n<p>It can be very useful to have those MD5. When do you plan to post them?</p>\n<p><br>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1875858,
      "author_name": "Chris Akiki",
      "author_url": "",
      "post_date": "2022-07-29T11:04:58.303000",
      "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> <a href=\"https://www.kaggle.com/marianr\" target=\"_blank\">@marianr</a> Sorry to be reviving such an old thread, but could you perhaps let me know if the dataset is shareable (and under which license) outside of Kaggle now that the competition is over? I would like to add it to Hugging Face datasets (<a href=\"https://huggingface.co/datasets\" target=\"_blank\">https://huggingface.co/datasets</a>) which would also support gated access subject to agreeing to additional terms and being logged in with an email address.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1876491,
          "author_name": "Will Cukierski",
          "author_url": "",
          "post_date": "2022-07-29T19:50:30.200000",
          "content": "<p><a href=\"https://www.kaggle.com/cakiki\" target=\"_blank\">@cakiki</a> Kaggle staff unfortunately have no special ownership in the datasets we use for competitions. I'd just have to refer you to the rules <a href=\"https://www.kaggle.com/competitions/malware-classification/rules\" target=\"_blank\">https://www.kaggle.com/competitions/malware-classification/rules</a>. Apologies I can't be of help!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 783369,
      "author_name": "B Raja Narasimhan",
      "author_url": "",
      "post_date": "2020-03-23T08:56:01.663000",
      "content": "<p>Hello Guys! \nCan anyone post a mini dataset of around 2 GB zip as its not feasible for students like me to have data connection to download 50GB dataset. Thanks in advance </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 367203,
      "author_name": "xiefangjun",
      "author_url": "",
      "post_date": "2018-08-07T09:51:03.323000",
      "content": "<p>how to download data?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 194580,
      "author_name": "vinus",
      "author_url": "",
      "post_date": "2017-06-21T04:58:23.677000",
      "content": "<p>HI, \n    Why is the file I downloaded is empty? Is the tag data not showing? Can you provide feature dataset for me and label about test data to validate my approach? Thank you very much!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 96863,
      "author_name": "WWW BIG - Cup Committee",
      "author_url": "",
      "post_date": "2015-10-20T20:34:57.650000",
      "content": "<p>Hi Arun,\nWe are also unable to send you a message.\nPlease send your contact info to: malwarecompetition@outlook.com</p>",
      "votes": 0,
      "replies": [
        {
          "id": 305030,
          "author_name": "Ricardo Ramos",
          "author_url": "",
          "post_date": "2018-03-28T11:39:17.610000",
          "content": "<p>Hi, I'm trying to do a research project and would need the hashes of the malware samples in the dataset, is it possible for you to send it to me? I've already sent 2 emails to the address malwarecompetition@outlook.com but still have not received any replies.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 96637,
      "author_name": "Arun L",
      "author_url": "",
      "post_date": "2015-10-19T10:48:32.817000",
      "content": "<p>Admin -- Thanks for the reply. Kaggle tells me I don't have enough points to send a private message.   Could you send me a message instead?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1216633,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-24T11:41:18.957000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 96632,
      "author_name": "WWW BIG - Cup Committee",
      "author_url": "",
      "post_date": "2015-10-19T09:11:19.963000",
      "content": "<p>Dear Arun,</p>\n\n<p>We did not upload the md5 to the public site.</p>\n\n<p>Please contact us through a private message including your mail. </p>\n\n<p>Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 96579,
      "author_name": "Arun L",
      "author_url": "",
      "post_date": "2015-10-18T19:07:53.917000",
      "content": "<p>I learned about this competition while reviewing a paper that used this data. The paper also mentioned availability of md5. I see a promise of uploading md5 in the discussion. But I don't see the data on the website.</p>\n\n<p>Are the md5s available at some other location?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 64747,
      "author_name": "NxGTR",
      "author_url": "",
      "post_date": "2015-02-22T22:32:16.330000",
      "content": "<p>Still nothing on this???:&nbsp;<br>http://www.kaggle.com/c/malware-classification/forums/t/12537/organizers-validity-of-train-labels</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 64733,
      "author_name": "Dan Journo",
      "author_url": "",
      "post_date": "2015-02-22T12:14:18.100000",
      "content": "<p>[quote=Abhishek;64732]</p>\n<p>its already on data page.</p>\n<p>[/quote]</p>\n<p>I thought about the MD5 hash of each sample, not on the entire set&nbsp;:)</p>\n<p>Do you think it is possible to have them?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 64665,
      "author_name": "defiant",
      "author_url": "",
      "post_date": "2015-02-20T20:04:33.157000",
      "content": "<p>[quote=Marian Radu;63589]</p>\n<p>Yes, all are x86. You can disassemble the code directly&nbsp;from the bytes files but it is not straight forward(PE header is missing), nor will you get the same level of detail.</p>\n<p>[/quote]</p>\n<p>Given that disassembling code from&nbsp;the bytes files is hard without the corresponding headers, can you make the corresponding IDA generated .idb files available, so one&nbsp;can employ IDA to extract a much richer set&nbsp;of features to work with? &nbsp;I can understand that distributing the raw binaries is unsafe due to their malicious nature, but having to work with&nbsp;lossy,&nbsp;ASCII representations of those binaries&nbsp;seems like an unnecessary, artificial constraint. Making the corresponding .idb files would greatly alleviate that constraint.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 64627,
      "author_name": "Lakshman Nataraj",
      "author_url": "",
      "post_date": "2015-02-20T04:30:28.443000",
      "content": "<p>Hello,</p>\n<p>Many files have GAP(??) in both bytes and asm. Is this deliberate? Do others have this too?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 64470,
      "author_name": "UD1989",
      "author_url": "",
      "post_date": "2015-02-18T08:17:12.307000",
      "content": "<p>[quote=William Cukierski;63607]</p>\n<p>[quote=Abhishek;63606]</p>\n<p>We dont need predictions for all the files in the test set?</p>\n<p>[/quote]</p>\n<p>I count 21746 files in test set (two per malware).</p>\n<p>sampleSubmission has 10873 ids, 10873 * 2 = 21746</p>\n<p>What are you seeing?</p>\n<p>[/quote]</p>\n\n<p>There are 21736 files in the training set , right? Just want to confirm that the extraction is fully successful. Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63617,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2015-02-04T22:27:30.270000",
      "content": "<p>[quote=William Cukierski;63610]</p>\n<p>I think we're talking about different things?</p>\n<p>[/quote]</p>\n<p>Yep. It all makes sense now. Thx.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63615,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-02-04T22:17:14.417000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63609,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-02-04T21:27:44.333000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63606,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-02-04T20:34:01.197000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63590,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-02-04T15:47:30.933000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63589,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-02-04T15:09:42.180000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63580,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-02-04T14:01:29.927000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63573,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-02-04T13:44:50.417000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 63508,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-02-03T21:14:02.747000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 442788,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-20T14:05:25.643000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "63494": "",
    "63592": "",
    "63565": "",
    "63576": "",
    "63509": "",
    "64732": "",
    "63879": "",
    "63610": "",
    "63607": "",
    "63569": "",
    "63567": "",
    "63562": "",
    "63510": "",
    "64491": "",
    "63571": "",
    "64731": "",
    "1875858": "@wcukierski @marianr Sorry to be reviving such an old thread, but could you perhaps let me know if the dataset is shareable (and under which license) outside of Kaggle now that the competition is over? I would like to add it to Hugging Face datasets (https://huggingface.co/datasets) which would also support gated access subject to agreeing to additional terms and being logged in with an email address.",
    "783369": "Hello Guys! \nCan anyone post a mini dataset of around 2 GB zip as its not feasible for students like me to have data connection to download 50GB dataset. Thanks in advance ",
    "367203": "how to download data?",
    "194580": "HI, \n    Why is the file I downloaded is empty? Is the tag data not showing? Can you provide feature dataset for me and label about test data to validate my approach? Thank you very much!",
    "96863": "Hi Arun,\r\nWe are also unable to send you a message.\r\nPlease send your contact info to: malwarecompetition@outlook.com",
    "96637": "Admin -- Thanks for the reply. Kaggle tells me I don't have enough points to send a private message.   Could you send me a message instead?",
    "96632": "Dear Arun,\r\n\r\nWe did not upload the md5 to the public site.\r\n\r\nPlease contact us through a private message including your mail. \r\n\r\nThanks!",
    "96579": "I learned about this competition while reviewing a paper that used this data. The paper also mentioned availability of md5. I see a promise of uploading md5 in the discussion. But I don't see the data on the website.\r\n\r\nAre the md5s available at some other location?\r\n\r\n",
    "64747": "",
    "64733": "",
    "64665": "",
    "64627": "",
    "64470": "",
    "63617": "",
    "63615": "",
    "63609": "",
    "63606": "",
    "63590": "",
    "63589": "",
    "63580": "",
    "63573": "",
    "63508": "",
    "442788": ""
  }
}