{
  "id": 12646,
  "title": "Fresher Requirement understanding not clear",
  "url": "/competitions/malware-classification/discussion/12646",
  "author_name": "",
  "post_date": "2015-03-01T12:44:44.593Z",
  "votes": null,
  "comment_count": 1,
  "views": 1074,
  "content": "<p>I am fresher to this Data Science altogether and when exploring this topic, I just got a doubt here (It may be too dumb).&nbsp;</p>\n\n<p>I understand that the given test data would be analyzed and using a logistic regression or any other algo, an equation is formed to classify the files.</p>\n<p>Once the equation is formed to calculate the probability we can test it using the test data. I am missing out at one point here, in the test data an altogether new file comes to get processed, how can we relate it. I mean how can we get the relationship between two malware files.</p>",
  "messages": [
    {
      "id": "65176",
      "postDate": "03/01/2015 12:44:44",
      "content": "<p>I am fresher to this Data Science altogether and when exploring this topic, I just got a doubt here (It may be too dumb).&nbsp;</p>\n\n<p>I understand that the given test data would be analyzed and using a logistic regression or any other algo, an equation is formed to classify the files.</p>\n<p>Once the equation is formed to calculate the probability we can test it using the test data. I am missing out at one point here, in the test data an altogether new file comes to get processed, how can we relate it. I mean how can we get the relationship between two malware files.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66266",
      "postDate": "03/15/2015 20:11:38",
      "content": "<p>It's not very clear what you're asking. You have the training data where each file is assigned a malware type (each file has two parts, for example: <em>foobar123.byte&nbsp;</em>and a<em> foobar123.asm</em>, but you can think of these two files as being one big file with information about that malware).</p>\n<p>So you extract information from<em> foobar123.byte</em> and <em>foobar123.asm&nbsp;</em>and&nbsp;look at <em>trainLabels.csv</em> to see which malware type&nbsp;<em>foobar123 </em>belongs to. This will give your training set Y, x1, x2, ..., xn, where each x is a feature extracted from <em>foo123 </em>and Y is the malware Class.</p>\n<p>Then you fit a model and do the same feature extraction using&nbsp;the <em>test.7z</em> file. Then apply your model to predict the labels and submit it to Kaggle.</p>\n<p>This is the standard way to approach this problem, hope it helps.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 66266,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "03/15/2015 20:11:38",
      "content": "<p>It's not very clear what you're asking. You have the training data where each file is assigned a malware type (each file has two parts, for example: <em>foobar123.byte&nbsp;</em>and a<em> foobar123.asm</em>, but you can think of these two files as being one big file with information about that malware).</p>\n<p>So you extract information from<em> foobar123.byte</em> and <em>foobar123.asm&nbsp;</em>and&nbsp;look at <em>trainLabels.csv</em> to see which malware type&nbsp;<em>foobar123 </em>belongs to. This will give your training set Y, x1, x2, ..., xn, where each x is a feature extracted from <em>foo123 </em>and Y is the malware Class.</p>\n<p>Then you fit a model and do the same feature extraction using&nbsp;the <em>test.7z</em> file. Then apply your model to predict the labels and submit it to Kaggle.</p>\n<p>This is the standard way to approach this problem, hope it helps.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "65176": "",
    "66266": ""
  },
  "source": "meta"
}