{
  "id": 12605,
  "title": "Multiple Models",
  "url": "/competitions/malware-classification/discussion/12605",
  "author_name": "",
  "post_date": "2015-02-25T01:38:37.990Z",
  "votes": null,
  "comment_count": 7,
  "views": 1908,
  "content": "<p>Hi guys,</p>\n<p>I would like to know a little bit more about how to approach this problem. Do you normally extract some features from the given data to obtain the dataset and then compute a 2D representation of two features. This gives us the limitation of only two features actually being relevant in the data representation.</p>\n<p>Regardless of the statement above, do you determine the values of certain extracted features and split the data accordingly and then use multiple classification models in order to classify samples. This means that based on some criteria, you split the dataset into two parts and send the first part to first classifier model and the second part to second classifier model?&nbsp;</p>\n<p>If you can tell me more about it, I would be forever grateful.</p>\n<p><br>Thank you</p>",
  "messages": [
    {
      "id": "64842",
      "postDate": "02/25/2015 01:38:37",
      "content": "<p>Hi guys,</p>\n<p>I would like to know a little bit more about how to approach this problem. Do you normally extract some features from the given data to obtain the dataset and then compute a 2D representation of two features. This gives us the limitation of only two features actually being relevant in the data representation.</p>\n<p>Regardless of the statement above, do you determine the values of certain extracted features and split the data accordingly and then use multiple classification models in order to classify samples. This means that based on some criteria, you split the dataset into two parts and send the first part to first classifier model and the second part to second classifier model?&nbsp;</p>\n<p>If you can tell me more about it, I would be forever grateful.</p>\n<p><br>Thank you</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64853",
      "postDate": "02/25/2015 08:42:18",
      "content": "<p>The quick answers are no and no.</p>\n<p>You should extract features from the dataset but you don't have to compress them. You could try to get a low-rank approximation of the data matrix (surely keeping much more than just 2 dimensions) and see if this helps. Have a look at the 'beat the benchmark' code. It's easy to follow and will give you an idea.<br><br>Using different classifiers for different parts of the dataset could be an approach but I wouldn't say it's the general case. You can have just one classifier on the whole dataset (or many classifiers trained on the whole dataset).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64948",
      "postDate": "02/26/2015 14:26:40",
      "content": "<p>Hi,</p>\n<p>First of all I would like to thank you for your answers. I did look at the beat the benchmark before and it has been a joyful ride, since I'm kind of new to machine learning.&nbsp;</p>\n<p>In any case, currently I'm using&nbsp;RandomForestClassifier, but I'm getting the following scores: this was done on 500 samples of data, so the numbers are not an actual representation of the entire data.</p>\n<p><code>Training score: 0.987<br>Testing score: 0.960<br></code></p>\n<p>The problem is that the CV score is computed like this:</p>\n<p><code> cm = confusion_matrix(Y, Y_pred)<br>print 'CV: %.3f' % log_loss(Y, Y_prob)<br></code><code>print cm</code></p>\n<p>Which outputs the following:</p>\n<p><code>CV: 0.924<br>[[ 55 0 1 0 0 0 1 0 1]<br> [ 4 116 0 0 0 0 0 1 0]<br> [ 0 1 149 0 0 1 0 0 0]<br> [ 0 0 0 1 0 1 0 2 1]<br> [ 0 0 0 0 2 1 0 0 0]<br> [ 0 1 1 1 1 27 0 2 3]<br> [ 2 0 0 0 0 1 20 0 0]<br> [ 2 1 0 1 1 1 0 51 1]<br> [ 2 0 2 1 0 0 0 0 40]]<br></code></p>\n<p>The problem I'm having is that the CV score is 0.924, but the testing score is&nbsp;0.960. Shouldn't be the CV considerably lower, since the training/testing score is between 92%-96% correct?</p>\n<p><br>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65003",
      "postDate": "02/27/2015 01:33:52",
      "content": "<p>I think the CV score output there is not accuracy, which seems to be how you understand it. The CV score is logloss I think, and logloss is the smaller the better.</p>\n<p>Check the following page:</p>\n<p>http://www.kaggle.com/c/malware-classification/details/evaluation</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65012",
      "postDate": "02/27/2015 08:21:13",
      "content": "<p>[quote=Protean;64948]</p>\n<p>Hi,</p>\n<p>First of all I would like to thank you for your answers. I did look at the beat the benchmark before and it has been a joyful ride, since I'm kind of new to machine learning.&nbsp;</p>\n<p>In any case, currently I'm using&nbsp;RandomForestClassifier, but I'm getting the following scores: this was done on 500 samples of data, so the numbers are not an actual representation of the entire data.</p>\n<p><code>Training score: 0.987<br>Testing score: 0.960<br></code></p>\n<p>The problem is that the CV score is computed like this:</p>\n<p><code> cm = confusion_matrix(Y, Y_pred)<br>print 'CV: %.3f' % log_loss(Y, Y_prob)<br></code><code>print cm</code></p>\n<p>Which outputs the following:</p>\n<p><code>CV: 0.924<br>[[ 55 0 1 0 0 0 1 0 1]<br> [ 4 116 0 0 0 0 0 1 0]<br> [ 0 1 149 0 0 1 0 0 0]<br> [ 0 0 0 1 0 1 0 2 1]<br> [ 0 0 0 0 2 1 0 0 0]<br> [ 0 1 1 1 1 27 0 2 3]<br> [ 2 0 0 0 0 1 20 0 0]<br> [ 2 1 0 1 1 1 0 51 1]<br> [ 2 0 2 1 0 0 0 0 40]]<br></code></p>\n<p>The problem I'm having is that the CV score is 0.924, but the testing score is&nbsp;0.960. Shouldn't be the CV considerably lower, since the training/testing score is between 92%-96% correct?</p>\n<p><br>Thanks</p>\n<p>[/quote]</p>\n<p>With Random Forest you can achieve 0.98 CV score. With GBDT you can achieve 0.99</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65013",
      "postDate": "02/27/2015 08:56:06",
      "content": "<p>Note that multinomial log loss (metric for this competition) and sklearn.metrics.logloss are not equal.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65037",
      "postDate": "02/27/2015 15:28:05",
      "content": "<p>From the sklearn docs:</p>\n<p>sklearn.metrics.log_loss(y_true, y_pred, eps=1e-15, normalize=True)</p>\n<p>Log loss, aka logistic loss or cross-entropy loss.<br>This is the loss function used in (multinomial) logistic regression and extensions of it such as neural networks, defined as the negative log-likelihood of the true labels given a probabilistic classifier&#8217;s predictions. For a single sample with true label yt in {0,1} and estimated probability yp that yt = 1, the log loss is</p>\n<p><br>-log P(yt|yp) = -(yt log(yp) + (1 - yt) log(1 - yp))</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65045",
      "postDate": "02/27/2015 16:31:58",
      "content": "<p>Thank you for all your answers, you've been most helpful.&nbsp;</p>\n<p>I have another question regarding the data used for learning. I've extracted two features from the dataset: the .asm/.bytes sizes ((10868, 2) as described in the BTB code and some other feature&nbsp;((10868, 2) ). I'm building the input training set by using the following code to get the&nbsp;((10868, 3) ) array.</p>\n\n<p><code>X1 = read_filesizes(train)<br>X2 = extract_feature(train)<br>X =&nbsp;numpy.column_stack((X1,X2))</code></p>\n\n<p>Then I'm feeding that input array to the learning model and I'm getting substantially worse results than when only using the X1 file sizes. I guess the problem is that the data doesn't fit together.&nbsp;</p>\n<p>My question is: how can I extract multiple features from the dataset and join the data together into one big array, which is fed to the learning algorithm. What are usual approaches when joining multiple features from dataset and feeding it to the learning model?</p>\n<p>Thank you</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 64853,
      "author_name": "asterios",
      "author_url": "",
      "post_date": "02/25/2015 08:42:18",
      "content": "<p>The quick answers are no and no.</p>\n<p>You should extract features from the dataset but you don't have to compress them. You could try to get a low-rank approximation of the data matrix (surely keeping much more than just 2 dimensions) and see if this helps. Have a look at the 'beat the benchmark' code. It's easy to follow and will give you an idea.<br><br>Using different classifiers for different parts of the dataset could be an approach but I wouldn't say it's the general case. You can have just one classifier on the whole dataset (or many classifiers trained on the whole dataset).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64948,
      "author_name": "protean",
      "author_url": "",
      "post_date": "02/26/2015 14:26:40",
      "content": "<p>Hi,</p>\n<p>First of all I would like to thank you for your answers. I did look at the beat the benchmark before and it has been a joyful ride, since I'm kind of new to machine learning.&nbsp;</p>\n<p>In any case, currently I'm using&nbsp;RandomForestClassifier, but I'm getting the following scores: this was done on 500 samples of data, so the numbers are not an actual representation of the entire data.</p>\n<p><code>Training score: 0.987<br>Testing score: 0.960<br></code></p>\n<p>The problem is that the CV score is computed like this:</p>\n<p><code> cm = confusion_matrix(Y, Y_pred)<br>print 'CV: %.3f' % log_loss(Y, Y_prob)<br></code><code>print cm</code></p>\n<p>Which outputs the following:</p>\n<p><code>CV: 0.924<br>[[ 55 0 1 0 0 0 1 0 1]<br> [ 4 116 0 0 0 0 0 1 0]<br> [ 0 1 149 0 0 1 0 0 0]<br> [ 0 0 0 1 0 1 0 2 1]<br> [ 0 0 0 0 2 1 0 0 0]<br> [ 0 1 1 1 1 27 0 2 3]<br> [ 2 0 0 0 0 1 20 0 0]<br> [ 2 1 0 1 1 1 0 51 1]<br> [ 2 0 2 1 0 0 0 0 40]]<br></code></p>\n<p>The problem I'm having is that the CV score is 0.924, but the testing score is&nbsp;0.960. Shouldn't be the CV considerably lower, since the training/testing score is between 92%-96% correct?</p>\n<p><br>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65003,
      "author_name": "jieqchen",
      "author_url": "",
      "post_date": "02/27/2015 01:33:52",
      "content": "<p>I think the CV score output there is not accuracy, which seems to be how you understand it. The CV score is logloss I think, and logloss is the smaller the better.</p>\n<p>Check the following page:</p>\n<p>http://www.kaggle.com/c/malware-classification/details/evaluation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65012,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "02/27/2015 08:21:13",
      "content": "<p>[quote=Protean;64948]</p>\n<p>Hi,</p>\n<p>First of all I would like to thank you for your answers. I did look at the beat the benchmark before and it has been a joyful ride, since I'm kind of new to machine learning.&nbsp;</p>\n<p>In any case, currently I'm using&nbsp;RandomForestClassifier, but I'm getting the following scores: this was done on 500 samples of data, so the numbers are not an actual representation of the entire data.</p>\n<p><code>Training score: 0.987<br>Testing score: 0.960<br></code></p>\n<p>The problem is that the CV score is computed like this:</p>\n<p><code> cm = confusion_matrix(Y, Y_pred)<br>print 'CV: %.3f' % log_loss(Y, Y_prob)<br></code><code>print cm</code></p>\n<p>Which outputs the following:</p>\n<p><code>CV: 0.924<br>[[ 55 0 1 0 0 0 1 0 1]<br> [ 4 116 0 0 0 0 0 1 0]<br> [ 0 1 149 0 0 1 0 0 0]<br> [ 0 0 0 1 0 1 0 2 1]<br> [ 0 0 0 0 2 1 0 0 0]<br> [ 0 1 1 1 1 27 0 2 3]<br> [ 2 0 0 0 0 1 20 0 0]<br> [ 2 1 0 1 1 1 0 51 1]<br> [ 2 0 2 1 0 0 0 0 40]]<br></code></p>\n<p>The problem I'm having is that the CV score is 0.924, but the testing score is&nbsp;0.960. Shouldn't be the CV considerably lower, since the training/testing score is between 92%-96% correct?</p>\n<p><br>Thanks</p>\n<p>[/quote]</p>\n<p>With Random Forest you can achieve 0.98 CV score. With GBDT you can achieve 0.99</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65013,
      "author_name": "mikhailtrofimov",
      "author_url": "",
      "post_date": "02/27/2015 08:56:06",
      "content": "<p>Note that multinomial log loss (metric for this competition) and sklearn.metrics.logloss are not equal.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65037,
      "author_name": "x75a40890",
      "author_url": "",
      "post_date": "02/27/2015 15:28:05",
      "content": "<p>From the sklearn docs:</p>\n<p>sklearn.metrics.log_loss(y_true, y_pred, eps=1e-15, normalize=True)</p>\n<p>Log loss, aka logistic loss or cross-entropy loss.<br>This is the loss function used in (multinomial) logistic regression and extensions of it such as neural networks, defined as the negative log-likelihood of the true labels given a probabilistic classifier&#8217;s predictions. For a single sample with true label yt in {0,1} and estimated probability yp that yt = 1, the log loss is</p>\n<p><br>-log P(yt|yp) = -(yt log(yp) + (1 - yt) log(1 - yp))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65045,
      "author_name": "protean",
      "author_url": "",
      "post_date": "02/27/2015 16:31:58",
      "content": "<p>Thank you for all your answers, you've been most helpful.&nbsp;</p>\n<p>I have another question regarding the data used for learning. I've extracted two features from the dataset: the .asm/.bytes sizes ((10868, 2) as described in the BTB code and some other feature&nbsp;((10868, 2) ). I'm building the input training set by using the following code to get the&nbsp;((10868, 3) ) array.</p>\n\n<p><code>X1 = read_filesizes(train)<br>X2 = extract_feature(train)<br>X =&nbsp;numpy.column_stack((X1,X2))</code></p>\n\n<p>Then I'm feeding that input array to the learning model and I'm getting substantially worse results than when only using the X1 file sizes. I guess the problem is that the data doesn't fit together.&nbsp;</p>\n<p>My question is: how can I extract multiple features from the dataset and join the data together into one big array, which is fed to the learning algorithm. What are usual approaches when joining multiple features from dataset and feeding it to the learning model?</p>\n<p>Thank you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "64842": "",
    "64853": "",
    "64948": "",
    "65003": "",
    "65012": "",
    "65013": "",
    "65037": "",
    "65045": ""
  },
  "source": "meta"
}