{
  "id": 2538,
  "title": "basic_benchmark feature vaules",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2538",
  "author_name": "",
  "post_date": "2012-09-01T20:22:05.790Z",
  "votes": null,
  "comment_count": 2,
  "views": 1357,
  "content": "<p>Hello,</p>\r\n<p>I have never used pandas and don't quite understand the DataFrame object, but I was looking at the baseline feature values and something seemed strange.&nbsp; After this line:</p>\r\n<blockquote>\r\n<p>fea = features.extract_features(feature_names, dataTrain)</p>\r\n</blockquote>\r\n<p>I tried to convert 'fea' into a simple numpy matrix array of values:</p>\r\n<blockquote>\r\n<p>feaX = np.array([fea.ix[i].tolist() for i in xrange(fea.shape[0])])</p>\r\n</blockquote>\r\n<p>feaX.shape is (140272,6) as expected, but some of the derived values look wrong.&nbsp; For example, BodyLength:</p>\r\n<blockquote>\r\n<p>print set(feaX[:,0])<br>\r\nset([140272.0])</p>\r\n</blockquote>\r\n<p>i.e. there is only one unique value of this feature!?&nbsp; The same goes for title length feature.&nbsp; Am I converting to numpy matrix wrong? or is there something wrong with the feature extraction code?</p>\r\n<p><br>\r\n<br>\r\n<br>\r\n</p>\r\n<p><br>\r\n<br>\r\n</p>\r\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "13773",
      "postDate": "09/01/2012 20:22:05",
      "content": "<p>Hello,</p>\r\n<p>I have never used pandas and don't quite understand the DataFrame object, but I was looking at the baseline feature values and something seemed strange.&nbsp; After this line:</p>\r\n<blockquote>\r\n<p>fea = features.extract_features(feature_names, dataTrain)</p>\r\n</blockquote>\r\n<p>I tried to convert 'fea' into a simple numpy matrix array of values:</p>\r\n<blockquote>\r\n<p>feaX = np.array([fea.ix[i].tolist() for i in xrange(fea.shape[0])])</p>\r\n</blockquote>\r\n<p>feaX.shape is (140272,6) as expected, but some of the derived values look wrong.&nbsp; For example, BodyLength:</p>\r\n<blockquote>\r\n<p>print set(feaX[:,0])<br>\r\nset([140272.0])</p>\r\n</blockquote>\r\n<p>i.e. there is only one unique value of this feature!?&nbsp; The same goes for title length feature.&nbsp; Am I converting to numpy matrix wrong? or is there something wrong with the feature extraction code?</p>\r\n<p><br>\r\n<br>\r\n<br>\r\n</p>\r\n<p><br>\r\n<br>\r\n</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13774",
      "postDate": "09/01/2012 20:24:26",
      "content": "<p>Update: I looked at how sklearn uses the input data (e.g. in the rf.fit() function), and it does:</p>\r\n<p>X = np.atleast_2d(fea)</p>\r\n<p>And doing this also gives constant values columns 0 and 3.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13775",
      "postDate": "09/01/2012 20:36:01",
      "content": "<p>I features.py, </p>\r\n<p>def body_length(data):<br>\r\nreturn data[&quot;BodyMarkdown&quot;].apply(len)</p>\r\n<p>maybe this applies len() to data[&quot;BodyMarkdown&quot;] rather than the individual elements in data[&quot;BodyMarkdown&quot;] ??</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13774,
      "author_name": "kp2357",
      "author_url": "",
      "post_date": "09/01/2012 20:24:26",
      "content": "<p>Update: I looked at how sklearn uses the input data (e.g. in the rf.fit() function), and it does:</p>\r\n<p>X = np.atleast_2d(fea)</p>\r\n<p>And doing this also gives constant values columns 0 and 3.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13775,
      "author_name": "kp2357",
      "author_url": "",
      "post_date": "09/01/2012 20:36:01",
      "content": "<p>I features.py, </p>\r\n<p>def body_length(data):<br>\r\nreturn data[&quot;BodyMarkdown&quot;].apply(len)</p>\r\n<p>maybe this applies len() to data[&quot;BodyMarkdown&quot;] rather than the individual elements in data[&quot;BodyMarkdown&quot;] ??</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13773": "",
    "13774": "",
    "13775": ""
  },
  "source": "meta"
}