{
  "id": 20250,
  "title": "Logistic Regression gives weird results",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20250",
  "author_name": "",
  "post_date": "2016-04-19T07:10:14.523Z",
  "votes": null,
  "comment_count": 3,
  "views": 700,
  "content": "<p>I wanted to see how well logistic regression with stochastic gradient descent performs on this dataset. So, I used a train_dataset size of 16000, test_dataset size of 16000 and validation_dataset size of 3000. (I downsampled the images by 20 to train the model faster) Here's the code snippet</p>\n\n<pre><code>logistic = linear_model.LogisticRegression()\nprint('LogisticRegression score_1: %f'\n      % logistic.fit(X_train_reshaped, Y_train).score(X_test_reshaped, Y_test))\nprint('LogisticRegression score_2: %f'\n      % logistic.fit(X_train_reshaped, Y_train).score(X_valid_test_reshaped, Y_valid_test))\n</code></pre>\n\n<p>Here are the results I got:\nscore_1: 0.11\nscore_2 : 0.98</p>\n\n<p>An accuracy of 11 percent?! What could probably be going wrong? (I have attached the script as well if somebody wants to have a look at it)</p>\n\n<p>Any help will be really appreciated! Thanks</p>",
  "messages": [
    {
      "id": "115571",
      "postDate": "04/19/2016 07:10:14",
      "content": "<p>I wanted to see how well logistic regression with stochastic gradient descent performs on this dataset. So, I used a train_dataset size of 16000, test_dataset size of 16000 and validation_dataset size of 3000. (I downsampled the images by 20 to train the model faster) Here's the code snippet</p>\n\n<pre><code>logistic = linear_model.LogisticRegression()\nprint('LogisticRegression score_1: %f'\n      % logistic.fit(X_train_reshaped, Y_train).score(X_test_reshaped, Y_test))\nprint('LogisticRegression score_2: %f'\n      % logistic.fit(X_train_reshaped, Y_train).score(X_valid_test_reshaped, Y_valid_test))\n</code></pre>\n\n<p>Here are the results I got:\nscore_1: 0.11\nscore_2 : 0.98</p>\n\n<p>An accuracy of 11 percent?! What could probably be going wrong? (I have attached the script as well if somebody wants to have a look at it)</p>\n\n<p>Any help will be really appreciated! Thanks</p>",
      "rawMarkdown": "I wanted to see how well logistic regression with stochastic gradient descent performs on this dataset. So, I used a train_dataset size of 16000, test_dataset size of 16000 and validation_dataset size of 3000. (I downsampled the images by 20 to train the model faster) Here's the code snippet\r\n\r\n    logistic = linear_model.LogisticRegression()\r\n    print('LogisticRegression score_1: %f'\r\n          % logistic.fit(X_train_reshaped, Y_train).score(X_test_reshaped, Y_test))\r\n    print('LogisticRegression score_2: %f'\r\n          % logistic.fit(X_train_reshaped, Y_train).score(X_valid_test_reshaped, Y_valid_test))\r\n\r\nHere are the results I got:\r\nscore_1: 0.11\r\nscore_2 : 0.98\r\n\r\nAn accuracy of 11 percent?! What could probably be going wrong? (I have attached the script as well if somebody wants to have a look at it)\r\n\r\nAny help will be really appreciated! Thanks",
      "votes": null
    },
    {
      "id": "115575",
      "postDate": "04/19/2016 07:24:39",
      "content": "<p>If you are using logistic regression directly on pixel values from these images, that's the kind of result I would expect. Each individual pixel value is going to have very low correlation with any specific class. Untreated pixels make very poor predictors in linear models.</p>\n\n<p>You might get better results if you pre-processed the image into features. E.g. a &quot;visual bag of words&quot; might give you better accuracy (not competitive, but better than 11%).</p>",
      "rawMarkdown": "If you are using logistic regression directly on pixel values from these images, that's the kind of result I would expect. Each individual pixel value is going to have very low correlation with any specific class. Untreated pixels make very poor predictors in linear models.\r\n\r\nYou might get better results if you pre-processed the image into features. E.g. a \"visual bag of words\" might give you better accuracy (not competitive, but better than 11%).",
      "votes": null
    },
    {
      "id": "115577",
      "postDate": "04/19/2016 07:32:09",
      "content": "<p>Thanks for your reply Neil.</p>\n\n<p>I'm converting the entire dataset into a 3D array (image index, x, y) of floating point values, normalized to have approximately zero mean and standard deviation ~0.5. I did use this approach in training the classifier on MNIST dataset and it did just fine.</p>\n\n<p>Here'e the code snippet that shows how I pre process the image:</p>\n\n<pre><code>downsample = 20\nimage_width = 640 / downsample #640\nimage_height = 480 / downsample# Pixel width and height.\npixel_depth = 255.0  # Number of levels per pixel.\n\ndef load_letter(folder, min_num_images):\n  &quot;&quot;&quot;Load the data for a single letter label.&quot;&quot;&quot;\n  image_files = os.listdir(folder)\n  #image_files = ['img_134.png']\n  #print(image_files)  \n  dataset = np.ndarray(shape=(len(image_files), image_height, image_width,3),\n                         dtype=np.float32)\n  print(folder)\n  i = 0\n  for image_index, image in enumerate(image_files):\n\n    image_file = os.path.join(folder, image)\n\n    try:\n      img =   ndimage.imread(image_file)\n      resized_img = imresize(img,(image_height,image_width))\n      image_data = (resized_img.astype(float) - \n                    pixel_depth / 2) / pixel_depth\n      #print (image_data.shape)\n\n      if image_data.shape  != (image_height, image_width,3):\n        raise Exception('Unexpected image shape: %s' % str(image_data.shape))\n      #print(image_data)\n      dataset[image_index, :, :] = image_data\n      i = i + 1\n      if((i % 1000) == 0):\n          print(i)  \n    except IOError as e:\n      print('Could not read:', image_file, ':', e, '- it\\'s ok, skipping.')\n\n  num_images = image_index + 1\n  dataset = dataset[0:num_images, :, :]\n  if num_images &lt; min_num_images:\n    raise Exception('Many fewer images than expected: %d &lt; %d' %\n                    (num_images, min_num_images))\n\n  print('Full dataset tensor:', dataset.shape)\n  print('Mean:', np.mean(dataset))\n  print('Standard deviation:', np.std(dataset))\n  return dataset\n\ndef maybe_pickle(data_folders, min_num_images_per_class, force=False):\n  dataset_names = []\n  for folder in data_folders:\n    set_filename = folder + '.pickle'\n    dataset_names.append(set_filename)\n    if os.path.exists(set_filename) and not force:\n      # You may override by setting force=True.\n      print('%s already present - Skipping pickling.' % set_filename)\n    else:\n      print('Pickling %s.' % set_filename)\n\n      dataset = load_letter(folder, min_num_images_per_class)\n      print(&quot;PICKLING..&quot;)\n      try:\n        with open(set_filename, 'wb') as f:\n          print(&quot;DUMPING...&quot;)\n          pickle.dump(dataset, f, pickle.HIGHEST_PROTOCOL)\n      except Exception as e:\n        print('Unable to save data to', set_filename, ':', e)\n\n  return dataset_names\n</code></pre>\n\n<p>Can you tell me what do you mean by &quot;visual bag of words&quot; and how I could get started with that? Thanks</p>",
      "rawMarkdown": "Thanks for your reply Neil.\r\n\r\nI'm converting the entire dataset into a 3D array (image index, x, y) of floating point values, normalized to have approximately zero mean and standard deviation ~0.5. I did use this approach in training the classifier on MNIST dataset and it did just fine.\r\n\r\nHere'e the code snippet that shows how I pre process the image:\r\n\r\n    downsample = 20\r\n    image_width = 640 / downsample #640\r\n    image_height = 480 / downsample# Pixel width and height.\r\n    pixel_depth = 255.0  # Number of levels per pixel.\r\n    \r\n    def load_letter(folder, min_num_images):\r\n      \"\"\"Load the data for a single letter label.\"\"\"\r\n      image_files = os.listdir(folder)\r\n      #image_files = ['img_134.png']\r\n      #print(image_files)  \r\n      dataset = np.ndarray(shape=(len(image_files), image_height, image_width,3),\r\n                             dtype=np.float32)\r\n      print(folder)\r\n      i = 0\r\n      for image_index, image in enumerate(image_files):\r\n        \r\n        image_file = os.path.join(folder, image)\r\n        \r\n        try:\r\n          img =   ndimage.imread(image_file)\r\n          resized_img = imresize(img,(image_height,image_width))\r\n          image_data = (resized_img.astype(float) - \r\n                        pixel_depth / 2) / pixel_depth\r\n          #print (image_data.shape)\r\n          \r\n          if image_data.shape  != (image_height, image_width,3):\r\n            raise Exception('Unexpected image shape: %s' % str(image_data.shape))\r\n          #print(image_data)\r\n          dataset[image_index, :, :] = image_data\r\n          i = i + 1\r\n          if((i % 1000) == 0):\r\n              print(i)  \r\n        except IOError as e:\r\n          print('Could not read:', image_file, ':', e, '- it\\'s ok, skipping.')\r\n        \r\n      num_images = image_index + 1\r\n      dataset = dataset[0:num_images, :, :]\r\n      if num_images < min_num_images:\r\n        raise Exception('Many fewer images than expected: %d < %d' %\r\n                        (num_images, min_num_images))\r\n        \r\n      print('Full dataset tensor:', dataset.shape)\r\n      print('Mean:', np.mean(dataset))\r\n      print('Standard deviation:', np.std(dataset))\r\n      return dataset\r\n            \r\n    def maybe_pickle(data_folders, min_num_images_per_class, force=False):\r\n      dataset_names = []\r\n      for folder in data_folders:\r\n        set_filename = folder + '.pickle'\r\n        dataset_names.append(set_filename)\r\n        if os.path.exists(set_filename) and not force:\r\n          # You may override by setting force=True.\r\n          print('%s already present - Skipping pickling.' % set_filename)\r\n        else:\r\n          print('Pickling %s.' % set_filename)\r\n         \r\n          dataset = load_letter(folder, min_num_images_per_class)\r\n          print(\"PICKLING..\")\r\n          try:\r\n            with open(set_filename, 'wb') as f:\r\n              print(\"DUMPING...\")\r\n              pickle.dump(dataset, f, pickle.HIGHEST_PROTOCOL)\r\n          except Exception as e:\r\n            print('Unable to save data to', set_filename, ':', e)\r\n             \r\n      return dataset_names\r\n\r\nCan you tell me what do you mean by \"visual bag of words\" and how I could get started with that? Thanks",
      "votes": null
    },
    {
      "id": "115592",
      "postDate": "04/19/2016 08:45:03",
      "content": "<p>The MNIST data is much simpler and cleaner than the photographs in this competition, so there is <em>some</em> correlation between pixel values and likely class in MNIST.</p>\n\n<p>A visual bag of words uses counts of matching simple pixel patterns (e.g. edges in various orientations) from the image as the features to learn from. Here is a Python implentation: <a href=\"https://github.com/shackenberg/Minimal-Bag-of-Visual-Words-Image-Classifier\">https://github.com/shackenberg/Minimal-Bag-of-Visual-Words-Image-Classifier</a></p>",
      "rawMarkdown": "The MNIST data is much simpler and cleaner than the photographs in this competition, so there is *some* correlation between pixel values and likely class in MNIST.\r\n\r\nA visual bag of words uses counts of matching simple pixel patterns (e.g. edges in various orientations) from the image as the features to learn from. Here is a Python implentation: https://github.com/shackenberg/Minimal-Bag-of-Visual-Words-Image-Classifier",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 115575,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/19/2016 07:24:39",
      "content": "<p>If you are using logistic regression directly on pixel values from these images, that's the kind of result I would expect. Each individual pixel value is going to have very low correlation with any specific class. Untreated pixels make very poor predictors in linear models.</p>\n\n<p>You might get better results if you pre-processed the image into features. E.g. a &quot;visual bag of words&quot; might give you better accuracy (not competitive, but better than 11%).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115577,
      "author_name": "abhijayvuyyuru",
      "author_url": "",
      "post_date": "04/19/2016 07:32:09",
      "content": "<p>Thanks for your reply Neil.</p>\n\n<p>I'm converting the entire dataset into a 3D array (image index, x, y) of floating point values, normalized to have approximately zero mean and standard deviation ~0.5. I did use this approach in training the classifier on MNIST dataset and it did just fine.</p>\n\n<p>Here'e the code snippet that shows how I pre process the image:</p>\n\n<pre><code>downsample = 20\nimage_width = 640 / downsample #640\nimage_height = 480 / downsample# Pixel width and height.\npixel_depth = 255.0  # Number of levels per pixel.\n\ndef load_letter(folder, min_num_images):\n  &quot;&quot;&quot;Load the data for a single letter label.&quot;&quot;&quot;\n  image_files = os.listdir(folder)\n  #image_files = ['img_134.png']\n  #print(image_files)  \n  dataset = np.ndarray(shape=(len(image_files), image_height, image_width,3),\n                         dtype=np.float32)\n  print(folder)\n  i = 0\n  for image_index, image in enumerate(image_files):\n\n    image_file = os.path.join(folder, image)\n\n    try:\n      img =   ndimage.imread(image_file)\n      resized_img = imresize(img,(image_height,image_width))\n      image_data = (resized_img.astype(float) - \n                    pixel_depth / 2) / pixel_depth\n      #print (image_data.shape)\n\n      if image_data.shape  != (image_height, image_width,3):\n        raise Exception('Unexpected image shape: %s' % str(image_data.shape))\n      #print(image_data)\n      dataset[image_index, :, :] = image_data\n      i = i + 1\n      if((i % 1000) == 0):\n          print(i)  \n    except IOError as e:\n      print('Could not read:', image_file, ':', e, '- it\\'s ok, skipping.')\n\n  num_images = image_index + 1\n  dataset = dataset[0:num_images, :, :]\n  if num_images &lt; min_num_images:\n    raise Exception('Many fewer images than expected: %d &lt; %d' %\n                    (num_images, min_num_images))\n\n  print('Full dataset tensor:', dataset.shape)\n  print('Mean:', np.mean(dataset))\n  print('Standard deviation:', np.std(dataset))\n  return dataset\n\ndef maybe_pickle(data_folders, min_num_images_per_class, force=False):\n  dataset_names = []\n  for folder in data_folders:\n    set_filename = folder + '.pickle'\n    dataset_names.append(set_filename)\n    if os.path.exists(set_filename) and not force:\n      # You may override by setting force=True.\n      print('%s already present - Skipping pickling.' % set_filename)\n    else:\n      print('Pickling %s.' % set_filename)\n\n      dataset = load_letter(folder, min_num_images_per_class)\n      print(&quot;PICKLING..&quot;)\n      try:\n        with open(set_filename, 'wb') as f:\n          print(&quot;DUMPING...&quot;)\n          pickle.dump(dataset, f, pickle.HIGHEST_PROTOCOL)\n      except Exception as e:\n        print('Unable to save data to', set_filename, ':', e)\n\n  return dataset_names\n</code></pre>\n\n<p>Can you tell me what do you mean by &quot;visual bag of words&quot; and how I could get started with that? Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115592,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/19/2016 08:45:03",
      "content": "<p>The MNIST data is much simpler and cleaner than the photographs in this competition, so there is <em>some</em> correlation between pixel values and likely class in MNIST.</p>\n\n<p>A visual bag of words uses counts of matching simple pixel patterns (e.g. edges in various orientations) from the image as the features to learn from. Here is a Python implentation: <a href=\"https://github.com/shackenberg/Minimal-Bag-of-Visual-Words-Image-Classifier\">https://github.com/shackenberg/Minimal-Bag-of-Visual-Words-Image-Classifier</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "115571": "I wanted to see how well logistic regression with stochastic gradient descent performs on this dataset. So, I used a train_dataset size of 16000, test_dataset size of 16000 and validation_dataset size of 3000. (I downsampled the images by 20 to train the model faster) Here's the code snippet\r\n\r\n    logistic = linear_model.LogisticRegression()\r\n    print('LogisticRegression score_1: %f'\r\n          % logistic.fit(X_train_reshaped, Y_train).score(X_test_reshaped, Y_test))\r\n    print('LogisticRegression score_2: %f'\r\n          % logistic.fit(X_train_reshaped, Y_train).score(X_valid_test_reshaped, Y_valid_test))\r\n\r\nHere are the results I got:\r\nscore_1: 0.11\r\nscore_2 : 0.98\r\n\r\nAn accuracy of 11 percent?! What could probably be going wrong? (I have attached the script as well if somebody wants to have a look at it)\r\n\r\nAny help will be really appreciated! Thanks",
    "115575": "If you are using logistic regression directly on pixel values from these images, that's the kind of result I would expect. Each individual pixel value is going to have very low correlation with any specific class. Untreated pixels make very poor predictors in linear models.\r\n\r\nYou might get better results if you pre-processed the image into features. E.g. a \"visual bag of words\" might give you better accuracy (not competitive, but better than 11%).",
    "115577": "Thanks for your reply Neil.\r\n\r\nI'm converting the entire dataset into a 3D array (image index, x, y) of floating point values, normalized to have approximately zero mean and standard deviation ~0.5. I did use this approach in training the classifier on MNIST dataset and it did just fine.\r\n\r\nHere'e the code snippet that shows how I pre process the image:\r\n\r\n    downsample = 20\r\n    image_width = 640 / downsample #640\r\n    image_height = 480 / downsample# Pixel width and height.\r\n    pixel_depth = 255.0  # Number of levels per pixel.\r\n    \r\n    def load_letter(folder, min_num_images):\r\n      \"\"\"Load the data for a single letter label.\"\"\"\r\n      image_files = os.listdir(folder)\r\n      #image_files = ['img_134.png']\r\n      #print(image_files)  \r\n      dataset = np.ndarray(shape=(len(image_files), image_height, image_width,3),\r\n                             dtype=np.float32)\r\n      print(folder)\r\n      i = 0\r\n      for image_index, image in enumerate(image_files):\r\n        \r\n        image_file = os.path.join(folder, image)\r\n        \r\n        try:\r\n          img =   ndimage.imread(image_file)\r\n          resized_img = imresize(img,(image_height,image_width))\r\n          image_data = (resized_img.astype(float) - \r\n                        pixel_depth / 2) / pixel_depth\r\n          #print (image_data.shape)\r\n          \r\n          if image_data.shape  != (image_height, image_width,3):\r\n            raise Exception('Unexpected image shape: %s' % str(image_data.shape))\r\n          #print(image_data)\r\n          dataset[image_index, :, :] = image_data\r\n          i = i + 1\r\n          if((i % 1000) == 0):\r\n              print(i)  \r\n        except IOError as e:\r\n          print('Could not read:', image_file, ':', e, '- it\\'s ok, skipping.')\r\n        \r\n      num_images = image_index + 1\r\n      dataset = dataset[0:num_images, :, :]\r\n      if num_images < min_num_images:\r\n        raise Exception('Many fewer images than expected: %d < %d' %\r\n                        (num_images, min_num_images))\r\n        \r\n      print('Full dataset tensor:', dataset.shape)\r\n      print('Mean:', np.mean(dataset))\r\n      print('Standard deviation:', np.std(dataset))\r\n      return dataset\r\n            \r\n    def maybe_pickle(data_folders, min_num_images_per_class, force=False):\r\n      dataset_names = []\r\n      for folder in data_folders:\r\n        set_filename = folder + '.pickle'\r\n        dataset_names.append(set_filename)\r\n        if os.path.exists(set_filename) and not force:\r\n          # You may override by setting force=True.\r\n          print('%s already present - Skipping pickling.' % set_filename)\r\n        else:\r\n          print('Pickling %s.' % set_filename)\r\n         \r\n          dataset = load_letter(folder, min_num_images_per_class)\r\n          print(\"PICKLING..\")\r\n          try:\r\n            with open(set_filename, 'wb') as f:\r\n              print(\"DUMPING...\")\r\n              pickle.dump(dataset, f, pickle.HIGHEST_PROTOCOL)\r\n          except Exception as e:\r\n            print('Unable to save data to', set_filename, ':', e)\r\n             \r\n      return dataset_names\r\n\r\nCan you tell me what do you mean by \"visual bag of words\" and how I could get started with that? Thanks",
    "115592": "The MNIST data is much simpler and cleaner than the photographs in this competition, so there is *some* correlation between pixel values and likely class in MNIST.\r\n\r\nA visual bag of words uses counts of matching simple pixel patterns (e.g. edges in various orientations) from the image as the features to learn from. Here is a Python implentation: https://github.com/shackenberg/Minimal-Bag-of-Visual-Words-Image-Classifier"
  },
  "source": "meta"
}