{
  "id": 57394,
  "title": "Want to discuss on how to process image features",
  "url": "/competitions/avito-demand-prediction/discussion/57394",
  "author_name": "",
  "post_date": "2018-05-23T08:54:27.068881Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm quite interested in some of the kernals on image feature extraction. I think there are several features that could be used as the model features - for example, the 'confidence' of image classification model could be used as a parameter of showing 'clearness', or quality of image. Or we could use blurness or dullness of image as a feature showing image quality. </p>\n\n<p>However, I think the the most significant problem is that there are too much images - over 1,000,000 images - so it take too much storage and processing capacity. (It's too time consuming). For the storage problem, I was trying to make a loop of loading capable amount of images - say, 1000 at a time -  score them, and store them on dataframe, empty the directory / reload images each time. However even I get over the problem of storage, still there's a problem of processing time. I found that scoring a image might take at least 0.5 sec, so processing entire image would take over a week. </p>\n\n<p>Any ideas on how to time-efficiently score images?</p>",
  "messages": [
    {
      "id": "332497",
      "postDate": "05/23/2018 08:54:27",
      "content": "<p>I'm quite interested in some of the kernals on image feature extraction. I think there are several features that could be used as the model features - for example, the 'confidence' of image classification model could be used as a parameter of showing 'clearness', or quality of image. Or we could use blurness or dullness of image as a feature showing image quality. </p>\n\n<p>However, I think the the most significant problem is that there are too much images - over 1,000,000 images - so it take too much storage and processing capacity. (It's too time consuming). For the storage problem, I was trying to make a loop of loading capable amount of images - say, 1000 at a time -  score them, and store them on dataframe, empty the directory / reload images each time. However even I get over the problem of storage, still there's a problem of processing time. I found that scoring a image might take at least 0.5 sec, so processing entire image would take over a week. </p>\n\n<p>Any ideas on how to time-efficiently score images?</p>",
      "rawMarkdown": "I'm quite interested in some of the kernals on image feature extraction. I think there are several features that could be used as the model features - for example, the 'confidence' of image classification model could be used as a parameter of showing 'clearness', or quality of image. Or we could use blurness or dullness of image as a feature showing image quality. \n\nHowever, I think the the most significant problem is that there are too much images - over 1,000,000 images - so it take too much storage and processing capacity. (It's too time consuming). For the storage problem, I was trying to make a loop of loading capable amount of images - say, 1000 at a time -  score them, and store them on dataframe, empty the directory / reload images each time. However even I get over the problem of storage, still there's a problem of processing time. I found that scoring a image might take at least 0.5 sec, so processing entire image would take over a week. \n\nAny ideas on how to time-efficiently score images?",
      "votes": null
    },
    {
      "id": "332505",
      "postDate": "05/23/2018 09:15:35",
      "content": "<p>No idea, I've waiting for my feature extraction process for several days. It seems that would not have much time left for tuning processes.  </p>",
      "rawMarkdown": "No idea, I've waiting for my feature extraction process for several days. It seems that would not have much time left for tuning processes.",
      "votes": null
    },
    {
      "id": "332874",
      "postDate": "05/24/2018 00:55:38",
      "content": "<p>Maybe <a href=\"https://www.kaggle.com/jpmiller/comparing-test-images-to-train-images\">my kernel</a> here can help? I only get the image hashes, but in a similar kernel I pulled 4-5 features in a couple hours. Plus I pull from the zip file to save space and not have to unzip 1.4M images.</p>",
      "rawMarkdown": "Maybe [my kernel](https://www.kaggle.com/jpmiller/comparing-test-images-to-train-images) here can help? I only get the image hashes, but in a similar kernel I pulled 4-5 features in a couple hours. Plus I pull from the zip file to save space and not have to unzip 1.4M images.",
      "votes": null
    },
    {
      "id": "332879",
      "postDate": "05/24/2018 01:20:34",
      "content": "<p>Oh thanks! I think that would help a lot! Learning a lot from yours!</p>",
      "rawMarkdown": "Oh thanks! I think that would help a lot! Learning a lot from yours!",
      "votes": null
    },
    {
      "id": "332880",
      "postDate": "05/24/2018 01:20:58",
      "content": "<p>I was also waiting! takes too long...</p>",
      "rawMarkdown": "I was also waiting! takes too long...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 332505,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "05/23/2018 09:15:35",
      "content": "<p>No idea, I've waiting for my feature extraction process for several days. It seems that would not have much time left for tuning processes.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 332880,
          "author_name": "sukhyun9673",
          "author_url": "",
          "post_date": "05/24/2018 01:20:58",
          "content": "<p>I was also waiting! takes too long...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332874,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "05/24/2018 00:55:38",
      "content": "<p>Maybe <a href=\"https://www.kaggle.com/jpmiller/comparing-test-images-to-train-images\">my kernel</a> here can help? I only get the image hashes, but in a similar kernel I pulled 4-5 features in a couple hours. Plus I pull from the zip file to save space and not have to unzip 1.4M images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 332879,
          "author_name": "sukhyun9673",
          "author_url": "",
          "post_date": "05/24/2018 01:20:34",
          "content": "<p>Oh thanks! I think that would help a lot! Learning a lot from yours!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "332497": "I'm quite interested in some of the kernals on image feature extraction. I think there are several features that could be used as the model features - for example, the 'confidence' of image classification model could be used as a parameter of showing 'clearness', or quality of image. Or we could use blurness or dullness of image as a feature showing image quality. \n\nHowever, I think the the most significant problem is that there are too much images - over 1,000,000 images - so it take too much storage and processing capacity. (It's too time consuming). For the storage problem, I was trying to make a loop of loading capable amount of images - say, 1000 at a time -  score them, and store them on dataframe, empty the directory / reload images each time. However even I get over the problem of storage, still there's a problem of processing time. I found that scoring a image might take at least 0.5 sec, so processing entire image would take over a week. \n\nAny ideas on how to time-efficiently score images?",
    "332505": "No idea, I've waiting for my feature extraction process for several days. It seems that would not have much time left for tuning processes.",
    "332874": "Maybe [my kernel](https://www.kaggle.com/jpmiller/comparing-test-images-to-train-images) here can help? I only get the image hashes, but in a similar kernel I pulled 4-5 features in a couple hours. Plus I pull from the zip file to save space and not have to unzip 1.4M images.",
    "332879": "Oh thanks! I think that would help a lot! Learning a lot from yours!",
    "332880": "I was also waiting! takes too long..."
  },
  "source": "meta"
}