{
  "id": 55315,
  "title": "How to deal with image data?",
  "url": "/competitions/avito-demand-prediction/discussion/55315",
  "author_name": "",
  "post_date": "2018-04-25T05:50:07.293067400Z",
  "votes": 15,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I recommend that extract image features via a CNN model like googlenet or others.</p>\n\n<p>I know there are some papers about how to predict CTR via deep-learning.</p>\n\n<p>Deep CTR Prediction in Display Advertising <a href=\"https://arxiv.org/abs/1609.06018\">https://arxiv.org/abs/1609.06018</a></p>\n\n<p>Would you like share some other method to deal with image data?</p>",
  "messages": [
    {
      "id": "319057",
      "postDate": "04/25/2018 05:50:07",
      "content": "<p>I recommend that extract image features via a CNN model like googlenet or others.</p>\n\n<p>I know there are some papers about how to predict CTR via deep-learning.</p>\n\n<p>Deep CTR Prediction in Display Advertising <a href=\"https://arxiv.org/abs/1609.06018\">https://arxiv.org/abs/1609.06018</a></p>\n\n<p>Would you like share some other method to deal with image data?</p>",
      "rawMarkdown": "I recommend that extract image features via a CNN model like googlenet or others.\n\nI know there are some papers about how to predict CTR via deep-learning.\n\nDeep CTR Prediction in Display Advertising https://arxiv.org/abs/1609.06018\n\nWould you like share some other method to deal with image data?",
      "votes": null
    },
    {
      "id": "319068",
      "postDate": "04/25/2018 06:15:22",
      "content": "<p>I think there are at least two ways.\nOne way is to extract image features using some pretrained network (like ResNet or VGG), optionally compress them using PCA, and consider them as just another feature set for LightGBM/FM/etc.\nAnother possible way, which is less straightforward and more trickier, is to train a new CNN classifier to directly estimate the target variable from images, then use prediction as an independent feature somehow.</p>",
      "rawMarkdown": "I think there are at least two ways.\nOne way is to extract image features using some pretrained network (like ResNet or VGG), optionally compress them using PCA, and consider them as just another feature set for LightGBM/FM/etc.\nAnother possible way, which is less straightforward and more trickier, is to train a new CNN classifier to directly estimate the target variable from images, then use prediction as an independent feature somehow.",
      "votes": null
    },
    {
      "id": "319092",
      "postDate": "04/25/2018 07:54:21",
      "content": "<p>Really good tips.</p>\n\n<p>But I don't make sure the image feature will help get better result, and we should take more effect on it.</p>",
      "rawMarkdown": "Really good tips.\n\nBut I don't make sure the image feature will help get better result, and we should take more effect on it.",
      "votes": null
    },
    {
      "id": "319263",
      "postDate": "04/25/2018 16:12:43",
      "content": "<p>I have try to extract image features via keras vgg16 pretrain-model.</p>\n\n<p>Look at the kernel <a href=\"https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16\">https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16</a></p>",
      "rawMarkdown": "I have try to extract image features via keras vgg16 pretrain-model.\n\nLook at the kernel https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16",
      "votes": null
    },
    {
      "id": "319298",
      "postDate": "04/25/2018 18:07:50",
      "content": "<p>Wonderful! I think you should definitely try to incorporate these embeddings into your classification pipeline.</p>",
      "rawMarkdown": "Wonderful! I think you should definitely try to incorporate these embeddings into your classification pipeline.",
      "votes": null
    },
    {
      "id": "319444",
      "postDate": "04/26/2018 03:41:30",
      "content": "<p>I'm working on it.</p>",
      "rawMarkdown": "I'm working on it.",
      "votes": null
    },
    {
      "id": "319577",
      "postDate": "04/26/2018 10:57:33",
      "content": "<p>I think, that wide &amp; deep approach make work here <a href=\"https://www.tensorflow.org/tutorials/wide_and_deep\">https://www.tensorflow.org/tutorials/wide_and_deep</a> where you have for example 3 networks (one for images, one for text and one for the rest) combined in one. </p>",
      "rawMarkdown": "I think, that wide &amp; deep approach make work here https://www.tensorflow.org/tutorials/wide_and_deep where you have for example 3 networks (one for images, one for text and one for the rest) combined in one.",
      "votes": null
    },
    {
      "id": "320093",
      "postDate": "04/27/2018 13:04:03",
      "content": "<p>I try to unzip each file and delete it just after to get some features, but this process is too long but I don't know how is it possible to deal with those images ? I just want to have some small features like mean, median, variance, min, max and image dimension in order to give a \"quality score\" to the picture.  </p>",
      "rawMarkdown": "I try to unzip each file and delete it just after to get some features, but this process is too long but I don't know how is it possible to deal with those images ? I just want to have some small features like mean, median, variance, min, max and image dimension in order to give a \"quality score\" to the picture.",
      "votes": null
    },
    {
      "id": "320189",
      "postDate": "04/27/2018 20:17:15",
      "content": "<p>I have a idea is that, we can extract feature via pre-train cnn model like e.g. vgg, resnet , then our target value is image_top_1 in train set columns, then train a classifiler model , you can custom your layer( like use 10 dims in pre-last layer), at last, we extract 10 dims features from our model. If give me a GPU machine, I really want to try this solution.</p>",
      "rawMarkdown": "I have a idea is that, we can extract feature via pre-train cnn model like e.g. vgg, resnet , then our target value is image_top_1 in train set columns, then train a classifiler model , you can custom your layer( like use 10 dims in pre-last layer), at last, we extract 10 dims features from our model. If give me a GPU machine, I really want to try this solution.",
      "votes": null
    },
    {
      "id": "320593",
      "postDate": "04/29/2018 07:07:46",
      "content": "<p>Will compressed features from an autoencoder work better than PCA from pre-trained model?</p>",
      "rawMarkdown": "Will compressed features from an autoencoder work better than PCA from pre-trained model?",
      "votes": null
    },
    {
      "id": "320796",
      "postDate": "04/29/2018 21:59:33",
      "content": "<p>Great!</p>",
      "rawMarkdown": "Great!",
      "votes": null
    },
    {
      "id": "320828",
      "postDate": "04/30/2018 01:10:26",
      "content": "<p>I just published a viable image-features generator inside a Kernel: <a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features\">https://www.kaggle.com/bguberfain/vgg16-train-features</a></p>",
      "rawMarkdown": "I just published a viable image-features generator inside a Kernel: https://www.kaggle.com/bguberfain/vgg16-train-features",
      "votes": null
    },
    {
      "id": "321609",
      "postDate": "05/01/2018 16:52:51",
      "content": "<p>With autoencoder, you can have extra depth and by adding extra constraints on the feature space (for example, by adding sparsity loss), you can have a representation that is more suitable for a problem you want to solve. The downside can be that it might be harder to train and you might get junk representations if you don't do it correctly.</p>",
      "rawMarkdown": "With autoencoder, you can have extra depth and by adding extra constraints on the feature space (for example, by adding sparsity loss), you can have a representation that is more suitable for a problem you want to solve. The downside can be that it might be harder to train and you might get junk representations if you don't do it correctly.",
      "votes": null
    },
    {
      "id": "321809",
      "postDate": "05/02/2018 00:05:40",
      "content": "<p>I recall a meetup where a presenter, working in movie recommendation systems, explained what features turned out to be important to viewers. He explained that the main colour of a movie's poster was most important: if a movie poster's main colour was the same, or similar to the last movie watched, then it was most likely to be selected as the next movie watched. While the setting of this competition is different, I think a basic feature like colour might be useful.</p>\n\n<p>As an amateur photographer, I pay attention to things such as the histogram (for contrast), colours, composition (of objects), and saturation of colours, among other things. Maybe these could be extracted from images too, to see whether they are important?</p>",
      "rawMarkdown": "I recall a meetup where a presenter, working in movie recommendation systems, explained what features turned out to be important to viewers. He explained that the main colour of a movie's poster was most important: if a movie poster's main colour was the same, or similar to the last movie watched, then it was most likely to be selected as the next movie watched. While the setting of this competition is different, I think a basic feature like colour might be useful.\n\nAs an amateur photographer, I pay attention to things such as the histogram (for contrast), colours, composition (of objects), and saturation of colours, among other things. Maybe these could be extracted from images too, to see whether they are important?",
      "votes": null
    },
    {
      "id": "321827",
      "postDate": "05/02/2018 01:33:49",
      "content": "<p>Quentin,\nHere's some code you can modify to get image data. I was able to do resizing, color conversions, and feature extractions on the train set in about 30 minutes.</p>\n\n<pre><code>import cv2\nfrom dask import bag, threaded\nfrom dask.diagnostics import ProgressBar\n\ndef get_props(fname):\n    exfile = zipped.read(fname)\n    arr = np.frombuffer(exfile, np.uint8)\n    if arr.size &gt; 0:   # exclude dirs and blanks\n        avgpx = np.mean(arr)\n    else: \n        avgpx = 0\n    return fname, arr.size, avgpx\n\ndef get_image_data(archive_name):\n    global zipped\n    zipped = ZipFile(archive_name)\n    names = zipped.namelist()\n    propbag = bag.from_sequence(names).map(get_props)\n    with ProgressBar():\n        props = propbag.compute(get=threaded.get)\n    props_df = pd.DataFrame(props, columns=['fname', 'arr_size', 'avgpx'])\n    return props_df\n\ntrain_img_data = get_image_data('./train_jpg.zip')\n</code></pre>",
      "rawMarkdown": "Quentin,\nHere's some code you can modify to get image data. I was able to do resizing, color conversions, and feature extractions on the train set in about 30 minutes.\n\n    import cv2\n    from dask import bag, threaded\n    from dask.diagnostics import ProgressBar\n\n    def get_props(fname):\n        exfile = zipped.read(fname)\n        arr = np.frombuffer(exfile, np.uint8)\n        if arr.size &gt; 0:   # exclude dirs and blanks\n            avgpx = np.mean(arr)\n        else: \n            avgpx = 0\n        return fname, arr.size, avgpx\n\n    def get_image_data(archive_name):\n        global zipped\n        zipped = ZipFile(archive_name)\n        names = zipped.namelist()\n        propbag = bag.from_sequence(names).map(get_props)\n        with ProgressBar():\n            props = propbag.compute(get=threaded.get)\n        props_df = pd.DataFrame(props, columns=['fname', 'arr_size', 'avgpx'])\n        return props_df\n\n    train_img_data = get_image_data('./train_jpg.zip')",
      "votes": null
    },
    {
      "id": "335300",
      "postDate": "05/29/2018 14:43:16",
      "content": "<p>Here is some links how to extract:\n<a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features\">https://www.kaggle.com/bguberfain/vgg16-train-features</a>\n<a href=\"https://www.kaggle.com/insaff/img-feature-extraction-with-pretrained-resnet\">https://www.kaggle.com/insaff/img-feature-extraction-with-pretrained-resnet</a>\n<a href=\"https://www.kaggle.com/insaff/vgg-feature-extraction-by-pretrained-model\">https://www.kaggle.com/insaff/vgg-feature-extraction-by-pretrained-model</a></p>",
      "rawMarkdown": "Here is some links how to extract:\nhttps://www.kaggle.com/bguberfain/vgg16-train-features\nhttps://www.kaggle.com/insaff/img-feature-extraction-with-pretrained-resnet\nhttps://www.kaggle.com/insaff/vgg-feature-extraction-by-pretrained-model",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 319068,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "04/25/2018 06:15:22",
      "content": "<p>I think there are at least two ways.\nOne way is to extract image features using some pretrained network (like ResNet or VGG), optionally compress them using PCA, and consider them as just another feature set for LightGBM/FM/etc.\nAnother possible way, which is less straightforward and more trickier, is to train a new CNN classifier to directly estimate the target variable from images, then use prediction as an independent feature somehow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 319092,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "04/25/2018 07:54:21",
          "content": "<p>Really good tips.</p>\n\n<p>But I don't make sure the image feature will help get better result, and we should take more effect on it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335300,
          "author_name": "insaff",
          "author_url": "",
          "post_date": "05/29/2018 14:43:16",
          "content": "<p>Here is some links how to extract:\n<a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features\">https://www.kaggle.com/bguberfain/vgg16-train-features</a>\n<a href=\"https://www.kaggle.com/insaff/img-feature-extraction-with-pretrained-resnet\">https://www.kaggle.com/insaff/img-feature-extraction-with-pretrained-resnet</a>\n<a href=\"https://www.kaggle.com/insaff/vgg-feature-extraction-by-pretrained-model\">https://www.kaggle.com/insaff/vgg-feature-extraction-by-pretrained-model</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 319263,
      "author_name": "classtag",
      "author_url": "",
      "post_date": "04/25/2018 16:12:43",
      "content": "<p>I have try to extract image features via keras vgg16 pretrain-model.</p>\n\n<p>Look at the kernel <a href=\"https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16\">https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 319298,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "04/25/2018 18:07:50",
          "content": "<p>Wonderful! I think you should definitely try to incorporate these embeddings into your classification pipeline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319444,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "04/26/2018 03:41:30",
          "content": "<p>I'm working on it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 319577,
      "author_name": "adamszalucha",
      "author_url": "",
      "post_date": "04/26/2018 10:57:33",
      "content": "<p>I think, that wide &amp; deep approach make work here <a href=\"https://www.tensorflow.org/tutorials/wide_and_deep\">https://www.tensorflow.org/tutorials/wide_and_deep</a> where you have for example 3 networks (one for images, one for text and one for the rest) combined in one. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320093,
      "author_name": "ptitmoustique",
      "author_url": "",
      "post_date": "04/27/2018 13:04:03",
      "content": "<p>I try to unzip each file and delete it just after to get some features, but this process is too long but I don't know how is it possible to deal with those images ? I just want to have some small features like mean, median, variance, min, max and image dimension in order to give a \"quality score\" to the picture.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 320189,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "04/27/2018 20:17:15",
          "content": "<p>I have a idea is that, we can extract feature via pre-train cnn model like e.g. vgg, resnet , then our target value is image_top_1 in train set columns, then train a classifiler model , you can custom your layer( like use 10 dims in pre-last layer), at last, we extract 10 dims features from our model. If give me a GPU machine, I really want to try this solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321827,
          "author_name": "jpmiller",
          "author_url": "",
          "post_date": "05/02/2018 01:33:49",
          "content": "<p>Quentin,\nHere's some code you can modify to get image data. I was able to do resizing, color conversions, and feature extractions on the train set in about 30 minutes.</p>\n\n<pre><code>import cv2\nfrom dask import bag, threaded\nfrom dask.diagnostics import ProgressBar\n\ndef get_props(fname):\n    exfile = zipped.read(fname)\n    arr = np.frombuffer(exfile, np.uint8)\n    if arr.size &gt; 0:   # exclude dirs and blanks\n        avgpx = np.mean(arr)\n    else: \n        avgpx = 0\n    return fname, arr.size, avgpx\n\ndef get_image_data(archive_name):\n    global zipped\n    zipped = ZipFile(archive_name)\n    names = zipped.namelist()\n    propbag = bag.from_sequence(names).map(get_props)\n    with ProgressBar():\n        props = propbag.compute(get=threaded.get)\n    props_df = pd.DataFrame(props, columns=['fname', 'arr_size', 'avgpx'])\n    return props_df\n\ntrain_img_data = get_image_data('./train_jpg.zip')\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320593,
      "author_name": "gunnvant",
      "author_url": "",
      "post_date": "04/29/2018 07:07:46",
      "content": "<p>Will compressed features from an autoencoder work better than PCA from pre-trained model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 321609,
          "author_name": "hauserquaid",
          "author_url": "",
          "post_date": "05/01/2018 16:52:51",
          "content": "<p>With autoencoder, you can have extra depth and by adding extra constraints on the feature space (for example, by adding sparsity loss), you can have a representation that is more suitable for a problem you want to solve. The downside can be that it might be harder to train and you might get junk representations if you don't do it correctly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320796,
      "author_name": "dukeacureds",
      "author_url": "",
      "post_date": "04/29/2018 21:59:33",
      "content": "<p>Great!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320828,
      "author_name": "bguberfain",
      "author_url": "",
      "post_date": "04/30/2018 01:10:26",
      "content": "<p>I just published a viable image-features generator inside a Kernel: <a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features\">https://www.kaggle.com/bguberfain/vgg16-train-features</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 321809,
      "author_name": "amhchiu",
      "author_url": "",
      "post_date": "05/02/2018 00:05:40",
      "content": "<p>I recall a meetup where a presenter, working in movie recommendation systems, explained what features turned out to be important to viewers. He explained that the main colour of a movie's poster was most important: if a movie poster's main colour was the same, or similar to the last movie watched, then it was most likely to be selected as the next movie watched. While the setting of this competition is different, I think a basic feature like colour might be useful.</p>\n\n<p>As an amateur photographer, I pay attention to things such as the histogram (for contrast), colours, composition (of objects), and saturation of colours, among other things. Maybe these could be extracted from images too, to see whether they are important?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "319057": "I recommend that extract image features via a CNN model like googlenet or others.\n\nI know there are some papers about how to predict CTR via deep-learning.\n\nDeep CTR Prediction in Display Advertising https://arxiv.org/abs/1609.06018\n\nWould you like share some other method to deal with image data?",
    "319068": "I think there are at least two ways.\nOne way is to extract image features using some pretrained network (like ResNet or VGG), optionally compress them using PCA, and consider them as just another feature set for LightGBM/FM/etc.\nAnother possible way, which is less straightforward and more trickier, is to train a new CNN classifier to directly estimate the target variable from images, then use prediction as an independent feature somehow.",
    "319092": "Really good tips.\n\nBut I don't make sure the image feature will help get better result, and we should take more effect on it.",
    "319263": "I have try to extract image features via keras vgg16 pretrain-model.\n\nLook at the kernel https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16",
    "319298": "Wonderful! I think you should definitely try to incorporate these embeddings into your classification pipeline.",
    "319444": "I'm working on it.",
    "319577": "I think, that wide &amp; deep approach make work here https://www.tensorflow.org/tutorials/wide_and_deep where you have for example 3 networks (one for images, one for text and one for the rest) combined in one.",
    "320093": "I try to unzip each file and delete it just after to get some features, but this process is too long but I don't know how is it possible to deal with those images ? I just want to have some small features like mean, median, variance, min, max and image dimension in order to give a \"quality score\" to the picture.",
    "320189": "I have a idea is that, we can extract feature via pre-train cnn model like e.g. vgg, resnet , then our target value is image_top_1 in train set columns, then train a classifiler model , you can custom your layer( like use 10 dims in pre-last layer), at last, we extract 10 dims features from our model. If give me a GPU machine, I really want to try this solution.",
    "320593": "Will compressed features from an autoencoder work better than PCA from pre-trained model?",
    "320796": "Great!",
    "320828": "I just published a viable image-features generator inside a Kernel: https://www.kaggle.com/bguberfain/vgg16-train-features",
    "321609": "With autoencoder, you can have extra depth and by adding extra constraints on the feature space (for example, by adding sparsity loss), you can have a representation that is more suitable for a problem you want to solve. The downside can be that it might be harder to train and you might get junk representations if you don't do it correctly.",
    "321809": "I recall a meetup where a presenter, working in movie recommendation systems, explained what features turned out to be important to viewers. He explained that the main colour of a movie's poster was most important: if a movie poster's main colour was the same, or similar to the last movie watched, then it was most likely to be selected as the next movie watched. While the setting of this competition is different, I think a basic feature like colour might be useful.\n\nAs an amateur photographer, I pay attention to things such as the histogram (for contrast), colours, composition (of objects), and saturation of colours, among other things. Maybe these could be extracted from images too, to see whether they are important?",
    "321827": "Quentin,\nHere's some code you can modify to get image data. I was able to do resizing, color conversions, and feature extractions on the train set in about 30 minutes.\n\n    import cv2\n    from dask import bag, threaded\n    from dask.diagnostics import ProgressBar\n\n    def get_props(fname):\n        exfile = zipped.read(fname)\n        arr = np.frombuffer(exfile, np.uint8)\n        if arr.size &gt; 0:   # exclude dirs and blanks\n            avgpx = np.mean(arr)\n        else: \n            avgpx = 0\n        return fname, arr.size, avgpx\n\n    def get_image_data(archive_name):\n        global zipped\n        zipped = ZipFile(archive_name)\n        names = zipped.namelist()\n        propbag = bag.from_sequence(names).map(get_props)\n        with ProgressBar():\n            props = propbag.compute(get=threaded.get)\n        props_df = pd.DataFrame(props, columns=['fname', 'arr_size', 'avgpx'])\n        return props_df\n\n    train_img_data = get_image_data('./train_jpg.zip')",
    "335300": "Here is some links how to extract:\nhttps://www.kaggle.com/bguberfain/vgg16-train-features\nhttps://www.kaggle.com/insaff/img-feature-extraction-with-pretrained-resnet\nhttps://www.kaggle.com/insaff/vgg-feature-extraction-by-pretrained-model"
  },
  "source": "meta"
}