{
  "id": 55981,
  "title": "How to go about training a model with images?",
  "url": "/competitions/avito-demand-prediction/discussion/55981",
  "author_name": "",
  "post_date": "2018-05-04T02:28:53.795232700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey all,</p>\n\n<p>I'm fairly new to DL so forgive me if i ask dumb questions lol</p>\n\n<p>My goal is to create model that can take both text and images, but i'm a little stuck as to how to go about building the model.</p>\n\n<p>I've trained a model using the text features, but I want to include images.  I thought of two ways I could do this...\n1. Use a pre-trained model(without the fully connected layers) to extract features from the images.  Then use those features as an input to your model along with the text features, concatenate them, then have a few fully connected layers on top with the last layer having a single output (predicted deal probability). </p>\n\n<ol>\n<li>Use a two input layer that takes in images and text. Use the convolution layers of pre-trained model(such as VGG16/19) for the image input , and a few fully connected layers to the text input. Concatenate these features and add a few fully connected layers on top with the last layer having a single output. </li>\n</ol>\n\n<p>Any thoughts? Hopefully what I said makes sense if not please let me know. \nThanks!</p>",
  "messages": [
    {
      "id": "322955",
      "postDate": "05/04/2018 02:28:53",
      "content": "<p>Hey all,</p>\n\n<p>I'm fairly new to DL so forgive me if i ask dumb questions lol</p>\n\n<p>My goal is to create model that can take both text and images, but i'm a little stuck as to how to go about building the model.</p>\n\n<p>I've trained a model using the text features, but I want to include images.  I thought of two ways I could do this...\n1. Use a pre-trained model(without the fully connected layers) to extract features from the images.  Then use those features as an input to your model along with the text features, concatenate them, then have a few fully connected layers on top with the last layer having a single output (predicted deal probability). </p>\n\n<ol>\n<li>Use a two input layer that takes in images and text. Use the convolution layers of pre-trained model(such as VGG16/19) for the image input , and a few fully connected layers to the text input. Concatenate these features and add a few fully connected layers on top with the last layer having a single output. </li>\n</ol>\n\n<p>Any thoughts? Hopefully what I said makes sense if not please let me know. \nThanks!</p>",
      "rawMarkdown": "Hey all,\n\nI'm fairly new to DL so forgive me if i ask dumb questions lol\n\nMy goal is to create model that can take both text and images, but i'm a little stuck as to how to go about building the model.\n\nI've trained a model using the text features, but I want to include images.  I thought of two ways I could do this...\n1. Use a pre-trained model(without the fully connected layers) to extract features from the images.  Then use those features as an input to your model along with the text features, concatenate them, then have a few fully connected layers on top with the last layer having a single output (predicted deal probability). \n\n2. Use a two input layer that takes in images and text. Use the convolution layers of pre-trained model(such as VGG16/19) for the image input , and a few fully connected layers to the text input. Concatenate these features and add a few fully connected layers on top with the last layer having a single output. \n\nAny thoughts? Hopefully what I said makes sense if not please let me know. \nThanks!",
      "votes": null
    },
    {
      "id": "324003",
      "postDate": "05/06/2018 22:22:14",
      "content": "<p>Something another competitor suggested that seems to work okay is to predict what each image is (power drill, purse, car, etcetera).  The higher the maximum probability would be used as a proxy for how reliable each image is.  I bet you could train your own image classifier on something more specific to this competition, but there's that strategy for now; it seems to be working for me thus far.</p>",
      "rawMarkdown": "Something another competitor suggested that seems to work okay is to predict what each image is (power drill, purse, car, etcetera).  The higher the maximum probability would be used as a proxy for how reliable each image is.  I bet you could train your own image classifier on something more specific to this competition, but there's that strategy for now; it seems to be working for me thus far.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 324003,
      "author_name": "matthewa313",
      "author_url": "",
      "post_date": "05/06/2018 22:22:14",
      "content": "<p>Something another competitor suggested that seems to work okay is to predict what each image is (power drill, purse, car, etcetera).  The higher the maximum probability would be used as a proxy for how reliable each image is.  I bet you could train your own image classifier on something more specific to this competition, but there's that strategy for now; it seems to be working for me thus far.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "322955": "Hey all,\n\nI'm fairly new to DL so forgive me if i ask dumb questions lol\n\nMy goal is to create model that can take both text and images, but i'm a little stuck as to how to go about building the model.\n\nI've trained a model using the text features, but I want to include images.  I thought of two ways I could do this...\n1. Use a pre-trained model(without the fully connected layers) to extract features from the images.  Then use those features as an input to your model along with the text features, concatenate them, then have a few fully connected layers on top with the last layer having a single output (predicted deal probability). \n\n2. Use a two input layer that takes in images and text. Use the convolution layers of pre-trained model(such as VGG16/19) for the image input , and a few fully connected layers to the text input. Concatenate these features and add a few fully connected layers on top with the last layer having a single output. \n\nAny thoughts? Hopefully what I said makes sense if not please let me know. \nThanks!",
    "324003": "Something another competitor suggested that seems to work okay is to predict what each image is (power drill, purse, car, etcetera).  The higher the maximum probability would be used as a proxy for how reliable each image is.  I bet you could train your own image classifier on something more specific to this competition, but there's that strategy for now; it seems to be working for me thus far."
  },
  "source": "meta"
}