{
  "id": 59917,
  "title": "share my NN solution.",
  "url": "/competitions/avito-demand-prediction/discussion/59917",
  "author_name": "Liu Jilong",
  "post_date": "2018-06-28T09:31:35.147000",
  "votes": 56,
  "comment_count": 13,
  "views": 0,
  "content": "<p>In the best single model thread, I shared some result of my NN model. The final LB is around 0.2180. It is not superior as the #1 team's model, but I'm glad to share my solution and wish it would help anyone.</p>\n\n<p>Here is my NN architecture:\n<img src=\"http://tuchang-1253208593.costj.myqcloud.com/NN%20architecture%20%281%29.png\" alt=\"NN architecture\"></p>\n\n<p>Some notes in the architecture:</p>\n\n<ol>\n<li>No fancy part at all, text sub model meanly comes from Toxic comment competitions, image sub model is serval Conv2d layers.</li>\n<li><code>user_id</code> is removed from category feature, due to too much cardinality</li>\n<li>It's especially prone to overfit, so i add BN before nearly every Dense.</li>\n<li>I tried public fasttext and custom trained fasttext embedding, both not work well. cbow word2vec helps a lot</li>\n<li>CNN and RNN both help text model. Attention helps.</li>\n</ol>\n\n<p>I use this main architecture without image sub-model  from the first day. The boosting between model with image and no image is about <code>0.0005</code>. Other boosting come from small changes in architecture, like add one more FC layer, change the dropout rate, use large <code>rnn_size</code> and <code>cnn_filters</code>.</p>\n\n<p>Something tried but not work:</p>\n\n<ol>\n<li>average of head K word's embedding. This should help by intuition, but it turns not work.</li>\n<li>CapsNet, CapsNet is slow to train and the result not show a boost.</li>\n<li>SkipConnection, I tried skip all continous feature to the mean Concatenate layer, but it not improve the performance, maybe need more practice. </li>\n</ol>\n\n<p>How I process the image pixels:\nThe competition provides 1.5M images, this makes long image loading time and high memory pressure. I tried keras <code>fit_generator</code> but <code>fit_generator</code> works worse then <code>fit</code>. </p>\n\n<p>My approach is sample: pre-load the image and resize to shape <code>64*64</code>, and save all the resized image in a numpy array. the shape of numpy array should be <code>num_image*64*64*3</code>. What's more: set <code>dtype=uint8</code> is key to save memory. Then the total memory of the train image should be around 20G. For me it's suitable for train.</p>\n\n<p>Thanks for all my teammates.</p>\n\n<p>I am glad for the first place except gold teams.</p>",
  "messages": [
    {
      "id": 349569,
      "postDate": "2018-06-28T09:31:35.147Z",
      "content": "<p>In the best single model thread, I shared some result of my NN model. The final LB is around 0.2180. It is not superior as the #1 team's model, but I'm glad to share my solution and wish it would help anyone.</p>\n\n<p>Here is my NN architecture:\n<img src=\"http://tuchang-1253208593.costj.myqcloud.com/NN%20architecture%20%281%29.png\" alt=\"NN architecture\"></p>\n\n<p>Some notes in the architecture:</p>\n\n<ol>\n<li>No fancy part at all, text sub model meanly comes from Toxic comment competitions, image sub model is serval Conv2d layers.</li>\n<li><code>user_id</code> is removed from category feature, due to too much cardinality</li>\n<li>It's especially prone to overfit, so i add BN before nearly every Dense.</li>\n<li>I tried public fasttext and custom trained fasttext embedding, both not work well. cbow word2vec helps a lot</li>\n<li>CNN and RNN both help text model. Attention helps.</li>\n</ol>\n\n<p>I use this main architecture without image sub-model  from the first day. The boosting between model with image and no image is about <code>0.0005</code>. Other boosting come from small changes in architecture, like add one more FC layer, change the dropout rate, use large <code>rnn_size</code> and <code>cnn_filters</code>.</p>\n\n<p>Something tried but not work:</p>\n\n<ol>\n<li>average of head K word's embedding. This should help by intuition, but it turns not work.</li>\n<li>CapsNet, CapsNet is slow to train and the result not show a boost.</li>\n<li>SkipConnection, I tried skip all continous feature to the mean Concatenate layer, but it not improve the performance, maybe need more practice. </li>\n</ol>\n\n<p>How I process the image pixels:\nThe competition provides 1.5M images, this makes long image loading time and high memory pressure. I tried keras <code>fit_generator</code> but <code>fit_generator</code> works worse then <code>fit</code>. </p>\n\n<p>My approach is sample: pre-load the image and resize to shape <code>64*64</code>, and save all the resized image in a numpy array. the shape of numpy array should be <code>num_image*64*64*3</code>. What's more: set <code>dtype=uint8</code> is key to save memory. Then the total memory of the train image should be around 20G. For me it's suitable for train.</p>\n\n<p>Thanks for all my teammates.</p>\n\n<p>I am glad for the first place except gold teams.</p>",
      "rawMarkdown": "In the best single model thread, I shared some result of my NN model. The final LB is around 0.2180. It is not superior as the #1 team's model, but I'm glad to share my solution and wish it would help anyone.\n\nHere is my NN architecture:\n![NN architecture][1]\n\n\n  [1]: http://tuchang-1253208593.costj.myqcloud.com/NN%20architecture%20%281%29.png\n\n\nSome notes in the architecture:\n\n1. No fancy part at all, text sub model meanly comes from Toxic comment competitions, image sub model is serval Conv2d layers.\n1. `user_id` is removed from category feature, due to too much cardinality\n2. It's especially prone to overfit, so i add BN before nearly every Dense.\n3. I tried public fasttext and custom trained fasttext embedding, both not work well. cbow word2vec helps a lot\n4. CNN and RNN both help text model. Attention helps.\n\nI use this main architecture without image sub-model  from the first day. The boosting between model with image and no image is about `0.0005`. Other boosting come from small changes in architecture, like add one more FC layer, change the dropout rate, use large `rnn_size` and `cnn_filters`.\n\n\nSomething tried but not work:\n\n1. average of head K word's embedding. This should help by intuition, but it turns not work.\n2. CapsNet, CapsNet is slow to train and the result not show a boost.\n3. SkipConnection, I tried skip all continous feature to the mean Concatenate layer, but it not improve the performance, maybe need more practice. \n\n\nHow I process the image pixels:\nThe competition provides 1.5M images, this makes long image loading time and high memory pressure. I tried keras `fit_generator` but `fit_generator` works worse then `fit`. \n\nMy approach is sample: pre-load the image and resize to shape `64*64`, and save all the resized image in a numpy array. the shape of numpy array should be `num_image*64*64*3`. What's more: set `dtype=uint8` is key to save memory. Then the total memory of the train image should be around 20G. For me it's suitable for train.\n\n\nThanks for all my teammates.\n\nI am glad for the first place except gold teams.",
      "votes": 56
    },
    {
      "id": 350103,
      "postDate": "2018-06-29T07:17:23.643Z",
      "content": "<p>Would you like to  share some cool code  about this nn? on github etc....</p>",
      "rawMarkdown": "Would you like to  share some cool code  about this nn? on github etc....",
      "votes": 1,
      "replies": [
        {
          "id": 447337,
          "postDate": "2018-12-29T17:13:05.307Z",
          "content": "<p>FYI the code for this NN is here: <a href=\"https://github.com/peterhurford/kaggle-avito_demand/blob/master/model_liu_nn.py\">https://github.com/peterhurford/kaggle-avito_demand/blob/master/model_liu_nn.py</a></p>",
          "rawMarkdown": "FYI the code for this NN is here: https://github.com/peterhurford/kaggle-avito_demand/blob/master/model_liu_nn.py",
          "votes": 2
        },
        {
          "id": 454702,
          "postDate": "2019-01-12T03:36:27.557Z",
          "content": "<p>thanks a lot, follow you~</p>",
          "rawMarkdown": "thanks a lot, follow you~"
        }
      ]
    },
    {
      "id": 349642,
      "postDate": "2018-06-28T12:31:00.337Z",
      "content": "<p>cool, I like it.</p>",
      "rawMarkdown": "cool, I like it.",
      "votes": 1
    },
    {
      "id": 349669,
      "postDate": "2018-06-28T13:14:45.817Z",
      "content": "<p>Great ! I'll try to replicate this and see if I get improvements. Would you be willing to share the Keras code or the hyperparameters used like number of units, dropout values, lr, epochs etc ? </p>\n\n<p>What was in your continuous variables ?</p>",
      "rawMarkdown": "Great ! I'll try to replicate this and see if I get improvements. Would you be willing to share the Keras code or the hyperparameters used like number of units, dropout values, lr, epochs etc ? \n\nWhat was in your continuous variables ?",
      "votes": 2
    },
    {
      "id": 350087,
      "postDate": "2018-06-29T06:37:12.583Z",
      "content": "<p>Nice work！666~</p>",
      "rawMarkdown": "Nice work！666~"
    },
    {
      "id": 350033,
      "postDate": "2018-06-29T03:38:13.757Z",
      "content": "<p>cool，can I ask where can I learn such DL knowledge？</p>",
      "rawMarkdown": "cool，can I ask where can I learn such DL knowledge？"
    },
    {
      "id": 349980,
      "postDate": "2018-06-29T01:25:58.213Z",
      "content": "<p>Nice work!! Jilong </p>",
      "rawMarkdown": "Nice work!! Jilong "
    },
    {
      "id": 349958,
      "postDate": "2018-06-29T00:16:40.273Z",
      "content": "<p>Congratulations @Liu Jilong for a very strong finish. Thanks for sharing your interesting NN architecture. What is your HW setup and how long did it take to run this model?</p>",
      "rawMarkdown": "Congratulations @Liu Jilong for a very strong finish. Thanks for sharing your interesting NN architecture. What is your HW setup and how long did it take to run this model?"
    },
    {
      "id": 349762,
      "postDate": "2018-06-28T15:55:07.107Z",
      "content": "<p>Great solution! What is the private LB score for this model?</p>",
      "rawMarkdown": "Great solution! What is the private LB score for this model?",
      "replies": [
        {
          "id": 349786,
          "postDate": "2018-06-28T16:37:26.657Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 349721,
      "postDate": "2018-06-28T14:35:40.240Z",
      "content": "<p>Great solution for resizing images! Using fit_generator to train CNN wasted me too much time...</p>",
      "rawMarkdown": "Great solution for resizing images! Using fit_generator to train CNN wasted me too much time..."
    },
    {
      "id": 349967,
      "postDate": "2018-06-29T00:48:19.403Z",
      "content": "<p>Thanks for sharing. </p>",
      "rawMarkdown": "Thanks for sharing. "
    }
  ],
  "comments": [
    {
      "id": 350103,
      "author_name": "high&mean",
      "author_url": "",
      "post_date": "2018-06-29T07:17:23.643000",
      "content": "<p>Would you like to  share some cool code  about this nn? on github etc....</p>",
      "votes": 1,
      "replies": [
        {
          "id": 447337,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-12-29T17:13:05.307000",
          "content": "<p>FYI the code for this NN is here: <a href=\"https://github.com/peterhurford/kaggle-avito_demand/blob/master/model_liu_nn.py\">https://github.com/peterhurford/kaggle-avito_demand/blob/master/model_liu_nn.py</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 454702,
          "author_name": "high&mean",
          "author_url": "",
          "post_date": "2019-01-12T03:36:27.557000",
          "content": "<p>thanks a lot, follow you~</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349642,
      "author_name": "Mengfei Li",
      "author_url": "",
      "post_date": "2018-06-28T12:31:00.337000",
      "content": "<p>cool, I like it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349669,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2018-06-28T13:14:45.817000",
      "content": "<p>Great ! I'll try to replicate this and see if I get improvements. Would you be willing to share the Keras code or the hyperparameters used like number of units, dropout values, lr, epochs etc ? </p>\n\n<p>What was in your continuous variables ?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 350087,
      "author_name": "uuulearn",
      "author_url": "",
      "post_date": "2018-06-29T06:37:12.583000",
      "content": "<p>Nice work！666~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 350033,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-06-29T03:38:13.757000",
      "content": "<p>cool，can I ask where can I learn such DL knowledge？</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349980,
      "author_name": "YoungLamb",
      "author_url": "",
      "post_date": "2018-06-29T01:25:58.213000",
      "content": "<p>Nice work!! Jilong </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349958,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-06-29T00:16:40.273000",
      "content": "<p>Congratulations @Liu Jilong for a very strong finish. Thanks for sharing your interesting NN architecture. What is your HW setup and how long did it take to run this model?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349762,
      "author_name": "Master",
      "author_url": "",
      "post_date": "2018-06-28T15:55:07.107000",
      "content": "<p>Great solution! What is the private LB score for this model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 349786,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-28T16:37:26.657000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349721,
      "author_name": "Webber",
      "author_url": "",
      "post_date": "2018-06-28T14:35:40.240000",
      "content": "<p>Great solution for resizing images! Using fit_generator to train CNN wasted me too much time...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349967,
      "author_name": "huiqin",
      "author_url": "",
      "post_date": "2018-06-29T00:48:19.403000",
      "content": "<p>Thanks for sharing. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349569": "In the best single model thread, I shared some result of my NN model. The final LB is around 0.2180. It is not superior as the #1 team's model, but I'm glad to share my solution and wish it would help anyone.\n\nHere is my NN architecture:\n![NN architecture][1]\n\n\n  [1]: http://tuchang-1253208593.costj.myqcloud.com/NN%20architecture%20%281%29.png\n\n\nSome notes in the architecture:\n\n1. No fancy part at all, text sub model meanly comes from Toxic comment competitions, image sub model is serval Conv2d layers.\n1. `user_id` is removed from category feature, due to too much cardinality\n2. It's especially prone to overfit, so i add BN before nearly every Dense.\n3. I tried public fasttext and custom trained fasttext embedding, both not work well. cbow word2vec helps a lot\n4. CNN and RNN both help text model. Attention helps.\n\nI use this main architecture without image sub-model  from the first day. The boosting between model with image and no image is about `0.0005`. Other boosting come from small changes in architecture, like add one more FC layer, change the dropout rate, use large `rnn_size` and `cnn_filters`.\n\n\nSomething tried but not work:\n\n1. average of head K word's embedding. This should help by intuition, but it turns not work.\n2. CapsNet, CapsNet is slow to train and the result not show a boost.\n3. SkipConnection, I tried skip all continous feature to the mean Concatenate layer, but it not improve the performance, maybe need more practice. \n\n\nHow I process the image pixels:\nThe competition provides 1.5M images, this makes long image loading time and high memory pressure. I tried keras `fit_generator` but `fit_generator` works worse then `fit`. \n\nMy approach is sample: pre-load the image and resize to shape `64*64`, and save all the resized image in a numpy array. the shape of numpy array should be `num_image*64*64*3`. What's more: set `dtype=uint8` is key to save memory. Then the total memory of the train image should be around 20G. For me it's suitable for train.\n\n\nThanks for all my teammates.\n\nI am glad for the first place except gold teams.",
    "350103": "Would you like to  share some cool code  about this nn? on github etc....",
    "349642": "cool, I like it.",
    "349669": "Great ! I'll try to replicate this and see if I get improvements. Would you be willing to share the Keras code or the hyperparameters used like number of units, dropout values, lr, epochs etc ? \n\nWhat was in your continuous variables ?",
    "350087": "Nice work！666~",
    "350033": "cool，can I ask where can I learn such DL knowledge？",
    "349980": "Nice work!! Jilong ",
    "349958": "Congratulations @Liu Jilong for a very strong finish. Thanks for sharing your interesting NN architecture. What is your HW setup and how long did it take to run this model?",
    "349762": "Great solution! What is the private LB score for this model?",
    "349721": "Great solution for resizing images! Using fit_generator to train CNN wasted me too much time...",
    "349967": "Thanks for sharing. "
  }
}