{
  "id": 56695,
  "title": "Using train_active to train text features",
  "url": "/competitions/avito-demand-prediction/discussion/56695",
  "author_name": "",
  "post_date": "2018-05-13T18:13:06.593628200Z",
  "votes": 24,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I added a kernel for showing <a href=\"https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\">how to train a Word2Vec model</a>  from train_active.csv and another one for showing how to <a href=\"https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\">use a self-trained model</a>  and to compare with using <a href=\"https://www.kaggle.com/christofhenkel/fasttext-starter-description-only\">pre-trained Fasttext embeddings</a></p>\n\n<p>Bottom line is that self-trained embeddings using the text in train_active.csv might significantly improve your model. </p>",
  "messages": [
    {
      "id": "328226",
      "postDate": "05/13/2018 18:13:06",
      "content": "<p>I added a kernel for showing <a href=\"https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\">how to train a Word2Vec model</a>  from train_active.csv and another one for showing how to <a href=\"https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\">use a self-trained model</a>  and to compare with using <a href=\"https://www.kaggle.com/christofhenkel/fasttext-starter-description-only\">pre-trained Fasttext embeddings</a></p>\n\n<p>Bottom line is that self-trained embeddings using the text in train_active.csv might significantly improve your model. </p>",
      "rawMarkdown": "I added a kernel for showing [how to train a Word2Vec model][1]  from train_active.csv and another one for showing how to [use a self-trained model][2]  and to compare with using [pre-trained Fasttext embeddings][3]\n\nBottom line is that self-trained embeddings using the text in train_active.csv might significantly improve your model. \n\n\n  [1]: https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\n  [2]: https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\n  [3]: https://www.kaggle.com/christofhenkel/fasttext-starter-description-only",
      "votes": null
    },
    {
      "id": "328334",
      "postDate": "05/14/2018 02:17:28",
      "content": "<p>So there is one question. How to jointly train word embed with other features. It seems hard to combine NN with Lightgbm.</p>\n\n<p>Maybe one should train a base NN without embedding and plugin the embeddings.</p>",
      "rawMarkdown": "So there is one question. How to jointly train word embed with other features. It seems hard to combine NN with Lightgbm.\n\nMaybe one should train a base NN without embedding and plugin the embeddings.",
      "votes": null
    },
    {
      "id": "328406",
      "postDate": "05/14/2018 08:27:24",
      "content": "<p>Thank you very much Dieter</p>\n\n<p>Self-trained embeddings gave me some improvement on CV ( I was using Fasttext trained on Wiki before )</p>",
      "rawMarkdown": "Thank you very much Dieter\n\nSelf-trained embeddings gave me some improvement on CV ( I was using Fasttext trained on Wiki before )",
      "votes": null
    },
    {
      "id": "328568",
      "postDate": "05/14/2018 16:10:08",
      "content": "<blockquote>\n  <p>Maybe one should train a base NN without embedding and plugin the embeddings</p>\n</blockquote>\n\n<p>Sorry, I don´t get your idea. Could you explain a bit more?</p>",
      "rawMarkdown": "&gt; Maybe one should train a base NN without embedding and plugin the embeddings\n\nSorry, I don´t get your idea. Could you explain a bit more?",
      "votes": null
    },
    {
      "id": "328573",
      "postDate": "05/14/2018 16:24:48",
      "content": "<p>The previous <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262\">competition</a>, which contains no deep learning features(image, text), shows NN is comparable to LGB based model.</p>\n\n<p>So training a NN model with origin features and add the embedding at some or other NN layer.\nA lot of things to try, but a good base NN model is the first step.</p>\n\n<p>Sorry for my English.</p>",
      "rawMarkdown": "The previous [competition](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262), which contains no deep learning features(image, text), shows NN is comparable to LGB based model.\n\nSo training a NN model with origin features and add the embedding at some or other NN layer.\nA lot of things to try, but a good base NN model is the first step.\n\nSorry for my English.",
      "votes": null
    },
    {
      "id": "328902",
      "postDate": "05/15/2018 10:33:20",
      "content": "<p>They are not \"comparable\"..</p>\n\n<p>NN models are way stronger when the data include (lot of) text features and/or images ...(unless you 're ready to do lot of feature engeneering and tuning for LGB ^^) </p>",
      "rawMarkdown": "They are not \"comparable\"..\n\nNN models are way stronger when the data include (lot of) text features and/or images ...(unless you 're ready to do lot of feature engeneering and tuning for LGB ^^)",
      "votes": null
    },
    {
      "id": "328931",
      "postDate": "05/15/2018 11:52:13",
      "content": "<p>Congrats. <br>\nUnfortunately, It does not work for me :D</p>",
      "rawMarkdown": "Congrats.  \nUnfortunately, It does not work for me :D",
      "votes": null
    },
    {
      "id": "331308",
      "postDate": "05/20/2018 23:29:14",
      "content": "<p>good, thank you very much.</p>",
      "rawMarkdown": "good, thank you very much.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 328334,
      "author_name": "liujilong",
      "author_url": "",
      "post_date": "05/14/2018 02:17:28",
      "content": "<p>So there is one question. How to jointly train word embed with other features. It seems hard to combine NN with Lightgbm.</p>\n\n<p>Maybe one should train a base NN without embedding and plugin the embeddings.</p>",
      "votes": null,
      "replies": [
        {
          "id": 328568,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "05/14/2018 16:10:08",
          "content": "<blockquote>\n  <p>Maybe one should train a base NN without embedding and plugin the embeddings</p>\n</blockquote>\n\n<p>Sorry, I don´t get your idea. Could you explain a bit more?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328573,
          "author_name": "liujilong",
          "author_url": "",
          "post_date": "05/14/2018 16:24:48",
          "content": "<p>The previous <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262\">competition</a>, which contains no deep learning features(image, text), shows NN is comparable to LGB based model.</p>\n\n<p>So training a NN model with origin features and add the embedding at some or other NN layer.\nA lot of things to try, but a good base NN model is the first step.</p>\n\n<p>Sorry for my English.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328902,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/15/2018 10:33:20",
          "content": "<p>They are not \"comparable\"..</p>\n\n<p>NN models are way stronger when the data include (lot of) text features and/or images ...(unless you 're ready to do lot of feature engeneering and tuning for LGB ^^) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 328406,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/14/2018 08:27:24",
      "content": "<p>Thank you very much Dieter</p>\n\n<p>Self-trained embeddings gave me some improvement on CV ( I was using Fasttext trained on Wiki before )</p>",
      "votes": null,
      "replies": [
        {
          "id": 328931,
          "author_name": "ngxbac",
          "author_url": "",
          "post_date": "05/15/2018 11:52:13",
          "content": "<p>Congrats. <br>\nUnfortunately, It does not work for me :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331308,
      "author_name": "jetouxu",
      "author_url": "",
      "post_date": "05/20/2018 23:29:14",
      "content": "<p>good, thank you very much.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "328226": "I added a kernel for showing [how to train a Word2Vec model][1]  from train_active.csv and another one for showing how to [use a self-trained model][2]  and to compare with using [pre-trained Fasttext embeddings][3]\n\nBottom line is that self-trained embeddings using the text in train_active.csv might significantly improve your model. \n\n\n  [1]: https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\n  [2]: https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\n  [3]: https://www.kaggle.com/christofhenkel/fasttext-starter-description-only",
    "328334": "So there is one question. How to jointly train word embed with other features. It seems hard to combine NN with Lightgbm.\n\nMaybe one should train a base NN without embedding and plugin the embeddings.",
    "328406": "Thank you very much Dieter\n\nSelf-trained embeddings gave me some improvement on CV ( I was using Fasttext trained on Wiki before )",
    "328568": "&gt; Maybe one should train a base NN without embedding and plugin the embeddings\n\nSorry, I don´t get your idea. Could you explain a bit more?",
    "328573": "The previous [competition](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262), which contains no deep learning features(image, text), shows NN is comparable to LGB based model.\n\nSo training a NN model with origin features and add the embedding at some or other NN layer.\nA lot of things to try, but a good base NN model is the first step.\n\nSorry for my English.",
    "328902": "They are not \"comparable\"..\n\nNN models are way stronger when the data include (lot of) text features and/or images ...(unless you 're ready to do lot of feature engeneering and tuning for LGB ^^)",
    "328931": "Congrats.  \nUnfortunately, It does not work for me :D",
    "331308": "good, thank you very much."
  },
  "source": "meta"
}