{
  "id": 57293,
  "title": "NN Embeddings - What Works and What Doesn't?",
  "url": "/competitions/avito-demand-prediction/discussion/57293",
  "author_name": "",
  "post_date": "2018-05-22T05:06:33.041603100Z",
  "votes": 15,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I am listing down somethings that worked for me in my RNN network. </p>\n\n<p>Approach for Word Processing: </p>\n\n<ul>\n<li>Single Text vector combining title and description </li>\n<li>Zero word processing </li>\n<li>Vocab size is fairly large at 200K </li>\n</ul>\n\n<p>Categorical Variables </p>\n\n<ul>\n<li><p>Label Encoding on Train+ Test ( I ended up combining train and test since I was getting errors \nduring the prediction phase that said new labels found )</p></li>\n<li><p>Log or Log1p conversions for numerical values such as price </p></li>\n</ul>\n\n<p>Embedding Approaches tried so far - </p>\n\n<ol>\n<li><p>Self Trained Word2Vec Embeddings -  300D\n Embedding and CV score improves with higher dimensions and more epochs used for training </p></li>\n<li><p>Architecture - Adding more dense layers is dropping Local CV. </p></li>\n<li><p>FastText for Russian  - Best Score So far for the RNN network I used.  300D</p></li>\n<li><p>Russian Glove  - Lowest score.  A very large portion of the vocab was found missing.  I am guessing that weakened the model since I was simply replacing missing words with a zeros (300D)</p></li>\n<li><p>Concatenation of FastText and Glove - 600 Vector embedding. \n  Dropped the score achieved by Fast Text. Clearly Glove doesn't work for whatever reason. </p></li>\n<li><p>Concatenation of FastText and Self Trained Word2vec </p>\n\n<p>I think there is a bit of a slight bump is CV with this.  In the process of generating 300D Vectors \n for Word2vec. </p></li>\n</ol>\n\n<p>Will update as and when I try new things. </p>\n\n<p>Regards\nShanth</p>",
  "messages": [
    {
      "id": "331911",
      "postDate": "05/22/2018 05:06:33",
      "content": "<p>I am listing down somethings that worked for me in my RNN network. </p>\n\n<p>Approach for Word Processing: </p>\n\n<ul>\n<li>Single Text vector combining title and description </li>\n<li>Zero word processing </li>\n<li>Vocab size is fairly large at 200K </li>\n</ul>\n\n<p>Categorical Variables </p>\n\n<ul>\n<li><p>Label Encoding on Train+ Test ( I ended up combining train and test since I was getting errors \nduring the prediction phase that said new labels found )</p></li>\n<li><p>Log or Log1p conversions for numerical values such as price </p></li>\n</ul>\n\n<p>Embedding Approaches tried so far - </p>\n\n<ol>\n<li><p>Self Trained Word2Vec Embeddings -  300D\n Embedding and CV score improves with higher dimensions and more epochs used for training </p></li>\n<li><p>Architecture - Adding more dense layers is dropping Local CV. </p></li>\n<li><p>FastText for Russian  - Best Score So far for the RNN network I used.  300D</p></li>\n<li><p>Russian Glove  - Lowest score.  A very large portion of the vocab was found missing.  I am guessing that weakened the model since I was simply replacing missing words with a zeros (300D)</p></li>\n<li><p>Concatenation of FastText and Glove - 600 Vector embedding. \n  Dropped the score achieved by Fast Text. Clearly Glove doesn't work for whatever reason. </p></li>\n<li><p>Concatenation of FastText and Self Trained Word2vec </p>\n\n<p>I think there is a bit of a slight bump is CV with this.  In the process of generating 300D Vectors \n for Word2vec. </p></li>\n</ol>\n\n<p>Will update as and when I try new things. </p>\n\n<p>Regards\nShanth</p>",
      "rawMarkdown": "I am listing down somethings that worked for me in my RNN network. \n\nApproach for Word Processing: \n\n- Single Text vector combining title and description \n- Zero word processing \n- Vocab size is fairly large at 200K \n\nCategorical Variables \n\n- Label Encoding on Train+ Test ( I ended up combining train and test since I was getting errors \n  during the prediction phase that said new labels found )\n\n- Log or Log1p conversions for numerical values such as price \n\nEmbedding Approaches tried so far - \n\n1.  Self Trained Word2Vec Embeddings -  300D\n     Embedding and CV score improves with higher dimensions and more epochs used for training \n\n2. Architecture - Adding more dense layers is dropping Local CV. \n\n3. FastText for Russian  - Best Score So far for the RNN network I used.  300D\n\n4. Russian Glove  - Lowest score.  A very large portion of the vocab was found missing.  I am guessing that weakened the model since I was simply replacing missing words with a zeros (300D)\n\n5.  Concatenation of FastText and Glove - 600 Vector embedding. \n      Dropped the score achieved by Fast Text. Clearly Glove doesn't work for whatever reason. \n\n6. Concatenation of FastText and Self Trained Word2vec \n\n    I think there is a bit of a slight bump is CV with this.  In the process of generating 300D Vectors \n     for Word2vec. \n\nWill update as and when I try new things. \n\nRegards\nShanth",
      "votes": null
    },
    {
      "id": "331977",
      "postDate": "05/22/2018 09:41:31",
      "content": "<p>Thanks shanth for the summary </p>\n\n<blockquote>\n  <p>Single Text vector combining title and description</p>\n</blockquote>\n\n<p>Did it improve your CV ( instead of using 2 text vectors separately ) ? </p>\n\n<blockquote>\n  <p>Self Trained Word2Vec Embeddings - 300D Embedding and CV score improves with higher dimensions and more epochs used for training ...</p>\n</blockquote>\n\n<p>I've just tried 100D Self Trained . May be I will try higher dimension.  However it gave me better CV and LB than Fasttext ( the one trained on wiki)</p>\n\n<blockquote>\n  <p>Concatenation of FastText and Glove - 600 Vector embedding. Dropped the score achieved by Fast Text. Clearly Glove doesn't work for whatever reason.</p>\n</blockquote>\n\n<p>I was using multi-embedding (Glove + fasttext) on Toxic competition but it was very time consumming for the training  ( I used 2M rows though). Given the bigger dataset here. I didn't try it.  Good to know it didn't improve. </p>",
      "rawMarkdown": "Thanks shanth for the summary \n\n&gt; Single Text vector combining title and description\n\nDid it improve your CV ( instead of using 2 text vectors separately ) ? \n\n&gt; Self Trained Word2Vec Embeddings - 300D Embedding and CV score improves with higher dimensions and more epochs used for training ...\n\nI've just tried 100D Self Trained . May be I will try higher dimension.  However it gave me better CV and LB than Fasttext ( the one trained on wiki)\n\n\n&gt;Concatenation of FastText and Glove - 600 Vector embedding. Dropped the score achieved by Fast Text. Clearly Glove doesn't work for whatever reason.\n\nI was using multi-embedding (Glove + fasttext) on Toxic competition but it was very time consumming for the training  ( I used 2M rows though). Given the bigger dataset here. I didn't try it.  Good to know it didn't improve.",
      "votes": null
    },
    {
      "id": "331989",
      "postDate": "05/22/2018 10:02:30",
      "content": "<p>@ Seringe </p>\n\n<p>I never tried individual vectors. I have been doing title+description all along. \n100D self trained actually reduced the CV for me significantly. Mebbe I have made some mistakes. Will go back and check. Thanks for the udpate. </p>\n\n<p>Also, for your embedding did you use all of train_active and test_active?</p>",
      "rawMarkdown": "Seringe \n\nI never tried individual vectors. I have been doing title+description all along. \n100D self trained actually reduced the CV for me significantly. Mebbe I have made some mistakes. Will go back and check. Thanks for the udpate. \n\nAlso, for your embedding did you use all of train_active and test_active?",
      "votes": null
    },
    {
      "id": "332145",
      "postDate": "05/22/2018 16:13:12",
      "content": "<p>Just a word of caution on your point, \"( I ended up combining train and test since I was getting errors during the prediction phase that said new labels found )\". Since there are labels unique to the test set, those values should not be embedded they should be deleted and replaced with the \"missing\" value. If you embed them, they will be initialized but never trained where as the \"missing\" embedding will be trained. In fact you should treat the validation set the same as test in this respect for the same reason. </p>",
      "rawMarkdown": "Just a word of caution on your point, \"( I ended up combining train and test since I was getting errors during the prediction phase that said new labels found )\". Since there are labels unique to the test set, those values should not be embedded they should be deleted and replaced with the \"missing\" value. If you embed them, they will be initialized but never trained where as the \"missing\" embedding will be trained. In fact you should treat the validation set the same as test in this respect for the same reason.",
      "votes": null
    },
    {
      "id": "332161",
      "postDate": "05/22/2018 16:28:15",
      "content": "<p>@ Michael \nThanks for that note. Will look into it. </p>",
      "rawMarkdown": "Michael \nThanks for that note. Will look into it.",
      "votes": null
    },
    {
      "id": "332256",
      "postDate": "05/22/2018 21:32:19",
      "content": "<p>I also get significantly better results with self trained 100D vs Russian fastText 300D (note some of the text we have is Ukrainian I think...). Did someone try Cbow vs SkipGram ?</p>\n\n<p>I'll train 300D during the night first.</p>",
      "rawMarkdown": "I also get significantly better results with self trained 100D vs Russian fastText 300D (note some of the text we have is Ukrainian I think...). Did someone try Cbow vs SkipGram ?\n\nI'll train 300D during the night first.",
      "votes": null
    },
    {
      "id": "332717",
      "postDate": "05/23/2018 16:01:29",
      "content": "<p>@ Seringe </p>\n\n<p>Ran a kernel for title and description as separate vectors and it converged better. But it take twice as much time though. </p>\n\n<p>My local CV had a boost of about 0.0005. Yet to submit. </p>\n\n<p>Regards\nShanth</p>",
      "rawMarkdown": "Seringe \n\nRan a kernel for title and description as separate vectors and it converged better. But it take twice as much time though. \n\nMy local CV had a boost of about 0.0005. Yet to submit. \n\nRegards\nShanth",
      "votes": null
    },
    {
      "id": "333091",
      "postDate": "05/24/2018 11:35:18",
      "content": "<p>Self-trained embeddings with the dimension of 64 worked fine for me. I tried to use frozen fasttext vectors pretrained on wikipedia, but they seem to have quite a small overlap with the vocabulary from avito dataset, so the performance was not very good.\nNow I'm thinking about simultaneously using two sets of embeddings: pretrained fasttext vectors on known words and self-trained embeddings on unknown words. </p>\n\n<p>Also I have found params fields and the title field to be complementary to each other. Many ads miss some important information in the title field, while it can be found in the params. Say, a title may not contain the word \"автомобиль\" (a car) which seems to be extremely important for the prediction, while param_1 can represent the whole category (cars). So I tried to process these fields together (just merged all into 1 text field) and it improved my score.</p>",
      "rawMarkdown": "Self-trained embeddings with the dimension of 64 worked fine for me. I tried to use frozen fasttext vectors pretrained on wikipedia, but they seem to have quite a small overlap with the vocabulary from avito dataset, so the performance was not very good.\nNow I'm thinking about simultaneously using two sets of embeddings: pretrained fasttext vectors on known words and self-trained embeddings on unknown words. \n\nAlso I have found params fields and the title field to be complementary to each other. Many ads miss some important information in the title field, while it can be found in the params. Say, a title may not contain the word \"автомобиль\" (a car) which seems to be extremely important for the prediction, while param_1 can represent the whole category (cars). So I tried to process these fields together (just merged all into 1 text field) and it improved my score.",
      "votes": null
    },
    {
      "id": "333098",
      "postDate": "05/24/2018 11:53:22",
      "content": "<p>@ Dimitriy.</p>\n\n<p>Joining title and param1 seems very intuitive. Good one. I will try it myself.\nThe other idea of using two vectors also sounds interesting. </p>\n\n<p>Regards\nShanth</p>",
      "rawMarkdown": "Dimitriy.\n\nJoining title and param1 seems very intuitive. Good one. I will try it myself.\nThe other idea of using two vectors also sounds interesting. \n\nRegards\nShanth",
      "votes": null
    },
    {
      "id": "333200",
      "postDate": "05/24/2018 16:13:51",
      "content": "<p>I also trained some embeddings using image_top_1 as target. You can check an example in this <a href=\"https://www.kaggle.com/christofhenkel/text2image-top-1\">kernel</a>.</p>",
      "rawMarkdown": "I also trained some embeddings using image_top_1 as target. You can check an example in this [kernel][1].\n\n\n  [1]: https://www.kaggle.com/christofhenkel/text2image-top-1",
      "votes": null
    },
    {
      "id": "333216",
      "postDate": "05/24/2018 17:16:55",
      "content": "<p>@ Dieter </p>\n\n<p>That's a great way to approach it. Will take a look at it. :)</p>",
      "rawMarkdown": "Dieter \n\nThat's a great way to approach it. Will take a look at it. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 331977,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/22/2018 09:41:31",
      "content": "<p>Thanks shanth for the summary </p>\n\n<blockquote>\n  <p>Single Text vector combining title and description</p>\n</blockquote>\n\n<p>Did it improve your CV ( instead of using 2 text vectors separately ) ? </p>\n\n<blockquote>\n  <p>Self Trained Word2Vec Embeddings - 300D Embedding and CV score improves with higher dimensions and more epochs used for training ...</p>\n</blockquote>\n\n<p>I've just tried 100D Self Trained . May be I will try higher dimension.  However it gave me better CV and LB than Fasttext ( the one trained on wiki)</p>\n\n<blockquote>\n  <p>Concatenation of FastText and Glove - 600 Vector embedding. Dropped the score achieved by Fast Text. Clearly Glove doesn't work for whatever reason.</p>\n</blockquote>\n\n<p>I was using multi-embedding (Glove + fasttext) on Toxic competition but it was very time consumming for the training  ( I used 2M rows though). Given the bigger dataset here. I didn't try it.  Good to know it didn't improve. </p>",
      "votes": null,
      "replies": [
        {
          "id": 331989,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "05/22/2018 10:02:30",
          "content": "<p>@ Seringe </p>\n\n<p>I never tried individual vectors. I have been doing title+description all along. \n100D self trained actually reduced the CV for me significantly. Mebbe I have made some mistakes. Will go back and check. Thanks for the udpate. </p>\n\n<p>Also, for your embedding did you use all of train_active and test_active?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332256,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "05/22/2018 21:32:19",
          "content": "<p>I also get significantly better results with self trained 100D vs Russian fastText 300D (note some of the text we have is Ukrainian I think...). Did someone try Cbow vs SkipGram ?</p>\n\n<p>I'll train 300D during the night first.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332717,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "05/23/2018 16:01:29",
          "content": "<p>@ Seringe </p>\n\n<p>Ran a kernel for title and description as separate vectors and it converged better. But it take twice as much time though. </p>\n\n<p>My local CV had a boost of about 0.0005. Yet to submit. </p>\n\n<p>Regards\nShanth</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332145,
      "author_name": "vannak",
      "author_url": "",
      "post_date": "05/22/2018 16:13:12",
      "content": "<p>Just a word of caution on your point, \"( I ended up combining train and test since I was getting errors during the prediction phase that said new labels found )\". Since there are labels unique to the test set, those values should not be embedded they should be deleted and replaced with the \"missing\" value. If you embed them, they will be initialized but never trained where as the \"missing\" embedding will be trained. In fact you should treat the validation set the same as test in this respect for the same reason. </p>",
      "votes": null,
      "replies": [
        {
          "id": 332161,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "05/22/2018 16:28:15",
          "content": "<p>@ Michael \nThanks for that note. Will look into it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 333091,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "05/24/2018 11:35:18",
      "content": "<p>Self-trained embeddings with the dimension of 64 worked fine for me. I tried to use frozen fasttext vectors pretrained on wikipedia, but they seem to have quite a small overlap with the vocabulary from avito dataset, so the performance was not very good.\nNow I'm thinking about simultaneously using two sets of embeddings: pretrained fasttext vectors on known words and self-trained embeddings on unknown words. </p>\n\n<p>Also I have found params fields and the title field to be complementary to each other. Many ads miss some important information in the title field, while it can be found in the params. Say, a title may not contain the word \"автомобиль\" (a car) which seems to be extremely important for the prediction, while param_1 can represent the whole category (cars). So I tried to process these fields together (just merged all into 1 text field) and it improved my score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 333098,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "05/24/2018 11:53:22",
          "content": "<p>@ Dimitriy.</p>\n\n<p>Joining title and param1 seems very intuitive. Good one. I will try it myself.\nThe other idea of using two vectors also sounds interesting. </p>\n\n<p>Regards\nShanth</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 333200,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "05/24/2018 16:13:51",
      "content": "<p>I also trained some embeddings using image_top_1 as target. You can check an example in this <a href=\"https://www.kaggle.com/christofhenkel/text2image-top-1\">kernel</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 333216,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "05/24/2018 17:16:55",
          "content": "<p>@ Dieter </p>\n\n<p>That's a great way to approach it. Will take a look at it. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "331911": "I am listing down somethings that worked for me in my RNN network. \n\nApproach for Word Processing: \n\n- Single Text vector combining title and description \n- Zero word processing \n- Vocab size is fairly large at 200K \n\nCategorical Variables \n\n- Label Encoding on Train+ Test ( I ended up combining train and test since I was getting errors \n  during the prediction phase that said new labels found )\n\n- Log or Log1p conversions for numerical values such as price \n\nEmbedding Approaches tried so far - \n\n1.  Self Trained Word2Vec Embeddings -  300D\n     Embedding and CV score improves with higher dimensions and more epochs used for training \n\n2. Architecture - Adding more dense layers is dropping Local CV. \n\n3. FastText for Russian  - Best Score So far for the RNN network I used.  300D\n\n4. Russian Glove  - Lowest score.  A very large portion of the vocab was found missing.  I am guessing that weakened the model since I was simply replacing missing words with a zeros (300D)\n\n5.  Concatenation of FastText and Glove - 600 Vector embedding. \n      Dropped the score achieved by Fast Text. Clearly Glove doesn't work for whatever reason. \n\n6. Concatenation of FastText and Self Trained Word2vec \n\n    I think there is a bit of a slight bump is CV with this.  In the process of generating 300D Vectors \n     for Word2vec. \n\nWill update as and when I try new things. \n\nRegards\nShanth",
    "331977": "Thanks shanth for the summary \n\n&gt; Single Text vector combining title and description\n\nDid it improve your CV ( instead of using 2 text vectors separately ) ? \n\n&gt; Self Trained Word2Vec Embeddings - 300D Embedding and CV score improves with higher dimensions and more epochs used for training ...\n\nI've just tried 100D Self Trained . May be I will try higher dimension.  However it gave me better CV and LB than Fasttext ( the one trained on wiki)\n\n\n&gt;Concatenation of FastText and Glove - 600 Vector embedding. Dropped the score achieved by Fast Text. Clearly Glove doesn't work for whatever reason.\n\nI was using multi-embedding (Glove + fasttext) on Toxic competition but it was very time consumming for the training  ( I used 2M rows though). Given the bigger dataset here. I didn't try it.  Good to know it didn't improve.",
    "331989": "Seringe \n\nI never tried individual vectors. I have been doing title+description all along. \n100D self trained actually reduced the CV for me significantly. Mebbe I have made some mistakes. Will go back and check. Thanks for the udpate. \n\nAlso, for your embedding did you use all of train_active and test_active?",
    "332145": "Just a word of caution on your point, \"( I ended up combining train and test since I was getting errors during the prediction phase that said new labels found )\". Since there are labels unique to the test set, those values should not be embedded they should be deleted and replaced with the \"missing\" value. If you embed them, they will be initialized but never trained where as the \"missing\" embedding will be trained. In fact you should treat the validation set the same as test in this respect for the same reason.",
    "332161": "Michael \nThanks for that note. Will look into it.",
    "332256": "I also get significantly better results with self trained 100D vs Russian fastText 300D (note some of the text we have is Ukrainian I think...). Did someone try Cbow vs SkipGram ?\n\nI'll train 300D during the night first.",
    "332717": "Seringe \n\nRan a kernel for title and description as separate vectors and it converged better. But it take twice as much time though. \n\nMy local CV had a boost of about 0.0005. Yet to submit. \n\nRegards\nShanth",
    "333091": "Self-trained embeddings with the dimension of 64 worked fine for me. I tried to use frozen fasttext vectors pretrained on wikipedia, but they seem to have quite a small overlap with the vocabulary from avito dataset, so the performance was not very good.\nNow I'm thinking about simultaneously using two sets of embeddings: pretrained fasttext vectors on known words and self-trained embeddings on unknown words. \n\nAlso I have found params fields and the title field to be complementary to each other. Many ads miss some important information in the title field, while it can be found in the params. Say, a title may not contain the word \"автомобиль\" (a car) which seems to be extremely important for the prediction, while param_1 can represent the whole category (cars). So I tried to process these fields together (just merged all into 1 text field) and it improved my score.",
    "333098": "Dimitriy.\n\nJoining title and param1 seems very intuitive. Good one. I will try it myself.\nThe other idea of using two vectors also sounds interesting. \n\nRegards\nShanth",
    "333200": "I also trained some embeddings using image_top_1 as target. You can check an example in this [kernel][1].\n\n\n  [1]: https://www.kaggle.com/christofhenkel/text2image-top-1",
    "333216": "Dieter \n\nThat's a great way to approach it. Will take a look at it. :)"
  },
  "source": "meta"
}