{
  "id": 81697,
  "title": "38th solution, data driven to find embedding weights",
  "url": "/competitions/quora-insincere-questions-classification/discussion/81697",
  "author_name": "Salon_sai",
  "post_date": "2019-02-24T09:04:32.954000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Model structure</h2>\n\n<p><img src=\"https://upload-images.jianshu.io/upload_images/6673934-43708f05b93e32fe.png?imageMogr2/auto-orient/strip%7CimageView2/2/w/1240\" alt=\"\"></p>\n\n<h2>Embedding weight</h2>\n\n<p>The weight of embedding: 0.5 * glove+0.4 * para+0.1 * wiki. The weights are obtained as follows: First, the vector <code>x</code> of the three embeddings in the sentence is obtained, and the shape is (batch_size, maxlen, embed_size, 3). and then I create weight(shape: (3, )) into the softmax layer and multiply it with the tensor <code>x</code>. I divide the data into 5-folds, and each fold is averaged by the embedding weight as the final embedding weight. The way that embedding weights are obtained is data driven.</p>\n\n<h2>Word embedding</h2>\n\n<p>First I try to preserve the case of the word, and then get the word vector through the pre-training word vector. If there is no match, then the word will be: lower upper capitalize processing to find the corresponding word vector, etc. If not found, the vector is randomly generated. But the general practice is not effective. The main reason for my guess is: There are some weird spellings like: \"tHe\", \"ThE\". I look up the text and find that they originally mean \"the\". But in The glove finds the corresponding embedding vector. These words appear relatively few times, so their confidence is low, when training glove, and they mean the same. So I changed them all into Lowcase.</p>\n\n<h2>CV evaluation</h2>\n\n<p>On discussion, everyone is worried about whether or not to trust local cv. But a lot of people are just trust single model local cv, not trust ensemble model local cv.. For the former, the assessment is the single model, and the second is the ensemble model.. The ensemble, used at the end of the game should be ensemble model local cv. The methods are as follows: we divide the training set into k folders, then train all the models that need to be blended and blend the prediction results, so that we can get a prediction result of fold, and so on. In addition, I use the corr to see the similarity of fold training, I hope the model is good and different. My approach is dangerous, after all, split folds is random, the more random seeds, the less stable the final result.</p>\n\n<h2>Loss Function</h2>\n\n<p>I tried to use Focal loss as loss function, the single model local cv is very good, there is 0.712. But don't forget just said, this is only single model local cv, not ensemble model local cv. Finally, due to time, it does not do its ensemble local CV, but I think it might be better to adjust the hyperparameters than the Binary Cross Entropy.</p>\n\n<p>May be my English is poorly. So I write the blog in my <a href=\"https://zhuanlan.zhihu.com/p/57550039\">知乎</a>, the Chinese Quora. Thx everyone share solutions and kernel. I hope that in the future I will share more interesting and effective technical knowledge like everyone else.</p>",
  "messages": [
    {
      "id": 477266,
      "postDate": "2019-02-24T09:04:32.953Z",
      "content": "<h2>Model structure</h2>\n\n<p><img src=\"https://upload-images.jianshu.io/upload_images/6673934-43708f05b93e32fe.png?imageMogr2/auto-orient/strip%7CimageView2/2/w/1240\" alt=\"\"></p>\n\n<h2>Embedding weight</h2>\n\n<p>The weight of embedding: 0.5 * glove+0.4 * para+0.1 * wiki. The weights are obtained as follows: First, the vector <code>x</code> of the three embeddings in the sentence is obtained, and the shape is (batch_size, maxlen, embed_size, 3). and then I create weight(shape: (3, )) into the softmax layer and multiply it with the tensor <code>x</code>. I divide the data into 5-folds, and each fold is averaged by the embedding weight as the final embedding weight. The way that embedding weights are obtained is data driven.</p>\n\n<h2>Word embedding</h2>\n\n<p>First I try to preserve the case of the word, and then get the word vector through the pre-training word vector. If there is no match, then the word will be: lower upper capitalize processing to find the corresponding word vector, etc. If not found, the vector is randomly generated. But the general practice is not effective. The main reason for my guess is: There are some weird spellings like: \"tHe\", \"ThE\". I look up the text and find that they originally mean \"the\". But in The glove finds the corresponding embedding vector. These words appear relatively few times, so their confidence is low, when training glove, and they mean the same. So I changed them all into Lowcase.</p>\n\n<h2>CV evaluation</h2>\n\n<p>On discussion, everyone is worried about whether or not to trust local cv. But a lot of people are just trust single model local cv, not trust ensemble model local cv.. For the former, the assessment is the single model, and the second is the ensemble model.. The ensemble, used at the end of the game should be ensemble model local cv. The methods are as follows: we divide the training set into k folders, then train all the models that need to be blended and blend the prediction results, so that we can get a prediction result of fold, and so on. In addition, I use the corr to see the similarity of fold training, I hope the model is good and different. My approach is dangerous, after all, split folds is random, the more random seeds, the less stable the final result.</p>\n\n<h2>Loss Function</h2>\n\n<p>I tried to use Focal loss as loss function, the single model local cv is very good, there is 0.712. But don't forget just said, this is only single model local cv, not ensemble model local cv. Finally, due to time, it does not do its ensemble local CV, but I think it might be better to adjust the hyperparameters than the Binary Cross Entropy.</p>\n\n<p>May be my English is poorly. So I write the blog in my <a href=\"https://zhuanlan.zhihu.com/p/57550039\">知乎</a>, the Chinese Quora. Thx everyone share solutions and kernel. I hope that in the future I will share more interesting and effective technical knowledge like everyone else.</p>",
      "rawMarkdown": "## Model structure\n\n![](https://upload-images.jianshu.io/upload_images/6673934-43708f05b93e32fe.png?imageMogr2/auto-orient/strip%7CimageView2/2/w/1240)\n\n## Embedding weight\nThe weight of embedding: 0.5 * glove+0.4 * para+0.1 * wiki. The weights are obtained as follows: First, the vector `x` of the three embeddings in the sentence is obtained, and the shape is (batch_size, maxlen, embed_size, 3). and then I create weight(shape: (3, )) into the softmax layer and multiply it with the tensor `x`. I divide the data into 5-folds, and each fold is averaged by the embedding weight as the final embedding weight. The way that embedding weights are obtained is data driven.\n\n## Word embedding\nFirst I try to preserve the case of the word, and then get the word vector through the pre-training word vector. If there is no match, then the word will be: lower upper capitalize processing to find the corresponding word vector, etc. If not found, the vector is randomly generated. But the general practice is not effective. The main reason for my guess is: There are some weird spellings like: \"tHe\", \"ThE\". I look up the text and find that they originally mean \"the\". But in The glove finds the corresponding embedding vector. These words appear relatively few times, so their confidence is low, when training glove, and they mean the same. So I changed them all into Lowcase.\n\n## CV evaluation\nOn discussion, everyone is worried about whether or not to trust local cv. But a lot of people are just trust single model local cv, not trust ensemble model local cv.. For the former, the assessment is the single model, and the second is the ensemble model.. The ensemble, used at the end of the game should be ensemble model local cv. The methods are as follows: we divide the training set into k folders, then train all the models that need to be blended and blend the prediction results, so that we can get a prediction result of fold, and so on. In addition, I use the corr to see the similarity of fold training, I hope the model is good and different. My approach is dangerous, after all, split folds is random, the more random seeds, the less stable the final result.\n\n## Loss Function\nI tried to use Focal loss as loss function, the single model local cv is very good, there is 0.712. But don't forget just said, this is only single model local cv, not ensemble model local cv. Finally, due to time, it does not do its ensemble local CV, but I think it might be better to adjust the hyperparameters than the Binary Cross Entropy.\n\nMay be my English is poorly. So I write the blog in my [知乎](https://zhuanlan.zhihu.com/p/57550039), the Chinese Quora. Thx everyone share solutions and kernel. I hope that in the future I will share more interesting and effective technical knowledge like everyone else.",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "477266": "## Model structure\n\n![](https://upload-images.jianshu.io/upload_images/6673934-43708f05b93e32fe.png?imageMogr2/auto-orient/strip%7CimageView2/2/w/1240)\n\n## Embedding weight\nThe weight of embedding: 0.5 * glove+0.4 * para+0.1 * wiki. The weights are obtained as follows: First, the vector `x` of the three embeddings in the sentence is obtained, and the shape is (batch_size, maxlen, embed_size, 3). and then I create weight(shape: (3, )) into the softmax layer and multiply it with the tensor `x`. I divide the data into 5-folds, and each fold is averaged by the embedding weight as the final embedding weight. The way that embedding weights are obtained is data driven.\n\n## Word embedding\nFirst I try to preserve the case of the word, and then get the word vector through the pre-training word vector. If there is no match, then the word will be: lower upper capitalize processing to find the corresponding word vector, etc. If not found, the vector is randomly generated. But the general practice is not effective. The main reason for my guess is: There are some weird spellings like: \"tHe\", \"ThE\". I look up the text and find that they originally mean \"the\". But in The glove finds the corresponding embedding vector. These words appear relatively few times, so their confidence is low, when training glove, and they mean the same. So I changed them all into Lowcase.\n\n## CV evaluation\nOn discussion, everyone is worried about whether or not to trust local cv. But a lot of people are just trust single model local cv, not trust ensemble model local cv.. For the former, the assessment is the single model, and the second is the ensemble model.. The ensemble, used at the end of the game should be ensemble model local cv. The methods are as follows: we divide the training set into k folders, then train all the models that need to be blended and blend the prediction results, so that we can get a prediction result of fold, and so on. In addition, I use the corr to see the similarity of fold training, I hope the model is good and different. My approach is dangerous, after all, split folds is random, the more random seeds, the less stable the final result.\n\n## Loss Function\nI tried to use Focal loss as loss function, the single model local cv is very good, there is 0.712. But don't forget just said, this is only single model local cv, not ensemble model local cv. Finally, due to time, it does not do its ensemble local CV, but I think it might be better to adjust the hyperparameters than the Binary Cross Entropy.\n\nMay be my English is poorly. So I write the blog in my [知乎](https://zhuanlan.zhihu.com/p/57550039), the Chinese Quora. Thx everyone share solutions and kernel. I hope that in the future I will share more interesting and effective technical knowledge like everyone else."
  }
}