{
  "id": 161997,
  "title": " 5th place solution",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/zlc-5th-place-solution",
  "author_name": "",
  "post_date": "2020-06-27T12:28:29.040Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Congratulations to all winners!</p>\n\n<p>I would like to thank Kaggle and Jigsaw team for hosting this interesting competition. I am also grateful to this great community. Without those great public kernels I do not know how and where to start this project. I feel honored to end up this competition with a solo gold medal. It is a fantastic learning experience. </p>\n\n<p>As the private leaderboard is finalized, I decide to summarize what I have learned during this process. My solution has a lot in common with other top solutions. The two helpful techniques/tricks are <strong>ensemble of diverse models</strong> and <strong>post-processing</strong>. At certain point I also thought of using monolingual models as discussed in the <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862\">1st place solution</a>. But I failed because I wrongly used XLM-R.</p>\n\n<h1>Ensemble of diverse models</h1>\n\n<h2>Diverse models</h2>\n\n<p>For my personal understanding, langugae models are diverse if the train sets, tokenization methods or model architectures are different. In fact, I did not develop any methods from scratch. Instead, I started from public kernels, and improve them by trial and error. Below is a list of kernels I took advantage of for ensembling in my final submission. </p>\n\n<ul>\n<li><p><a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">XLM-R TPU on PyTorch using Pseudolabeled (PL) opensubtitles</a> by  <a href=\"https://www.kaggle.com/shonenkov\">Alex Shonenkov</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">The first functioning XLM-R model</a> by <a href=\"https://www.kaggle.com/xhlulu\">xhlulu</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">XLM-R with fine-tuned MLM</a> by <a href=\"https://www.kaggle.com/riblidezso\">Dezso Ribli</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training\">Two-stage BERT using translated dataset</a> by <a href=\"https://www.kaggle.com/miklgr500\">Michael Kazachok</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/jhoward/nb-svm-strong-linear-baseline\">NB-SVM</a> by <a href=\"https://www.kaggle.com/jhoward\">Jeremy Howard</a></p></li>\n</ul>\n\n<p>The final submission is an ensemble of 11 models, including two NB-SVM models trained with <a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">Google API translated multilingual train set</a>, two BERT models trained with English train, validation and test sets, one XLM-R model trained with downsampled unintended bias train set and original validation set, one XLM-R model trained with PL test set, one XLM-R model trained using multilingual train set, and one XLM-R model with multilingual train set and PL test set, two XLM-R models with fined-tuned MLM, and three XLM-R models trained using PL opensubtitles. The ensemble weights are found by intuition and probing the public leaderboard score. The weights are</p>\n\n<p><code>python\nweights = {}\nweights['bert_1'] = weights['bert_2'] = 0.4/2\nweights['nbsvm_1'] = weights['nbsvm_2'] = 0.035/2 <br>\nweights['xlm_r_en'] =0.7; weights['xlm_r_pl_en'] = 0.3 <br>\nweights['xlm_r_multilingual']= 0.35; weights['xlm_r_pl_multilingual'] = 0.2\nweights['xlm_r_mlm_en'] = weights['xlm_r_mlm_en_2'] = 0.25/2 <br>\nweights['xlm_r_mlm_multilingual'] = weights['xlm_r_mlm_multilingual_2'] = 0.25/2\nweights['xlm_r_opensubtitle'] = weights['xlm_r_opensubtitle_2'] = weights['xlm_r_opensubtitle_3'] = 0.27/3 \n</code></p>\n\n<h1>Post-processing</h1>\n\n<p>As highlighted in many other top solutions, post-processing contributes a lot to AUC score, since what matters is not the accuracy of single prediction, but the rank/distribution of all predictions. As we need to predict six-language comments in the test set, it is likely that the model may over-predict toxicity for one language while under-predict toxicity for another language. Therefore, scaling the prediction for individual language might be helpful. I first tried to put a 1.3 scaling coefficient for <strong>fr, es and ru</strong>. It gave my a boost of around 0.002 Public LB score for my best ensemble model. I also tried this with two single model predictions and also saw a similar trend. It gave me confidence that this should work generally, at least for the models I used. So for as long as two weeks, I experimented on the scaling coefficients, in the end I found the following coefficients that work best for me.</p>\n\n<p><code>python\nscaling = {\n    'fr': 1.175, \n    'ru': 0.975,\n    'es': 1.475, \n    'it': 0.88, <br>\n    'pt': 0.8, \n    'tr': 1.0 \n}\n</code></p>\n\n<p>Another language-wise post-processing trick is to separate the prediction somehow, it leads to a tiny LB score increase.</p>\n\n<p>```python\ndef post_process(d):\n    # d: list of numbers representing a distribution\n    d = np.array(d)\n    q_low, q_high = np.quantile(d, q=0.55), np.quantile(d, q=0.95)\n    mask_nontoxic, mask_toxic = d&lt;=q_low, d&gt;=q_high\n    d[mask_nontoxic] = 0.9*d[mask_nontoxic]\n    d[mask_toxic] = 1.1*d[mask_toxic]\n    return d</p>\n\n<p>langs = test.lang.unique()\nfor language in langs:\n    mask = test.lang == language\n    sub.loc[mask, 'toxic'] = post_process(sub.loc[mask, 'toxic'])\n```</p>\n\n<h1>Things that worked</h1>\n\n<ul>\n<li>Ensemble</li>\n<li>Post-processing: language-wise scaling</li>\n<li>Add another dense layer before the final dense layer</li>\n<li>Train on more balanced data</li>\n<li>Train on PL test set</li>\n</ul>\n\n<h1>Things that did not work</h1>\n\n<ul>\n<li><p>Ensemble including logistic regression models trained and predicted on either English or multilingual data</p></li>\n<li><p>Monolingual models using XLM-R</p></li>\n<li><p>Add more dense layers</p></li>\n<li><p>Learning rate scheduler</p></li>\n<li><p>Preprocess the data</p></li>\n<li><p>other post-processing techniques</p></li>\n</ul>",
  "messages": [
    {
      "id": "903575",
      "postDate": "06/27/2020 00:47:32",
      "content": "<p>Congratulations to all winners!</p>\n\n<p>I would like to thank Kaggle and Jigsaw team for hosting this interesting competition. I am also grateful to this great community. Without those great public kernels I do not know how and where to start this project. I feel honored to end up this competition with a solo gold medal. It is a fantastic learning experience. </p>\n\n<p>As the private leaderboard is finalized, I decide to summarize what I have learned during this process. My solution has a lot in common with other top solutions. The two helpful techniques/tricks are <strong>ensemble of diverse models</strong> and <strong>post-processing</strong>. At certain point I also thought of using monolingual models as discussed in the <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862\">1st place solution</a>. But I failed because I wrongly used XLM-R.</p>\n\n<h1>Ensemble of diverse models</h1>\n\n<h2>Diverse models</h2>\n\n<p>For my personal understanding, langugae models are diverse if the train sets, tokenization methods or model architectures are different. In fact, I did not develop any methods from scratch. Instead, I started from public kernels, and improve them by trial and error. Below is a list of kernels I took advantage of for ensembling in my final submission. </p>\n\n<ul>\n<li><p><a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">XLM-R TPU on PyTorch using Pseudolabeled (PL) opensubtitles</a> by  <a href=\"https://www.kaggle.com/shonenkov\">Alex Shonenkov</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">The first functioning XLM-R model</a> by <a href=\"https://www.kaggle.com/xhlulu\">xhlulu</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">XLM-R with fine-tuned MLM</a> by <a href=\"https://www.kaggle.com/riblidezso\">Dezso Ribli</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training\">Two-stage BERT using translated dataset</a> by <a href=\"https://www.kaggle.com/miklgr500\">Michael Kazachok</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/jhoward/nb-svm-strong-linear-baseline\">NB-SVM</a> by <a href=\"https://www.kaggle.com/jhoward\">Jeremy Howard</a></p></li>\n</ul>\n\n<p>The final submission is an ensemble of 11 models, including two NB-SVM models trained with <a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">Google API translated multilingual train set</a>, two BERT models trained with English train, validation and test sets, one XLM-R model trained with downsampled unintended bias train set and original validation set, one XLM-R model trained with PL test set, one XLM-R model trained using multilingual train set, and one XLM-R model with multilingual train set and PL test set, two XLM-R models with fined-tuned MLM, and three XLM-R models trained using PL opensubtitles. The ensemble weights are found by intuition and probing the public leaderboard score. The weights are</p>\n\n<p><code>python\nweights = {}\nweights['bert_1'] = weights['bert_2'] = 0.4/2\nweights['nbsvm_1'] = weights['nbsvm_2'] = 0.035/2 <br>\nweights['xlm_r_en'] =0.7; weights['xlm_r_pl_en'] = 0.3 <br>\nweights['xlm_r_multilingual']= 0.35; weights['xlm_r_pl_multilingual'] = 0.2\nweights['xlm_r_mlm_en'] = weights['xlm_r_mlm_en_2'] = 0.25/2 <br>\nweights['xlm_r_mlm_multilingual'] = weights['xlm_r_mlm_multilingual_2'] = 0.25/2\nweights['xlm_r_opensubtitle'] = weights['xlm_r_opensubtitle_2'] = weights['xlm_r_opensubtitle_3'] = 0.27/3 \n</code></p>\n\n<h1>Post-processing</h1>\n\n<p>As highlighted in many other top solutions, post-processing contributes a lot to AUC score, since what matters is not the accuracy of single prediction, but the rank/distribution of all predictions. As we need to predict six-language comments in the test set, it is likely that the model may over-predict toxicity for one language while under-predict toxicity for another language. Therefore, scaling the prediction for individual language might be helpful. I first tried to put a 1.3 scaling coefficient for <strong>fr, es and ru</strong>. It gave my a boost of around 0.002 Public LB score for my best ensemble model. I also tried this with two single model predictions and also saw a similar trend. It gave me confidence that this should work generally, at least for the models I used. So for as long as two weeks, I experimented on the scaling coefficients, in the end I found the following coefficients that work best for me.</p>\n\n<p><code>python\nscaling = {\n    'fr': 1.175, \n    'ru': 0.975,\n    'es': 1.475, \n    'it': 0.88, <br>\n    'pt': 0.8, \n    'tr': 1.0 \n}\n</code></p>\n\n<p>Another language-wise post-processing trick is to separate the prediction somehow, it leads to a tiny LB score increase.</p>\n\n<p>```python\ndef post_process(d):\n    # d: list of numbers representing a distribution\n    d = np.array(d)\n    q_low, q_high = np.quantile(d, q=0.55), np.quantile(d, q=0.95)\n    mask_nontoxic, mask_toxic = d&lt;=q_low, d&gt;=q_high\n    d[mask_nontoxic] = 0.9*d[mask_nontoxic]\n    d[mask_toxic] = 1.1*d[mask_toxic]\n    return d</p>\n\n<p>langs = test.lang.unique()\nfor language in langs:\n    mask = test.lang == language\n    sub.loc[mask, 'toxic'] = post_process(sub.loc[mask, 'toxic'])\n```</p>\n\n<h1>Things that worked</h1>\n\n<ul>\n<li>Ensemble</li>\n<li>Post-processing: language-wise scaling</li>\n<li>Add another dense layer before the final dense layer</li>\n<li>Train on more balanced data</li>\n<li>Train on PL test set</li>\n</ul>\n\n<h1>Things that did not work</h1>\n\n<ul>\n<li><p>Ensemble including logistic regression models trained and predicted on either English or multilingual data</p></li>\n<li><p>Monolingual models using XLM-R</p></li>\n<li><p>Add more dense layers</p></li>\n<li><p>Learning rate scheduler</p></li>\n<li><p>Preprocess the data</p></li>\n<li><p>other post-processing techniques</p></li>\n</ul>",
      "rawMarkdown": "Congratulations to all winners!\n\nI would like to thank Kaggle and Jigsaw team for hosting this interesting competition. I am also grateful to this great community. Without those great public kernels I do not know how and where to start this project. I feel honored to end up this competition with a solo gold medal. It is a fantastic learning experience. \n\nAs the private leaderboard is finalized, I decide to summarize what I have learned during this process. My solution has a lot in common with other top solutions. The two helpful techniques/tricks are **ensemble of diverse models** and **post-processing**. At certain point I also thought of using monolingual models as discussed in the [1st place solution](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862). But I failed because I wrongly used XLM-R.\n\n# Ensemble of diverse models\n\n## Diverse models\n\nFor my personal understanding, langugae models are diverse if the train sets, tokenization methods or model architectures are different. In fact, I did not develop any methods from scratch. Instead, I started from public kernels, and improve them by trial and error. Below is a list of kernels I took advantage of for ensembling in my final submission. \n\n- [XLM-R TPU on PyTorch using Pseudolabeled (PL) opensubtitles](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta) by  [Alex Shonenkov](https://www.kaggle.com/shonenkov)\n\n- [The first functioning XLM-R model](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta) by [xhlulu](https://www.kaggle.com/xhlulu)\n\n- [XLM-R with fine-tuned MLM](https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large) by [Dezso Ribli](https://www.kaggle.com/riblidezso)\n\n- [Two-stage BERT using translated dataset](https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training) by [Michael Kazachok](https://www.kaggle.com/miklgr500)\n\n- [NB-SVM](https://www.kaggle.com/jhoward/nb-svm-strong-linear-baseline) by [Jeremy Howard](https://www.kaggle.com/jhoward)\n\nThe final submission is an ensemble of 11 models, including two NB-SVM models trained with [Google API translated multilingual train set](https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api), two BERT models trained with English train, validation and test sets, one XLM-R model trained with downsampled unintended bias train set and original validation set, one XLM-R model trained with PL test set, one XLM-R model trained using multilingual train set, and one XLM-R model with multilingual train set and PL test set, two XLM-R models with fined-tuned MLM, and three XLM-R models trained using PL opensubtitles. The ensemble weights are found by intuition and probing the public leaderboard score. The weights are\n\n```python\nweights = {}\nweights['bert_1'] = weights['bert_2'] = 0.4/2\nweights['nbsvm_1'] = weights['nbsvm_2'] = 0.035/2                         \nweights['xlm_r_en'] =0.7; weights['xlm_r_pl_en'] = 0.3                \nweights['xlm_r_multilingual']= 0.35; weights['xlm_r_pl_multilingual'] = 0.2\nweights['xlm_r_mlm_en'] = weights['xlm_r_mlm_en_2'] = 0.25/2             \nweights['xlm_r_mlm_multilingual'] = weights['xlm_r_mlm_multilingual_2'] = 0.25/2\nweights['xlm_r_opensubtitle'] = weights['xlm_r_opensubtitle_2'] = weights['xlm_r_opensubtitle_3'] = 0.27/3 \n```\n\n# Post-processing\n\nAs highlighted in many other top solutions, post-processing contributes a lot to AUC score, since what matters is not the accuracy of single prediction, but the rank/distribution of all predictions. As we need to predict six-language comments in the test set, it is likely that the model may over-predict toxicity for one language while under-predict toxicity for another language. Therefore, scaling the prediction for individual language might be helpful. I first tried to put a 1.3 scaling coefficient for **fr, es and ru**. It gave my a boost of around 0.002 Public LB score for my best ensemble model. I also tried this with two single model predictions and also saw a similar trend. It gave me confidence that this should work generally, at least for the models I used. So for as long as two weeks, I experimented on the scaling coefficients, in the end I found the following coefficients that work best for me.\n\n```python\nscaling = {\n    'fr': 1.175, \n    'ru': 0.975,\n    'es': 1.475, \n    'it': 0.88,  \n    'pt': 0.8, \n    'tr': 1.0 \n}\n```\n\nAnother language-wise post-processing trick is to separate the prediction somehow, it leads to a tiny LB score increase.\n\n```python\ndef post_process(d):\n    # d: list of numbers representing a distribution\n    d = np.array(d)\n    q_low, q_high = np.quantile(d, q=0.55), np.quantile(d, q=0.95)\n    mask_nontoxic, mask_toxic = d&lt;=q_low, d&gt;=q_high\n    d[mask_nontoxic] = 0.9*d[mask_nontoxic]\n    d[mask_toxic] = 1.1*d[mask_toxic]\n    return d\n\nlangs = test.lang.unique()\nfor language in langs:\n    mask = test.lang == language\n    sub.loc[mask, 'toxic'] = post_process(sub.loc[mask, 'toxic'])\n```\n\n# Things that worked\n\n- Ensemble\n- Post-processing: language-wise scaling\n- Add another dense layer before the final dense layer\n- Train on more balanced data\n- Train on PL test set\n\n# Things that did not work\n\n- Ensemble including logistic regression models trained and predicted on either English or multilingual data\n\n- Monolingual models using XLM-R\n\n- Add more dense layers\n\n- Learning rate scheduler\n\n- Preprocess the data\n\n- other post-processing techniques",
      "votes": null
    },
    {
      "id": "903597",
      "postDate": "06/27/2020 01:18:05",
      "content": "<p>congratulations! your work is amazing.</p>",
      "rawMarkdown": "congratulations! your work is amazing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 903597,
      "author_name": "ccbuzaijia",
      "author_url": "",
      "post_date": "06/27/2020 01:18:05",
      "content": "<p>congratulations! your work is amazing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "903575": "Congratulations to all winners!\n\nI would like to thank Kaggle and Jigsaw team for hosting this interesting competition. I am also grateful to this great community. Without those great public kernels I do not know how and where to start this project. I feel honored to end up this competition with a solo gold medal. It is a fantastic learning experience. \n\nAs the private leaderboard is finalized, I decide to summarize what I have learned during this process. My solution has a lot in common with other top solutions. The two helpful techniques/tricks are **ensemble of diverse models** and **post-processing**. At certain point I also thought of using monolingual models as discussed in the [1st place solution](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862). But I failed because I wrongly used XLM-R.\n\n# Ensemble of diverse models\n\n## Diverse models\n\nFor my personal understanding, langugae models are diverse if the train sets, tokenization methods or model architectures are different. In fact, I did not develop any methods from scratch. Instead, I started from public kernels, and improve them by trial and error. Below is a list of kernels I took advantage of for ensembling in my final submission. \n\n- [XLM-R TPU on PyTorch using Pseudolabeled (PL) opensubtitles](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta) by  [Alex Shonenkov](https://www.kaggle.com/shonenkov)\n\n- [The first functioning XLM-R model](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta) by [xhlulu](https://www.kaggle.com/xhlulu)\n\n- [XLM-R with fine-tuned MLM](https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large) by [Dezso Ribli](https://www.kaggle.com/riblidezso)\n\n- [Two-stage BERT using translated dataset](https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training) by [Michael Kazachok](https://www.kaggle.com/miklgr500)\n\n- [NB-SVM](https://www.kaggle.com/jhoward/nb-svm-strong-linear-baseline) by [Jeremy Howard](https://www.kaggle.com/jhoward)\n\nThe final submission is an ensemble of 11 models, including two NB-SVM models trained with [Google API translated multilingual train set](https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api), two BERT models trained with English train, validation and test sets, one XLM-R model trained with downsampled unintended bias train set and original validation set, one XLM-R model trained with PL test set, one XLM-R model trained using multilingual train set, and one XLM-R model with multilingual train set and PL test set, two XLM-R models with fined-tuned MLM, and three XLM-R models trained using PL opensubtitles. The ensemble weights are found by intuition and probing the public leaderboard score. The weights are\n\n```python\nweights = {}\nweights['bert_1'] = weights['bert_2'] = 0.4/2\nweights['nbsvm_1'] = weights['nbsvm_2'] = 0.035/2                         \nweights['xlm_r_en'] =0.7; weights['xlm_r_pl_en'] = 0.3                \nweights['xlm_r_multilingual']= 0.35; weights['xlm_r_pl_multilingual'] = 0.2\nweights['xlm_r_mlm_en'] = weights['xlm_r_mlm_en_2'] = 0.25/2             \nweights['xlm_r_mlm_multilingual'] = weights['xlm_r_mlm_multilingual_2'] = 0.25/2\nweights['xlm_r_opensubtitle'] = weights['xlm_r_opensubtitle_2'] = weights['xlm_r_opensubtitle_3'] = 0.27/3 \n```\n\n# Post-processing\n\nAs highlighted in many other top solutions, post-processing contributes a lot to AUC score, since what matters is not the accuracy of single prediction, but the rank/distribution of all predictions. As we need to predict six-language comments in the test set, it is likely that the model may over-predict toxicity for one language while under-predict toxicity for another language. Therefore, scaling the prediction for individual language might be helpful. I first tried to put a 1.3 scaling coefficient for **fr, es and ru**. It gave my a boost of around 0.002 Public LB score for my best ensemble model. I also tried this with two single model predictions and also saw a similar trend. It gave me confidence that this should work generally, at least for the models I used. So for as long as two weeks, I experimented on the scaling coefficients, in the end I found the following coefficients that work best for me.\n\n```python\nscaling = {\n    'fr': 1.175, \n    'ru': 0.975,\n    'es': 1.475, \n    'it': 0.88,  \n    'pt': 0.8, \n    'tr': 1.0 \n}\n```\n\nAnother language-wise post-processing trick is to separate the prediction somehow, it leads to a tiny LB score increase.\n\n```python\ndef post_process(d):\n    # d: list of numbers representing a distribution\n    d = np.array(d)\n    q_low, q_high = np.quantile(d, q=0.55), np.quantile(d, q=0.95)\n    mask_nontoxic, mask_toxic = d&lt;=q_low, d&gt;=q_high\n    d[mask_nontoxic] = 0.9*d[mask_nontoxic]\n    d[mask_toxic] = 1.1*d[mask_toxic]\n    return d\n\nlangs = test.lang.unique()\nfor language in langs:\n    mask = test.lang == language\n    sub.loc[mask, 'toxic'] = post_process(sub.loc[mask, 'toxic'])\n```\n\n# Things that worked\n\n- Ensemble\n- Post-processing: language-wise scaling\n- Add another dense layer before the final dense layer\n- Train on more balanced data\n- Train on PL test set\n\n# Things that did not work\n\n- Ensemble including logistic regression models trained and predicted on either English or multilingual data\n\n- Monolingual models using XLM-R\n\n- Add more dense layers\n\n- Learning rate scheduler\n\n- Preprocess the data\n\n- other post-processing techniques",
    "903597": "congratulations! your work is amazing."
  },
  "source": "meta"
}