{
  "id": 232158,
  "title": "Has someone meet Exposure bias problem?",
  "url": "/competitions/bms-molecular-translation/discussion/232158",
  "author_name": "",
  "post_date": "2021-04-12T12:23:00.405114900Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>My CNN-LSTM with Attention model has a huge gap between training and validation when training with fully teacher forcing. To mitigate such issue, I adopted scheduled sampling to<br>\nmatch model distribution and data distribution, then the resulting model achieved 9+ LB（could be further imporved） and the gap between training and CV is much close.</p>\n<p>I see that published <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">notebook</a> by Y.Nakama has no such exposure bias and perform very well on LB, though trained fully in teacher forcing. So I want to ask has someone experienced same issue like me?</p>",
  "messages": [
    {
      "id": "1271229",
      "postDate": "04/12/2021 12:23:00",
      "content": "<p>My CNN-LSTM with Attention model has a huge gap between training and validation when training with fully teacher forcing. To mitigate such issue, I adopted scheduled sampling to<br>\nmatch model distribution and data distribution, then the resulting model achieved 9+ LB（could be further imporved） and the gap between training and CV is much close.</p>\n<p>I see that published <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">notebook</a> by Y.Nakama has no such exposure bias and perform very well on LB, though trained fully in teacher forcing. So I want to ask has someone experienced same issue like me?</p>",
      "rawMarkdown": "My CNN-LSTM with Attention model has a huge gap between training and validation when training with fully teacher forcing. To mitigate such issue, I adopted scheduled sampling to\nmatch model distribution and data distribution, then the resulting model achieved 9+ LB（could be further imporved） and the gap between training and CV is much close.\n\nI see that published [notebook](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter) by Y.Nakama has no such exposure bias and perform very well on LB, though trained fully in teacher forcing. So I want to ask has someone experienced same issue like me?",
      "votes": null
    },
    {
      "id": "1271822",
      "postDate": "04/13/2021 00:08:23",
      "content": "<p>once your model accuracy is high enough, exposure bias is less of a problem. <br>\nI think CNN can get about &gt;95% single token classification rate. Together with k-beam search, this almost solved the problem. you should monitor loss, top-k classification (per token) in your validation too</p>\n<p>I recall reading one paper saying that even if you take care of exposure bias (using better training method), your gain is only about 2 to 3%</p>\n<p>For this dataset, i think there is not much of exposure bias. the train/valid/lb scores are pretty close<br>\ni have a bug previously, that cause my train/cv/lb gap.</p>",
      "rawMarkdown": "once your model accuracy is high enough, exposure bias is less of a problem. \nI think CNN can get about >95% single token classification rate. Together with k-beam search, this almost solved the problem. you should monitor loss, top-k classification (per token) in your validation too\n\nI recall reading one paper saying that even if you take care of exposure bias (using better training method), your gain is only about 2 to 3%\n\nFor this dataset, i think there is not much of exposure bias. the train/valid/lb scores are pretty close\ni have a bug previously, that cause my train/cv/lb gap.",
      "votes": null
    },
    {
      "id": "1271868",
      "postDate": "04/13/2021 02:12:38",
      "content": "<p>Thans for your sharing, really helpful. BTW,  could you post that paper, name or link, very appreciate!</p>",
      "rawMarkdown": "Thans for your sharing, really helpful. BTW,  could you post that paper, name or link, very appreciate!",
      "votes": null
    },
    {
      "id": "1271892",
      "postDate": "04/13/2021 03:01:17",
      "content": "<p>However, according to our measurements in controlled experiments, there's only around 3% performance gain when the training-inference discrepancy is completely removed</p>\n<p><a href=\"https://openreview.net/forum?id=rJg2fTNtwr\" target=\"_blank\">https://openreview.net/forum?id=rJg2fTNtwr</a></p>",
      "rawMarkdown": "However, according to our measurements in controlled experiments, there's only around 3% performance gain when the training-inference discrepancy is completely removed\n\nhttps://openreview.net/forum?id=rJg2fTNtwr",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1271822,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/13/2021 00:08:23",
      "content": "<p>once your model accuracy is high enough, exposure bias is less of a problem. <br>\nI think CNN can get about &gt;95% single token classification rate. Together with k-beam search, this almost solved the problem. you should monitor loss, top-k classification (per token) in your validation too</p>\n<p>I recall reading one paper saying that even if you take care of exposure bias (using better training method), your gain is only about 2 to 3%</p>\n<p>For this dataset, i think there is not much of exposure bias. the train/valid/lb scores are pretty close<br>\ni have a bug previously, that cause my train/cv/lb gap.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1271868,
          "author_name": "wantsu",
          "author_url": "",
          "post_date": "04/13/2021 02:12:38",
          "content": "<p>Thans for your sharing, really helpful. BTW,  could you post that paper, name or link, very appreciate!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1271892,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/13/2021 03:01:17",
          "content": "<p>However, according to our measurements in controlled experiments, there's only around 3% performance gain when the training-inference discrepancy is completely removed</p>\n<p><a href=\"https://openreview.net/forum?id=rJg2fTNtwr\" target=\"_blank\">https://openreview.net/forum?id=rJg2fTNtwr</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1271229": "My CNN-LSTM with Attention model has a huge gap between training and validation when training with fully teacher forcing. To mitigate such issue, I adopted scheduled sampling to\nmatch model distribution and data distribution, then the resulting model achieved 9+ LB（could be further imporved） and the gap between training and CV is much close.\n\nI see that published [notebook](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter) by Y.Nakama has no such exposure bias and perform very well on LB, though trained fully in teacher forcing. So I want to ask has someone experienced same issue like me?",
    "1271822": "once your model accuracy is high enough, exposure bias is less of a problem. \nI think CNN can get about >95% single token classification rate. Together with k-beam search, this almost solved the problem. you should monitor loss, top-k classification (per token) in your validation too\n\nI recall reading one paper saying that even if you take care of exposure bias (using better training method), your gain is only about 2 to 3%\n\nFor this dataset, i think there is not much of exposure bias. the train/valid/lb scores are pretty close\ni have a bug previously, that cause my train/cv/lb gap.",
    "1271868": "Thans for your sharing, really helpful. BTW,  could you post that paper, name or link, very appreciate!",
    "1271892": "However, according to our measurements in controlled experiments, there's only around 3% performance gain when the training-inference discrepancy is completely removed\n\nhttps://openreview.net/forum?id=rJg2fTNtwr"
  },
  "source": "meta"
}