{
  "id": 126609,
  "title": "bert-joint model answers thresholds",
  "url": "/competitions/tensorflow2-question-answering/discussion/126609",
  "author_name": "",
  "post_date": "2020-01-18T19:16:58.113265600Z",
  "votes": 4,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi everyone, I am using Bert-joint model and got predictions with confidence scores. I am confused about the settings of long answer and short answer thresholds. Is there anyone has some ideas about setting thresholds to get a better f1 score? I would be appreciate if someone could share your ideas or experience on that. Thanks a lot.</p>",
  "messages": [
    {
      "id": "722593",
      "postDate": "01/18/2020 19:16:58",
      "content": "<p>Hi everyone, I am using Bert-joint model and got predictions with confidence scores. I am confused about the settings of long answer and short answer thresholds. Is there anyone has some ideas about setting thresholds to get a better f1 score? I would be appreciate if someone could share your ideas or experience on that. Thanks a lot.</p>",
      "rawMarkdown": "Hi everyone, I am using Bert-joint model and got predictions with confidence scores. I am confused about the settings of long answer and short answer thresholds. Is there anyone has some ideas about setting thresholds to get a better f1 score? I would be appreciate if someone could share your ideas or experience on that. Thanks a lot.",
      "votes": null
    },
    {
      "id": "722776",
      "postDate": "01/19/2020 05:01:58",
      "content": "<p><a href=\"/cztestforkaggle\">@cztestforkaggle</a> You are 8th in this competition right now, I'd expect you have more insights to share than other 1000+ participants in this competition :) Can you share your tips and experience so far? Regarding the bert-joint thresholds, they can be tuned to get a better score on public lb, you can tune them on your dev set as well.</p>",
      "rawMarkdown": "cztestforkaggle You are 8th in this competition right now, I'd expect you have more insights to share than other 1000+ participants in this competition :) Can you share your tips and experience so far? Regarding the bert-joint thresholds, they can be tuned to get a better score on public lb, you can tune them on your dev set as well.",
      "votes": null
    },
    {
      "id": "722808",
      "postDate": "01/19/2020 06:20:05",
      "content": "<p><a href=\"/thedrcat\">@thedrcat</a> I wouldn't depend much on thresholds, it's likely to overfit on public LB which you'll regret on private test set :P</p>",
      "rawMarkdown": "thedrcat I wouldn't depend much on thresholds, it's likely to overfit on public LB which you'll regret on private test set :P",
      "votes": null
    },
    {
      "id": "722829",
      "postDate": "01/19/2020 06:52:39",
      "content": "<p><a href=\"/axel81\">@axel81</a> ,</p>\n\n<p>It would be nice to share a bit your non-regret approach :)</p>",
      "rawMarkdown": "axel81 ,\n\nIt would be nice to share a bit your non-regret approach :)",
      "votes": null
    },
    {
      "id": "722833",
      "postDate": "01/19/2020 06:56:28",
      "content": "<p><a href=\"/thedrcat\">@thedrcat</a> I didn't figure out how the metrics works. I tried several thresholds and the f1 score diffs from 0.56 to 0.68. So the thresholds really matters if you are using Bert-joint model.</p>",
      "rawMarkdown": "thedrcat I didn't figure out how the metrics works. I tried several thresholds and the f1 score diffs from 0.56 to 0.68. So the thresholds really matters if you are using Bert-joint model.",
      "votes": null
    },
    {
      "id": "722855",
      "postDate": "01/19/2020 07:59:40",
      "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a>  I didn't do hit and trial to figure out better thresholds as it might overfit the Public LB and  fail on private LB. So at start I did some validation on few samples and I am using a fixed threshold ever since :). With improvements in my model I can see improvement in LB without playing with thresholds. I learnt this from Severstal Steel Contest tbh, thresholds cost us our ranks there :P</p>",
      "rawMarkdown": "yihdarshieh  I didn't do hit and trial to figure out better thresholds as it might overfit the Public LB and  fail on private LB. So at start I did some validation on few samples and I am using a fixed threshold ever since :). With improvements in my model I can see improvement in LB without playing with thresholds. I learnt this from Severstal Steel Contest tbh, thresholds cost us our ranks there :P",
      "votes": null
    },
    {
      "id": "722861",
      "postDate": "01/19/2020 08:09:13",
      "content": "<p><a href=\"/axel81\">@axel81</a> , I am interested in your model improvement, I don't have any after I finished my three published tf kernel 😂</p>",
      "rawMarkdown": "axel81 , I am interested in your model improvement, I don't have any after I finished my three published tf kernel 😂",
      "votes": null
    },
    {
      "id": "722868",
      "postDate": "01/19/2020 08:21:14",
      "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> are you saying  you are using bert-joint baseline from 0.6 kernel and able to get 0.63 with threshold?</p>",
      "rawMarkdown": "yihdarshieh are you saying  you are using bert-joint baseline from 0.6 kernel and able to get 0.63 with threshold?",
      "votes": null
    },
    {
      "id": "722870",
      "postDate": "01/19/2020 08:29:38",
      "content": "<p>No, i used my kernels, but it's just bert joint. I used the ideas from 2 kernels scoring 0.48 and 0.57. I searched a lot of thresholds to get 0.63. I will definitely regret but I HAVE NO other idea on model construction. I will try the idea in the kernel scoring 0.60, but it's still post processing</p>",
      "rawMarkdown": "No, i used my kernels, but it's just bert joint. I used the ideas from 2 kernels scoring 0.48 and 0.57. I searched a lot of thresholds to get 0.63. I will definitely regret but I HAVE NO other idea on model construction. I will try the idea in the kernel scoring 0.60, but it's still post processing",
      "votes": null
    },
    {
      "id": "722896",
      "postDate": "01/19/2020 09:19:07",
      "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> the 0.6 kernel has same postprocessing it just plays with thresholds, better look for other things than trying that ;)</p>",
      "rawMarkdown": "yihdarshieh the 0.6 kernel has same postprocessing it just plays with thresholds, better look for other things than trying that ;)",
      "votes": null
    },
    {
      "id": "723085",
      "postDate": "01/19/2020 13:53:52",
      "content": "<p>Oh, you are doing pretty well. We are about to ask you what is the best way to apply thresholds? Or use classifier outputs as a toggle? </p>",
      "rawMarkdown": "Oh, you are doing pretty well. We are about to ask you what is the best way to apply thresholds? Or use classifier outputs as a toggle?",
      "votes": null
    },
    {
      "id": "723096",
      "postDate": "01/19/2020 14:05:07",
      "content": "<p>I have tried to add a bi-LSTM classifier at the last layer of BERT, but it is hard to train.</p>",
      "rawMarkdown": "I have tried to add a bi-LSTM classifier at the last layer of BERT, but it is hard to train.",
      "votes": null
    },
    {
      "id": "723443",
      "postDate": "01/20/2020 03:25:46",
      "content": "<p>I agree with <a href=\"/axel81\">@axel81</a> that thresholds are not dependable. With certain ones I get a dev score of 0.66 and think it's an improvement and then score 0.62 on the public LB. Out of curiosity, what threshold values did you try? </p>",
      "rawMarkdown": "I agree with @axel81 that thresholds are not dependable. With certain ones I get a dev score of 0.66 and think it's an improvement and then score 0.62 on the public LB. Out of curiosity, what threshold values did you try?",
      "votes": null
    },
    {
      "id": "723460",
      "postDate": "01/20/2020 04:24:26",
      "content": "<p>I also agree with <a href=\"/axel81\">@axel81</a> that better not to depend on thresholds (too much). The thing is that after spending a lot of time working on training, inference and, in particular, evaluation kernels, I didn't get much time on the model itself (also because of lacking some creativity, I guess).</p>\n\n<p>And all high scoring public kernels are about post-processing (let me know if I am wrong). It would be nice for someone to  share different ideas about the models, but probably they would like to keep it as private ...</p>\n\n<p>I didn't get much fun in this competition (other than coding well the kernels), especially due to the lack of a clear metric, and also being not able to have some new ideas about the models </p>",
      "rawMarkdown": "I also agree with @axel81 that better not to depend on thresholds (too much). The thing is that after spending a lot of time working on training, inference and, in particular, evaluation kernels, I didn't get much time on the model itself (also because of lacking some creativity, I guess).\n\nAnd all high scoring public kernels are about post-processing (let me know if I am wrong). It would be nice for someone to  share different ideas about the models, but probably they would like to keep it as private ...\n\nI didn't get much fun in this competition (other than coding well the kernels), especially due to the lack of a clear metric, and also being not able to have some new ideas about the models",
      "votes": null
    },
    {
      "id": "723462",
      "postDate": "01/20/2020 04:29:21",
      "content": "<p>So far it has been more about implementation for me too. But I still love it as am still new to NLP :\")</p>",
      "rawMarkdown": "So far it has been more about implementation for me too. But I still love it as am still new to NLP :\")",
      "votes": null
    },
    {
      "id": "723504",
      "postDate": "01/20/2020 05:26:47",
      "content": "<p>it is not right , only use  Bert large XLNET ROBERTA SPANBERT ALBERT and emsemble  them is a good solution</p>",
      "rawMarkdown": "it is not right , only use  Bert large XLNET ROBERTA SPANBERT ALBERT and emsemble  them is a good solution",
      "votes": null
    },
    {
      "id": "723574",
      "postDate": "01/20/2020 07:07:02",
      "content": "<p><a href=\"/yaoruoping\">@yaoruoping</a> What about the kernel's time limit?</p>",
      "rawMarkdown": "yaoruoping What about the kernel's time limit?",
      "votes": null
    },
    {
      "id": "724728",
      "postDate": "01/21/2020 12:34:54",
      "content": "<p><a href=\"/yaroshevskiy\">@yaroshevskiy</a> You seem to be doing well as well. What have you tried? 👀 </p>",
      "rawMarkdown": "yaroshevskiy You seem to be doing well as well. What have you tried? 👀",
      "votes": null
    },
    {
      "id": "724781",
      "postDate": "01/21/2020 13:29:10",
      "content": "<p>Implemented deafult joint in torch, added few obvious improvements, changed transformer encoder. Nothing cool really</p>",
      "rawMarkdown": "Implemented deafult joint in torch, added few obvious improvements, changed transformer encoder. Nothing cool really",
      "votes": null
    },
    {
      "id": "724813",
      "postDate": "01/21/2020 14:04:08",
      "content": "<p>Could you tell us what \"changed transformer encoder\" means?</p>",
      "rawMarkdown": "Could you tell us what \"changed transformer encoder\" means?",
      "votes": null
    },
    {
      "id": "724818",
      "postDate": "01/21/2020 14:12:36",
      "content": "<p>You Can USE a SIMPLE MODEL ,such as BM25, BERT-BASE CNN(our used) .. , and predict the  result (feature or candidate) with top scores).\nI can not say more until the competetion finish</p>",
      "rawMarkdown": "You Can USE a SIMPLE MODEL ,such as BM25, BERT-BASE CNN(our used) .. , and predict the  result (feature or candidate) with top scores).\nI can not say more until the competetion finish",
      "votes": null
    },
    {
      "id": "724834",
      "postDate": "01/21/2020 14:29:37",
      "content": "<p>there are bert, roberta, distilbert, albert, lot of them</p>",
      "rawMarkdown": "there are bert, roberta, distilbert, albert, lot of them",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 722776,
      "author_name": "thedrcat",
      "author_url": "",
      "post_date": "01/19/2020 05:01:58",
      "content": "<p><a href=\"/cztestforkaggle\">@cztestforkaggle</a> You are 8th in this competition right now, I'd expect you have more insights to share than other 1000+ participants in this competition :) Can you share your tips and experience so far? Regarding the bert-joint thresholds, they can be tuned to get a better score on public lb, you can tune them on your dev set as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 722808,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/19/2020 06:20:05",
          "content": "<p><a href=\"/thedrcat\">@thedrcat</a> I wouldn't depend much on thresholds, it's likely to overfit on public LB which you'll regret on private test set :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722829,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/19/2020 06:52:39",
          "content": "<p><a href=\"/axel81\">@axel81</a> ,</p>\n\n<p>It would be nice to share a bit your non-regret approach :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722833,
          "author_name": "cztestforkaggle",
          "author_url": "",
          "post_date": "01/19/2020 06:56:28",
          "content": "<p><a href=\"/thedrcat\">@thedrcat</a> I didn't figure out how the metrics works. I tried several thresholds and the f1 score diffs from 0.56 to 0.68. So the thresholds really matters if you are using Bert-joint model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722855,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/19/2020 07:59:40",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a>  I didn't do hit and trial to figure out better thresholds as it might overfit the Public LB and  fail on private LB. So at start I did some validation on few samples and I am using a fixed threshold ever since :). With improvements in my model I can see improvement in LB without playing with thresholds. I learnt this from Severstal Steel Contest tbh, thresholds cost us our ranks there :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722861,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/19/2020 08:09:13",
          "content": "<p><a href=\"/axel81\">@axel81</a> , I am interested in your model improvement, I don't have any after I finished my three published tf kernel 😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722868,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/19/2020 08:21:14",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> are you saying  you are using bert-joint baseline from 0.6 kernel and able to get 0.63 with threshold?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722870,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/19/2020 08:29:38",
          "content": "<p>No, i used my kernels, but it's just bert joint. I used the ideas from 2 kernels scoring 0.48 and 0.57. I searched a lot of thresholds to get 0.63. I will definitely regret but I HAVE NO other idea on model construction. I will try the idea in the kernel scoring 0.60, but it's still post processing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722896,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/19/2020 09:19:07",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> the 0.6 kernel has same postprocessing it just plays with thresholds, better look for other things than trying that ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 723085,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "01/19/2020 13:53:52",
      "content": "<p>Oh, you are doing pretty well. We are about to ask you what is the best way to apply thresholds? Or use classifier outputs as a toggle? </p>",
      "votes": null,
      "replies": [
        {
          "id": 723096,
          "author_name": "guozhiyu0914",
          "author_url": "",
          "post_date": "01/19/2020 14:05:07",
          "content": "<p>I have tried to add a bi-LSTM classifier at the last layer of BERT, but it is hard to train.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724728,
          "author_name": "msheriey",
          "author_url": "",
          "post_date": "01/21/2020 12:34:54",
          "content": "<p><a href=\"/yaroshevskiy\">@yaroshevskiy</a> You seem to be doing well as well. What have you tried? 👀 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724781,
          "author_name": "yaroshevskiy",
          "author_url": "",
          "post_date": "01/21/2020 13:29:10",
          "content": "<p>Implemented deafult joint in torch, added few obvious improvements, changed transformer encoder. Nothing cool really</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724813,
          "author_name": "mxmka87",
          "author_url": "",
          "post_date": "01/21/2020 14:04:08",
          "content": "<p>Could you tell us what \"changed transformer encoder\" means?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724834,
          "author_name": "yaroshevskiy",
          "author_url": "",
          "post_date": "01/21/2020 14:29:37",
          "content": "<p>there are bert, roberta, distilbert, albert, lot of them</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 723443,
      "author_name": "msheriey",
      "author_url": "",
      "post_date": "01/20/2020 03:25:46",
      "content": "<p>I agree with <a href=\"/axel81\">@axel81</a> that thresholds are not dependable. With certain ones I get a dev score of 0.66 and think it's an improvement and then score 0.62 on the public LB. Out of curiosity, what threshold values did you try? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 723460,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/20/2020 04:24:26",
      "content": "<p>I also agree with <a href=\"/axel81\">@axel81</a> that better not to depend on thresholds (too much). The thing is that after spending a lot of time working on training, inference and, in particular, evaluation kernels, I didn't get much time on the model itself (also because of lacking some creativity, I guess).</p>\n\n<p>And all high scoring public kernels are about post-processing (let me know if I am wrong). It would be nice for someone to  share different ideas about the models, but probably they would like to keep it as private ...</p>\n\n<p>I didn't get much fun in this competition (other than coding well the kernels), especially due to the lack of a clear metric, and also being not able to have some new ideas about the models </p>",
      "votes": null,
      "replies": [
        {
          "id": 723462,
          "author_name": "msheriey",
          "author_url": "",
          "post_date": "01/20/2020 04:29:21",
          "content": "<p>So far it has been more about implementation for me too. But I still love it as am still new to NLP :\")</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 723504,
          "author_name": "yaoruoping",
          "author_url": "",
          "post_date": "01/20/2020 05:26:47",
          "content": "<p>it is not right , only use  Bert large XLNET ROBERTA SPANBERT ALBERT and emsemble  them is a good solution</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 723574,
          "author_name": "msheriey",
          "author_url": "",
          "post_date": "01/20/2020 07:07:02",
          "content": "<p><a href=\"/yaoruoping\">@yaoruoping</a> What about the kernel's time limit?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724818,
          "author_name": "yaoruoping",
          "author_url": "",
          "post_date": "01/21/2020 14:12:36",
          "content": "<p>You Can USE a SIMPLE MODEL ,such as BM25, BERT-BASE CNN(our used) .. , and predict the  result (feature or candidate) with top scores).\nI can not say more until the competetion finish</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "722593": "Hi everyone, I am using Bert-joint model and got predictions with confidence scores. I am confused about the settings of long answer and short answer thresholds. Is there anyone has some ideas about setting thresholds to get a better f1 score? I would be appreciate if someone could share your ideas or experience on that. Thanks a lot.",
    "722776": "cztestforkaggle You are 8th in this competition right now, I'd expect you have more insights to share than other 1000+ participants in this competition :) Can you share your tips and experience so far? Regarding the bert-joint thresholds, they can be tuned to get a better score on public lb, you can tune them on your dev set as well.",
    "722808": "thedrcat I wouldn't depend much on thresholds, it's likely to overfit on public LB which you'll regret on private test set :P",
    "722829": "axel81 ,\n\nIt would be nice to share a bit your non-regret approach :)",
    "722833": "thedrcat I didn't figure out how the metrics works. I tried several thresholds and the f1 score diffs from 0.56 to 0.68. So the thresholds really matters if you are using Bert-joint model.",
    "722855": "yihdarshieh  I didn't do hit and trial to figure out better thresholds as it might overfit the Public LB and  fail on private LB. So at start I did some validation on few samples and I am using a fixed threshold ever since :). With improvements in my model I can see improvement in LB without playing with thresholds. I learnt this from Severstal Steel Contest tbh, thresholds cost us our ranks there :P",
    "722861": "axel81 , I am interested in your model improvement, I don't have any after I finished my three published tf kernel 😂",
    "722868": "yihdarshieh are you saying  you are using bert-joint baseline from 0.6 kernel and able to get 0.63 with threshold?",
    "722870": "No, i used my kernels, but it's just bert joint. I used the ideas from 2 kernels scoring 0.48 and 0.57. I searched a lot of thresholds to get 0.63. I will definitely regret but I HAVE NO other idea on model construction. I will try the idea in the kernel scoring 0.60, but it's still post processing",
    "722896": "yihdarshieh the 0.6 kernel has same postprocessing it just plays with thresholds, better look for other things than trying that ;)",
    "723085": "Oh, you are doing pretty well. We are about to ask you what is the best way to apply thresholds? Or use classifier outputs as a toggle?",
    "723096": "I have tried to add a bi-LSTM classifier at the last layer of BERT, but it is hard to train.",
    "723443": "I agree with @axel81 that thresholds are not dependable. With certain ones I get a dev score of 0.66 and think it's an improvement and then score 0.62 on the public LB. Out of curiosity, what threshold values did you try?",
    "723460": "I also agree with @axel81 that better not to depend on thresholds (too much). The thing is that after spending a lot of time working on training, inference and, in particular, evaluation kernels, I didn't get much time on the model itself (also because of lacking some creativity, I guess).\n\nAnd all high scoring public kernels are about post-processing (let me know if I am wrong). It would be nice for someone to  share different ideas about the models, but probably they would like to keep it as private ...\n\nI didn't get much fun in this competition (other than coding well the kernels), especially due to the lack of a clear metric, and also being not able to have some new ideas about the models",
    "723462": "So far it has been more about implementation for me too. But I still love it as am still new to NLP :\")",
    "723504": "it is not right , only use  Bert large XLNET ROBERTA SPANBERT ALBERT and emsemble  them is a good solution",
    "723574": "yaoruoping What about the kernel's time limit?",
    "724728": "yaroshevskiy You seem to be doing well as well. What have you tried? 👀",
    "724781": "Implemented deafult joint in torch, added few obvious improvements, changed transformer encoder. Nothing cool really",
    "724813": "Could you tell us what \"changed transformer encoder\" means?",
    "724818": "You Can USE a SIMPLE MODEL ,such as BM25, BERT-BASE CNN(our used) .. , and predict the  result (feature or candidate) with top scores).\nI can not say more until the competetion finish",
    "724834": "there are bert, roberta, distilbert, albert, lot of them"
  },
  "source": "meta"
}