{
  "id": 121852,
  "title": "Is it worth participating in this competition?",
  "url": "/competitions/tensorflow2-question-answering/discussion/121852",
  "author_name": "",
  "post_date": "2019-12-16T05:34:28.819436800Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello all, </p>\n\n<p>I have been interested in question answering as a field for a long time, ever since the days people were using memory networks on bAbi datasets! So, seeing the progress the field has made in the last couple of years, I was very excited to participate in this competition and try out the new techniques field. However, here's why I am hesitant to participate a month before the competition ends:</p>\n\n<ol>\n<li>The metric still doesn't seem to be fully understood with only a month left in the competition.</li>\n<li>There's no PyTorch starter scoring kernel. This is honestly disappointing. Literally any kernel on doing some simple training and inference with a simple BERT model would prove to be useful to many competitors (including myself) I suspect.</li>\n<li>Just the overall lack of EDA and good modeling kernels this far into the competition. TFQA currently only has 15 scoring kernels. And those kernels basically are the provided baseline, binary classification baseline, LSTM baseline, and the bert joint baseline (YES/NO kernel being a modification of this). On the other hand, Google QUEST challenge has about 53 scoring kernels. And there are many interesting modeling approaches, from different BERT models, to LSTM baselines, to Universal Sentence Encoder approaches, and even Ridge Regression! I like to take inspiration from existing kernels, and the lack of such kernels (and also the lack of variation) at this stage in the competition is disappointing.</li>\n</ol>\n\n<p>What do you guys think? Are these also valid concerns you have had? Are there opportunities for some of these problems to be resolved soon?</p>",
  "messages": [
    {
      "id": "696094",
      "postDate": "12/16/2019 05:34:28",
      "content": "<p>Hello all, </p>\n\n<p>I have been interested in question answering as a field for a long time, ever since the days people were using memory networks on bAbi datasets! So, seeing the progress the field has made in the last couple of years, I was very excited to participate in this competition and try out the new techniques field. However, here's why I am hesitant to participate a month before the competition ends:</p>\n\n<ol>\n<li>The metric still doesn't seem to be fully understood with only a month left in the competition.</li>\n<li>There's no PyTorch starter scoring kernel. This is honestly disappointing. Literally any kernel on doing some simple training and inference with a simple BERT model would prove to be useful to many competitors (including myself) I suspect.</li>\n<li>Just the overall lack of EDA and good modeling kernels this far into the competition. TFQA currently only has 15 scoring kernels. And those kernels basically are the provided baseline, binary classification baseline, LSTM baseline, and the bert joint baseline (YES/NO kernel being a modification of this). On the other hand, Google QUEST challenge has about 53 scoring kernels. And there are many interesting modeling approaches, from different BERT models, to LSTM baselines, to Universal Sentence Encoder approaches, and even Ridge Regression! I like to take inspiration from existing kernels, and the lack of such kernels (and also the lack of variation) at this stage in the competition is disappointing.</li>\n</ol>\n\n<p>What do you guys think? Are these also valid concerns you have had? Are there opportunities for some of these problems to be resolved soon?</p>",
      "rawMarkdown": "Hello all, \n\nI have been interested in question answering as a field for a long time, ever since the days people were using memory networks on bAbi datasets! So, seeing the progress the field has made in the last couple of years, I was very excited to participate in this competition and try out the new techniques field. However, here's why I am hesitant to participate a month before the competition ends:\n\n1. The metric still doesn't seem to be fully understood with only a month left in the competition.\n2. There's no PyTorch starter scoring kernel. This is honestly disappointing. Literally any kernel on doing some simple training and inference with a simple BERT model would prove to be useful to many competitors (including myself) I suspect.\n3. Just the overall lack of EDA and good modeling kernels this far into the competition. TFQA currently only has 15 scoring kernels. And those kernels basically are the provided baseline, binary classification baseline, LSTM baseline, and the bert joint baseline (YES/NO kernel being a modification of this). On the other hand, Google QUEST challenge has about 53 scoring kernels. And there are many interesting modeling approaches, from different BERT models, to LSTM baselines, to Universal Sentence Encoder approaches, and even Ridge Regression! I like to take inspiration from existing kernels, and the lack of such kernels (and also the lack of variation) at this stage in the competition is disappointing.\n\nWhat do you guys think? Are these also valid concerns you have had? Are there opportunities for some of these problems to be resolved soon?",
      "votes": null
    },
    {
      "id": "696271",
      "postDate": "12/16/2019 11:28:53",
      "content": "<ol>\n<li>Metric of course is the most annoying part. But some folks report that the metric works fine for them. If in doubt, wait for Phil to release the Python implementation. For me personally, <code>nq_eval</code> works well. </li>\n<li>It’s shared. Even if it were not, you can create your own PyTorch training script, good practice. </li>\n<li>So create your own EDA! I made a post on how to visualize NQ examples. Just exploring questions and answers you can come up with multiple improvements over your base model. </li>\n</ol>\n\n<p>QUEST challenge is for the wider audience, for sure. But this task is more meaningful and exciting imho. Don’t wait for someone to create smth for you, turn on your common sense, and gradually you’ll be progressing. \nPS. Your nick sort of suggests that you love attacking hard tasks without too much supervision. </p>",
      "rawMarkdown": "1. Metric of course is the most annoying part. But some folks report that the metric works fine for them. If in doubt, wait for Phil to release the Python implementation. For me personally, `nq_eval` works well. \n2. It’s shared. Even if it were not, you can create your own PyTorch training script, good practice. \n3. So create your own EDA! I made a post on how to visualize NQ examples. Just exploring questions and answers you can come up with multiple improvements over your base model. \n\nQUEST challenge is for the wider audience, for sure. But this task is more meaningful and exciting imho. Don’t wait for someone to create smth for you, turn on your common sense, and gradually you’ll be progressing. \nPS. Your nick sort of suggests that you love attacking hard tasks without too much supervision.",
      "votes": null
    },
    {
      "id": "696295",
      "postDate": "12/16/2019 12:19:28",
      "content": "<p>There is a pytorch starter here: <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline\">https://www.kaggle.com/sakami/tfqa-pytorch-baseline</a></p>",
      "rawMarkdown": "There is a pytorch starter here: https://www.kaggle.com/sakami/tfqa-pytorch-baseline",
      "votes": null
    },
    {
      "id": "697279",
      "postDate": "12/17/2019 17:12:48",
      "content": "<p>You seem to have a lot of idea about what's wrong with this competition - luckily all those things are fixable. I have recently posted an EDA of the target variable <a href=\"https://www.kaggle.com/rohitagarwal/eda-of-target-variable\">here</a>. It's not the most visually pleasing, but it's my first honest attempt.</p>\n\n<p>Would love to have more people sharing best practices.</p>",
      "rawMarkdown": "You seem to have a lot of idea about what's wrong with this competition - luckily all those things are fixable. I have recently posted an EDA of the target variable [here](https://www.kaggle.com/rohitagarwal/eda-of-target-variable). It's not the most visually pleasing, but it's my first honest attempt.\n\nWould love to have more people sharing best practices.",
      "votes": null
    },
    {
      "id": "697397",
      "postDate": "12/17/2019 20:53:16",
      "content": "<p>Thanks, but I think the inference is not working properly... I corrected my statement to \"Pytorch starter scoring kernel\". Nevertheless, I will look into it and hopefully it will give me some ideas.</p>",
      "rawMarkdown": "Thanks, but I think the inference is not working properly... I corrected my statement to \"Pytorch starter scoring kernel\". Nevertheless, I will look into it and hopefully it will give me some ideas.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 696271,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "12/16/2019 11:28:53",
      "content": "<ol>\n<li>Metric of course is the most annoying part. But some folks report that the metric works fine for them. If in doubt, wait for Phil to release the Python implementation. For me personally, <code>nq_eval</code> works well. </li>\n<li>It’s shared. Even if it were not, you can create your own PyTorch training script, good practice. </li>\n<li>So create your own EDA! I made a post on how to visualize NQ examples. Just exploring questions and answers you can come up with multiple improvements over your base model. </li>\n</ol>\n\n<p>QUEST challenge is for the wider audience, for sure. But this task is more meaningful and exciting imho. Don’t wait for someone to create smth for you, turn on your common sense, and gradually you’ll be progressing. \nPS. Your nick sort of suggests that you love attacking hard tasks without too much supervision. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 696295,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/16/2019 12:19:28",
      "content": "<p>There is a pytorch starter here: <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline\">https://www.kaggle.com/sakami/tfqa-pytorch-baseline</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 697397,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "12/17/2019 20:53:16",
          "content": "<p>Thanks, but I think the inference is not working properly... I corrected my statement to \"Pytorch starter scoring kernel\". Nevertheless, I will look into it and hopefully it will give me some ideas.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 697279,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "12/17/2019 17:12:48",
      "content": "<p>You seem to have a lot of idea about what's wrong with this competition - luckily all those things are fixable. I have recently posted an EDA of the target variable <a href=\"https://www.kaggle.com/rohitagarwal/eda-of-target-variable\">here</a>. It's not the most visually pleasing, but it's my first honest attempt.</p>\n\n<p>Would love to have more people sharing best practices.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "696094": "Hello all, \n\nI have been interested in question answering as a field for a long time, ever since the days people were using memory networks on bAbi datasets! So, seeing the progress the field has made in the last couple of years, I was very excited to participate in this competition and try out the new techniques field. However, here's why I am hesitant to participate a month before the competition ends:\n\n1. The metric still doesn't seem to be fully understood with only a month left in the competition.\n2. There's no PyTorch starter scoring kernel. This is honestly disappointing. Literally any kernel on doing some simple training and inference with a simple BERT model would prove to be useful to many competitors (including myself) I suspect.\n3. Just the overall lack of EDA and good modeling kernels this far into the competition. TFQA currently only has 15 scoring kernels. And those kernels basically are the provided baseline, binary classification baseline, LSTM baseline, and the bert joint baseline (YES/NO kernel being a modification of this). On the other hand, Google QUEST challenge has about 53 scoring kernels. And there are many interesting modeling approaches, from different BERT models, to LSTM baselines, to Universal Sentence Encoder approaches, and even Ridge Regression! I like to take inspiration from existing kernels, and the lack of such kernels (and also the lack of variation) at this stage in the competition is disappointing.\n\nWhat do you guys think? Are these also valid concerns you have had? Are there opportunities for some of these problems to be resolved soon?",
    "696271": "1. Metric of course is the most annoying part. But some folks report that the metric works fine for them. If in doubt, wait for Phil to release the Python implementation. For me personally, `nq_eval` works well. \n2. It’s shared. Even if it were not, you can create your own PyTorch training script, good practice. \n3. So create your own EDA! I made a post on how to visualize NQ examples. Just exploring questions and answers you can come up with multiple improvements over your base model. \n\nQUEST challenge is for the wider audience, for sure. But this task is more meaningful and exciting imho. Don’t wait for someone to create smth for you, turn on your common sense, and gradually you’ll be progressing. \nPS. Your nick sort of suggests that you love attacking hard tasks without too much supervision.",
    "696295": "There is a pytorch starter here: https://www.kaggle.com/sakami/tfqa-pytorch-baseline",
    "697279": "You seem to have a lot of idea about what's wrong with this competition - luckily all those things are fixable. I have recently posted an EDA of the target variable [here](https://www.kaggle.com/rohitagarwal/eda-of-target-variable). It's not the most visually pleasing, but it's my first honest attempt.\n\nWould love to have more people sharing best practices.",
    "697397": "Thanks, but I think the inference is not working properly... I corrected my statement to \"Pytorch starter scoring kernel\". Nevertheless, I will look into it and hopefully it will give me some ideas."
  },
  "source": "meta"
}