{
  "id": 128188,
  "title": "30th Place Solution",
  "url": "/competitions/tensorflow2-question-answering/discussion/128188",
  "author_name": "Pedro Azevedo",
  "post_date": "2020-01-29T13:39:47.380000",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello everyone.</p>\n\n<p>This was my first real experience in kaggle.</p>\n\n<p>Briefly, I want to say that I'm very glad that my team achieved 30th place in this competition. I want to thank my teammates <a href=\"/xiaokangwang\">@xiaokangwang</a>  @xiao-xiao for all the help and work during this competition. Also thanks for all the answers given during the competition explaining everything and giving awesome ideas!</p>\n\n<p>Without further ado, the solution mainly focused on three things:\n-&gt; Fine-tuning Albert Large\n-&gt; DataAugmentation\n-&gt; Understanding thresholds.</p>\n\n<p>The model used was <strong>Albert xLarge version 2</strong> trained on SQUAD2.0 and Fine-tuned on the tiny-dev. One of the important things was the <strong>DataAugmentation</strong> used for training, changing the document_text replacing words by synonyms using WordNet corpus.</p>\n\n<p><img src=\"https://i.imgur.com/IrA2CmI.png\" alt=\"Synonyms Examples\"></p>\n\n<p>In here there wasn't no much of magic. We also tried xxLarge only able to run with a huge doc_stride giving bad results. </p>\n\n<p>Removing html tags helped as well.</p>\n\n<p>Understand the data and the output of the model was an important step. \nThe long_answers appeared 50% of the times\nThe short_answers appeared 35% of the times.</p>\n\n<p><img src=\"https://i.imgur.com/eHTopaB.png\" alt=\"Output score\"></p>\n\n<p>In this regard,  we used threshold values based on the outputs using 50% of all the results for the long_answers e 35% for the short_answers. This was achieve by ordering the list of results, select the number that was in the middle and for the short_answer, 35%. For example, given an output list [1,2,3,4,5] we selected the threshold 3 for the long_answer and 2 for short_answer.</p>\n\n<p>I hope this gives good ideas for future work. </p>\n\n<p>Best Regards,\nPedro Azevedo</p>",
  "messages": [
    {
      "id": 732090,
      "postDate": "2020-01-29T13:39:47.380Z",
      "content": "<p>Hello everyone.</p>\n\n<p>This was my first real experience in kaggle.</p>\n\n<p>Briefly, I want to say that I'm very glad that my team achieved 30th place in this competition. I want to thank my teammates <a href=\"/xiaokangwang\">@xiaokangwang</a>  @xiao-xiao for all the help and work during this competition. Also thanks for all the answers given during the competition explaining everything and giving awesome ideas!</p>\n\n<p>Without further ado, the solution mainly focused on three things:\n-&gt; Fine-tuning Albert Large\n-&gt; DataAugmentation\n-&gt; Understanding thresholds.</p>\n\n<p>The model used was <strong>Albert xLarge version 2</strong> trained on SQUAD2.0 and Fine-tuned on the tiny-dev. One of the important things was the <strong>DataAugmentation</strong> used for training, changing the document_text replacing words by synonyms using WordNet corpus.</p>\n\n<p><img src=\"https://i.imgur.com/IrA2CmI.png\" alt=\"Synonyms Examples\"></p>\n\n<p>In here there wasn't no much of magic. We also tried xxLarge only able to run with a huge doc_stride giving bad results. </p>\n\n<p>Removing html tags helped as well.</p>\n\n<p>Understand the data and the output of the model was an important step. \nThe long_answers appeared 50% of the times\nThe short_answers appeared 35% of the times.</p>\n\n<p><img src=\"https://i.imgur.com/eHTopaB.png\" alt=\"Output score\"></p>\n\n<p>In this regard,  we used threshold values based on the outputs using 50% of all the results for the long_answers e 35% for the short_answers. This was achieve by ordering the list of results, select the number that was in the middle and for the short_answer, 35%. For example, given an output list [1,2,3,4,5] we selected the threshold 3 for the long_answer and 2 for short_answer.</p>\n\n<p>I hope this gives good ideas for future work. </p>\n\n<p>Best Regards,\nPedro Azevedo</p>",
      "rawMarkdown": "Hello everyone.\n\nThis was my first real experience in kaggle.\n\nBriefly, I want to say that I'm very glad that my team achieved 30th place in this competition. I want to thank my teammates @xiaokangwang  @xiao-xiao for all the help and work during this competition. Also thanks for all the answers given during the competition explaining everything and giving awesome ideas!\n\nWithout further ado, the solution mainly focused on three things:\n-&gt; Fine-tuning Albert Large\n-&gt; DataAugmentation\n-&gt; Understanding thresholds.\n\nThe model used was **Albert xLarge version 2** trained on SQUAD2.0 and Fine-tuned on the tiny-dev. One of the important things was the **DataAugmentation** used for training, changing the document_text replacing words by synonyms using WordNet corpus.\n\n![Synonyms Examples](https://i.imgur.com/IrA2CmI.png)\n\nIn here there wasn't no much of magic. We also tried xxLarge only able to run with a huge doc_stride giving bad results. \n\nRemoving html tags helped as well.\n\nUnderstand the data and the output of the model was an important step. \nThe long_answers appeared 50% of the times\nThe short_answers appeared 35% of the times.\n\n![Output score](https://i.imgur.com/eHTopaB.png)\n\nIn this regard,  we used threshold values based on the outputs using 50% of all the results for the long_answers e 35% for the short_answers. This was achieve by ordering the list of results, select the number that was in the middle and for the short_answer, 35%. For example, given an output list [1,2,3,4,5] we selected the threshold 3 for the long_answer and 2 for short_answer.\n\nI hope this gives good ideas for future work. \n\nBest Regards,\nPedro Azevedo",
      "votes": 8
    },
    {
      "id": 732879,
      "postDate": "2020-01-30T12:18:40.023Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": 4
    },
    {
      "id": 733206,
      "postDate": "2020-01-30T20:26:35.663Z",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!"
    },
    {
      "id": 732446,
      "postDate": "2020-01-29T20:45:14.160Z",
      "content": "<p>Sorry, the images were not showing. Fixed the bug! </p>",
      "rawMarkdown": "Sorry, the images were not showing. Fixed the bug! "
    },
    {
      "id": 732350,
      "postDate": "2020-01-29T18:01:58.433Z",
      "content": "<p>Congratulations!!\nThanks for sharing your Approach &amp; Insights <a href=\"/pedrojlazevedo\">@pedrojlazevedo</a> </p>",
      "rawMarkdown": "Congratulations!!\nThanks for sharing your Approach &amp; Insights @pedrojlazevedo "
    },
    {
      "id": 732142,
      "postDate": "2020-01-29T14:30:02.400Z",
      "content": "<p>Thanks for your sharing!\nI have a question about how to train Albert xLarge on SQUAD2.0 ?</p>",
      "rawMarkdown": "Thanks for your sharing!\nI have a question about how to train Albert xLarge on SQUAD2.0 ?",
      "replies": [
        {
          "id": 732154,
          "postDate": "2020-01-29T14:37:10.780Z",
          "content": "<p>I hope that you liked.</p>\n\n<p>You can see in their github:<a href=\"https://github.com/google-research/ALBERT/\">ALBERT</a>\nIn the read.me you can see all the instructions.</p>\n\n<p>Also check the article that the Albert development team did:\n<a href=\"https://arxiv.org/pdf/1909.11942.pdf\">Article</a></p>",
          "rawMarkdown": "I hope that you liked.\n\nYou can see in their github:[ALBERT](https://github.com/google-research/ALBERT/)\nIn the read.me you can see all the instructions.\n\nAlso check the article that the Albert development team did:\n[Article](https://arxiv.org/pdf/1909.11942.pdf)",
          "votes": 2
        },
        {
          "id": 732185,
          "postDate": "2020-01-29T14:51:14.300Z",
          "content": "<p>Thanks for your explanation.\nIt's very helpful for me.</p>",
          "rawMarkdown": "Thanks for your explanation.\nIt's very helpful for me.",
          "votes": 1
        }
      ]
    },
    {
      "id": 732447,
      "postDate": "2020-01-29T20:46:20.427Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 732879,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-30T12:18:40.023000",
      "content": "<p>Congrats!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 733206,
      "author_name": "Manish Nayak",
      "author_url": "",
      "post_date": "2020-01-30T20:26:35.663000",
      "content": "<p>Congratulations!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732446,
      "author_name": "Pedro Azevedo",
      "author_url": "",
      "post_date": "2020-01-29T20:45:14.160000",
      "content": "<p>Sorry, the images were not showing. Fixed the bug! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732350,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2020-01-29T18:01:58.433000",
      "content": "<p>Congratulations!!\nThanks for sharing your Approach &amp; Insights <a href=\"/pedrojlazevedo\">@pedrojlazevedo</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732142,
      "author_name": "Dean",
      "author_url": "",
      "post_date": "2020-01-29T14:30:02.400000",
      "content": "<p>Thanks for your sharing!\nI have a question about how to train Albert xLarge on SQUAD2.0 ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 732154,
          "author_name": "Pedro Azevedo",
          "author_url": "",
          "post_date": "2020-01-29T14:37:10.780000",
          "content": "<p>I hope that you liked.</p>\n\n<p>You can see in their github:<a href=\"https://github.com/google-research/ALBERT/\">ALBERT</a>\nIn the read.me you can see all the instructions.</p>\n\n<p>Also check the article that the Albert development team did:\n<a href=\"https://arxiv.org/pdf/1909.11942.pdf\">Article</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 732185,
          "author_name": "Dean",
          "author_url": "",
          "post_date": "2020-01-29T14:51:14.300000",
          "content": "<p>Thanks for your explanation.\nIt's very helpful for me.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 732447,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-29T20:46:20.427000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "732090": "Hello everyone.\n\nThis was my first real experience in kaggle.\n\nBriefly, I want to say that I'm very glad that my team achieved 30th place in this competition. I want to thank my teammates @xiaokangwang  @xiao-xiao for all the help and work during this competition. Also thanks for all the answers given during the competition explaining everything and giving awesome ideas!\n\nWithout further ado, the solution mainly focused on three things:\n-&gt; Fine-tuning Albert Large\n-&gt; DataAugmentation\n-&gt; Understanding thresholds.\n\nThe model used was **Albert xLarge version 2** trained on SQUAD2.0 and Fine-tuned on the tiny-dev. One of the important things was the **DataAugmentation** used for training, changing the document_text replacing words by synonyms using WordNet corpus.\n\n![Synonyms Examples](https://i.imgur.com/IrA2CmI.png)\n\nIn here there wasn't no much of magic. We also tried xxLarge only able to run with a huge doc_stride giving bad results. \n\nRemoving html tags helped as well.\n\nUnderstand the data and the output of the model was an important step. \nThe long_answers appeared 50% of the times\nThe short_answers appeared 35% of the times.\n\n![Output score](https://i.imgur.com/eHTopaB.png)\n\nIn this regard,  we used threshold values based on the outputs using 50% of all the results for the long_answers e 35% for the short_answers. This was achieve by ordering the list of results, select the number that was in the middle and for the short_answer, 35%. For example, given an output list [1,2,3,4,5] we selected the threshold 3 for the long_answer and 2 for short_answer.\n\nI hope this gives good ideas for future work. \n\nBest Regards,\nPedro Azevedo",
    "732879": "Congrats!",
    "733206": "Congratulations!!",
    "732446": "Sorry, the images were not showing. Fixed the bug! ",
    "732350": "Congratulations!!\nThanks for sharing your Approach &amp; Insights @pedrojlazevedo ",
    "732142": "Thanks for your sharing!\nI have a question about how to train Albert xLarge on SQUAD2.0 ?",
    "732447": ""
  }
}