{
  "id": 127258,
  "title": "45th solution, my journey and learnings, feeling grateful :) ",
  "url": "/competitions/tensorflow2-question-answering/writeups/good-answer-45th-solution-my-journey-and-learnings",
  "author_name": "",
  "post_date": "2020-01-23T04:52:10.235026800Z",
  "votes": 14,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I woke up at 4am today, as most days in the last weeks of this competition. The last time I checked yesterday I was 54th on the public LB, and I really counted on a medal and progressing to Kaggle Expert with this competition. I did my daily quick workout and finally opened up Kaggle to see that I moved a few places up and stayed in the silver zone. I smiled, relaxed, and decided to write up my learnings immediately. </p>\n\n<p>My solution is very simple:\n1. Started with the bert-joint kernel by prvi (thank you @prokaj)\n2. Trained a new bert-joint model, starting with bert-large-uncased (I believe the whole-word-masking gave it a 1-2 points boost over the available bert-joint checkpoint), finetuning 1 epoch on SQUAD, then finetuning 1 epoch (lr=3e-5) on the NQ dataset. I tried several other settings, with lr between 1e5 and 3e5, 1-3 epochs, but the initial setting worked best. \n3. I did extensive validation on the NQ dev set, compared the model outputs with ground truth (using a modified version of NQ browser), and used those insights to set the post-processing thresholds. </p>\n\n<p>Things I wished to do - I spent quite some time trying to put another layer (bi-LSTM) over the output features of bert-joint, to learn the post-processing rules rather than setting them by hand. In the end, this turned out to take more time that I could afford between work and family, so I dropped the idea. I’m looking forward to see the winning solutions, see if they implemented this and learn from them. </p>\n\n<p>Before I share my learnings, some context. I’ve been working in IT for 15 years, doing various roles across project and product management, operations and consulting, but no coding / data science work. I initially got interested in ML 2 years ago with Andrew Ng courses (signed up to Kaggle for the first time then), but haven’t really done anything practical until a few months ago when I discovered fast.ai. I did the part 1 of fast.ai Deep Learning course and came to Kaggle to practice the skills. Thank you @jhoward for the learnings and the motivation!</p>\n\n<p>My learnings:\n1. It’s ok to be overwhelmed. I said on many nights to my wife that this thing is too difficult for me… then I woke up in the morning, reviewed each line of code to understand the inputs/outputs, analyzed the errors, and came up with a solution. \n2. A little time every day is better than nothing. I have a full time job, wife and 2.5 year old daughter… I started doing Kaggle in the evenings, once my girls went to sleep, but after trying to get my daughter to sleep for 1-2 hours I had no energy left for coding… then I switched to going to sleep early and waking up early, and with 1-2 hours per day I felt like I can learn and make progress. \n3. The ML/Kaggle community is amazing! My go-to places for learning are the fast.ai forums, Kaggle discussion and ML Twitter. It’s amazing how open this community is, how much learning and sharing is going on. Thank you!!!</p>\n\n<p>With this, I’d like to express my gratitude to Kaggle and Google for organizing this competition and providing the TPU credits. Thank you to the Kaggle community (especially @prokaj, @kashitsky, @christofhenkel, @yihdarshieh) for sharing your code and insights, it’s amazing to be able to learn from so many talented people. And congratulations to the winners, medalists, and everyone that learned something during this competition!</p>\n\n<p>Last thing - I’ve done only solo competitions so far, but I’m looking forward to find partners for future competitions. If you’d like to team up in the future, please connect with me at darek.kleczek@gmail.com :) </p>",
  "messages": [
    {
      "id": "726567",
      "postDate": "01/23/2020 04:52:10",
      "content": "<p>I woke up at 4am today, as most days in the last weeks of this competition. The last time I checked yesterday I was 54th on the public LB, and I really counted on a medal and progressing to Kaggle Expert with this competition. I did my daily quick workout and finally opened up Kaggle to see that I moved a few places up and stayed in the silver zone. I smiled, relaxed, and decided to write up my learnings immediately. </p>\n\n<p>My solution is very simple:\n1. Started with the bert-joint kernel by prvi (thank you @prokaj)\n2. Trained a new bert-joint model, starting with bert-large-uncased (I believe the whole-word-masking gave it a 1-2 points boost over the available bert-joint checkpoint), finetuning 1 epoch on SQUAD, then finetuning 1 epoch (lr=3e-5) on the NQ dataset. I tried several other settings, with lr between 1e5 and 3e5, 1-3 epochs, but the initial setting worked best. \n3. I did extensive validation on the NQ dev set, compared the model outputs with ground truth (using a modified version of NQ browser), and used those insights to set the post-processing thresholds. </p>\n\n<p>Things I wished to do - I spent quite some time trying to put another layer (bi-LSTM) over the output features of bert-joint, to learn the post-processing rules rather than setting them by hand. In the end, this turned out to take more time that I could afford between work and family, so I dropped the idea. I’m looking forward to see the winning solutions, see if they implemented this and learn from them. </p>\n\n<p>Before I share my learnings, some context. I’ve been working in IT for 15 years, doing various roles across project and product management, operations and consulting, but no coding / data science work. I initially got interested in ML 2 years ago with Andrew Ng courses (signed up to Kaggle for the first time then), but haven’t really done anything practical until a few months ago when I discovered fast.ai. I did the part 1 of fast.ai Deep Learning course and came to Kaggle to practice the skills. Thank you @jhoward for the learnings and the motivation!</p>\n\n<p>My learnings:\n1. It’s ok to be overwhelmed. I said on many nights to my wife that this thing is too difficult for me… then I woke up in the morning, reviewed each line of code to understand the inputs/outputs, analyzed the errors, and came up with a solution. \n2. A little time every day is better than nothing. I have a full time job, wife and 2.5 year old daughter… I started doing Kaggle in the evenings, once my girls went to sleep, but after trying to get my daughter to sleep for 1-2 hours I had no energy left for coding… then I switched to going to sleep early and waking up early, and with 1-2 hours per day I felt like I can learn and make progress. \n3. The ML/Kaggle community is amazing! My go-to places for learning are the fast.ai forums, Kaggle discussion and ML Twitter. It’s amazing how open this community is, how much learning and sharing is going on. Thank you!!!</p>\n\n<p>With this, I’d like to express my gratitude to Kaggle and Google for organizing this competition and providing the TPU credits. Thank you to the Kaggle community (especially @prokaj, @kashitsky, @christofhenkel, @yihdarshieh) for sharing your code and insights, it’s amazing to be able to learn from so many talented people. And congratulations to the winners, medalists, and everyone that learned something during this competition!</p>\n\n<p>Last thing - I’ve done only solo competitions so far, but I’m looking forward to find partners for future competitions. If you’d like to team up in the future, please connect with me at darek.kleczek@gmail.com :) </p>",
      "rawMarkdown": "I woke up at 4am today, as most days in the last weeks of this competition. The last time I checked yesterday I was 54th on the public LB, and I really counted on a medal and progressing to Kaggle Expert with this competition. I did my daily quick workout and finally opened up Kaggle to see that I moved a few places up and stayed in the silver zone. I smiled, relaxed, and decided to write up my learnings immediately. \n\nMy solution is very simple:\n1. Started with the bert-joint kernel by prvi (thank you @prokaj)\n2. Trained a new bert-joint model, starting with bert-large-uncased (I believe the whole-word-masking gave it a 1-2 points boost over the available bert-joint checkpoint), finetuning 1 epoch on SQUAD, then finetuning 1 epoch (lr=3e-5) on the NQ dataset. I tried several other settings, with lr between 1e5 and 3e5, 1-3 epochs, but the initial setting worked best. \n3. I did extensive validation on the NQ dev set, compared the model outputs with ground truth (using a modified version of NQ browser), and used those insights to set the post-processing thresholds. \n\nThings I wished to do - I spent quite some time trying to put another layer (bi-LSTM) over the output features of bert-joint, to learn the post-processing rules rather than setting them by hand. In the end, this turned out to take more time that I could afford between work and family, so I dropped the idea. I’m looking forward to see the winning solutions, see if they implemented this and learn from them. \n\nBefore I share my learnings, some context. I’ve been working in IT for 15 years, doing various roles across project and product management, operations and consulting, but no coding / data science work. I initially got interested in ML 2 years ago with Andrew Ng courses (signed up to Kaggle for the first time then), but haven’t really done anything practical until a few months ago when I discovered fast.ai. I did the part 1 of fast.ai Deep Learning course and came to Kaggle to practice the skills. Thank you @jhoward for the learnings and the motivation!\n\nMy learnings:\n1. It’s ok to be overwhelmed. I said on many nights to my wife that this thing is too difficult for me… then I woke up in the morning, reviewed each line of code to understand the inputs/outputs, analyzed the errors, and came up with a solution. \n2. A little time every day is better than nothing. I have a full time job, wife and 2.5 year old daughter… I started doing Kaggle in the evenings, once my girls went to sleep, but after trying to get my daughter to sleep for 1-2 hours I had no energy left for coding… then I switched to going to sleep early and waking up early, and with 1-2 hours per day I felt like I can learn and make progress. \n3. The ML/Kaggle community is amazing! My go-to places for learning are the fast.ai forums, Kaggle discussion and ML Twitter. It’s amazing how open this community is, how much learning and sharing is going on. Thank you!!!\n\nWith this, I’d like to express my gratitude to Kaggle and Google for organizing this competition and providing the TPU credits. Thank you to the Kaggle community (especially @prokaj, @kashitsky, @christofhenkel, @yihdarshieh) for sharing your code and insights, it’s amazing to be able to learn from so many talented people. And congratulations to the winners, medalists, and everyone that learned something during this competition!\n\nLast thing - I’ve done only solo competitions so far, but I’m looking forward to find partners for future competitions. If you’d like to team up in the future, please connect with me at darek.kleczek@gmail.com :)",
      "votes": null
    },
    {
      "id": "726648",
      "postDate": "01/23/2020 06:28:42",
      "content": "<p>Congratulations\nNice Write-Up\nThanks for Sharing your Approach &amp; Insights!! <a href=\"/thedrcat\">@thedrcat</a> </p>",
      "rawMarkdown": "Congratulations\nNice Write-Up\nThanks for Sharing your Approach &amp; Insights!! @thedrcat",
      "votes": null
    },
    {
      "id": "727029",
      "postDate": "01/23/2020 11:49:38",
      "content": "<p>Inspirational Post. \nI am new to TPU. How is pytorch TPU pipeline in comparsion to TF TPU?</p>",
      "rawMarkdown": "Inspirational Post. \nI am new to TPU. How is pytorch TPU pipeline in comparsion to TF TPU?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 726648,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "01/23/2020 06:28:42",
      "content": "<p>Congratulations\nNice Write-Up\nThanks for Sharing your Approach &amp; Insights!! <a href=\"/thedrcat\">@thedrcat</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 727029,
      "author_name": "shayekh",
      "author_url": "",
      "post_date": "01/23/2020 11:49:38",
      "content": "<p>Inspirational Post. \nI am new to TPU. How is pytorch TPU pipeline in comparsion to TF TPU?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "726567": "I woke up at 4am today, as most days in the last weeks of this competition. The last time I checked yesterday I was 54th on the public LB, and I really counted on a medal and progressing to Kaggle Expert with this competition. I did my daily quick workout and finally opened up Kaggle to see that I moved a few places up and stayed in the silver zone. I smiled, relaxed, and decided to write up my learnings immediately. \n\nMy solution is very simple:\n1. Started with the bert-joint kernel by prvi (thank you @prokaj)\n2. Trained a new bert-joint model, starting with bert-large-uncased (I believe the whole-word-masking gave it a 1-2 points boost over the available bert-joint checkpoint), finetuning 1 epoch on SQUAD, then finetuning 1 epoch (lr=3e-5) on the NQ dataset. I tried several other settings, with lr between 1e5 and 3e5, 1-3 epochs, but the initial setting worked best. \n3. I did extensive validation on the NQ dev set, compared the model outputs with ground truth (using a modified version of NQ browser), and used those insights to set the post-processing thresholds. \n\nThings I wished to do - I spent quite some time trying to put another layer (bi-LSTM) over the output features of bert-joint, to learn the post-processing rules rather than setting them by hand. In the end, this turned out to take more time that I could afford between work and family, so I dropped the idea. I’m looking forward to see the winning solutions, see if they implemented this and learn from them. \n\nBefore I share my learnings, some context. I’ve been working in IT for 15 years, doing various roles across project and product management, operations and consulting, but no coding / data science work. I initially got interested in ML 2 years ago with Andrew Ng courses (signed up to Kaggle for the first time then), but haven’t really done anything practical until a few months ago when I discovered fast.ai. I did the part 1 of fast.ai Deep Learning course and came to Kaggle to practice the skills. Thank you @jhoward for the learnings and the motivation!\n\nMy learnings:\n1. It’s ok to be overwhelmed. I said on many nights to my wife that this thing is too difficult for me… then I woke up in the morning, reviewed each line of code to understand the inputs/outputs, analyzed the errors, and came up with a solution. \n2. A little time every day is better than nothing. I have a full time job, wife and 2.5 year old daughter… I started doing Kaggle in the evenings, once my girls went to sleep, but after trying to get my daughter to sleep for 1-2 hours I had no energy left for coding… then I switched to going to sleep early and waking up early, and with 1-2 hours per day I felt like I can learn and make progress. \n3. The ML/Kaggle community is amazing! My go-to places for learning are the fast.ai forums, Kaggle discussion and ML Twitter. It’s amazing how open this community is, how much learning and sharing is going on. Thank you!!!\n\nWith this, I’d like to express my gratitude to Kaggle and Google for organizing this competition and providing the TPU credits. Thank you to the Kaggle community (especially @prokaj, @kashitsky, @christofhenkel, @yihdarshieh) for sharing your code and insights, it’s amazing to be able to learn from so many talented people. And congratulations to the winners, medalists, and everyone that learned something during this competition!\n\nLast thing - I’ve done only solo competitions so far, but I’m looking forward to find partners for future competitions. If you’d like to team up in the future, please connect with me at darek.kleczek@gmail.com :)",
    "726648": "Congratulations\nNice Write-Up\nThanks for Sharing your Approach &amp; Insights!! @thedrcat",
    "727029": "Inspirational Post. \nI am new to TPU. How is pytorch TPU pipeline in comparsion to TF TPU?"
  },
  "source": "meta"
}