{
  "id": 79718,
  "title": "RFC on my first serious competition",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79718",
  "author_name": "",
  "post_date": "2019-02-06T21:27:25.048033200Z",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p>It's the first time I dedicated a serious amount of time to a Kaggle competition. I wanted to share a few thoughts about it and ask for comments and suggestions. Since the deadline has passed, I figured it's okay now to publish and discuss my work (and it's not anything near high-scoring anyway).</p>\n\n<p>At first though, I mistakenly tried an all-CNN approach in a notebook format:\n<a href=\"https://www.kaggle.com/asquarek/fork-of-word2vec-cnn\">https://www.kaggle.com/asquarek/fork-of-word2vec-cnn</a></p>\n\n<p>In the end, I ended up mostly using Cuda-backed GRUs because of the blazing fast iteration speed they allowed me, and also switched to a vanilla Python script format:\n<a href=\"https://www.kaggle.com/asquarek/from-scratch\">https://www.kaggle.com/asquarek/from-scratch</a>  (commit 106 was the last one before the deadline)</p>\n\n<p>Looking at these kernels side by side, I can clearly see two mistakes that I made initially:\n1. Using the Jupyter notebook. It encourages very linear programs a la 1950s Fortran, which makes it incredibly hard to manage changes and strategies, and also negates one advantage I have in a competition, namely being a good software engineer. Python scripts work so much better for me for anything except data exploration. The code in the second notebook just spun out of control and after taking a two month break from the competition, I couldn't salvage or even really understand it so I started over.\n2. While 1D CNNs look promising, it's impossible to quickly iterate on a script that takes ~2 hours to complete, so in a sense I had to pick the fastest option available. UPD: when I used both CNN and GRU (commit 109), I got the validation error of 0.69; I guess the lesson is don't leave it until the last few hours to do the work :)</p>\n\n<p>One thing that survived intact was the <code>DataGenerator</code> class. I didn't see anything like it in public kernels, and I found it useful for text processing. Maybe I'm wrong and there's a better way of working with text on a time and memory budget, please tell me if that's the case!</p>\n\n<p>I also found it very useful to use Numba on my data augmentation function, it's so much faster this way! Also, using it made me rewrite the whole thing in the array processing style instead of a bunch of generators. <code>yield</code> is slow, and it's sad.</p>\n\n<p>P.S.: My favorite false negative: \"Why didn't the United States intervene to save Jesus from the Romans?\" How <em>do</em> you teach a model what a joke is? :)</p>",
  "messages": [
    {
      "id": "467314",
      "postDate": "02/06/2019 21:27:25",
      "content": "<p>It's the first time I dedicated a serious amount of time to a Kaggle competition. I wanted to share a few thoughts about it and ask for comments and suggestions. Since the deadline has passed, I figured it's okay now to publish and discuss my work (and it's not anything near high-scoring anyway).</p>\n\n<p>At first though, I mistakenly tried an all-CNN approach in a notebook format:\n<a href=\"https://www.kaggle.com/asquarek/fork-of-word2vec-cnn\">https://www.kaggle.com/asquarek/fork-of-word2vec-cnn</a></p>\n\n<p>In the end, I ended up mostly using Cuda-backed GRUs because of the blazing fast iteration speed they allowed me, and also switched to a vanilla Python script format:\n<a href=\"https://www.kaggle.com/asquarek/from-scratch\">https://www.kaggle.com/asquarek/from-scratch</a>  (commit 106 was the last one before the deadline)</p>\n\n<p>Looking at these kernels side by side, I can clearly see two mistakes that I made initially:\n1. Using the Jupyter notebook. It encourages very linear programs a la 1950s Fortran, which makes it incredibly hard to manage changes and strategies, and also negates one advantage I have in a competition, namely being a good software engineer. Python scripts work so much better for me for anything except data exploration. The code in the second notebook just spun out of control and after taking a two month break from the competition, I couldn't salvage or even really understand it so I started over.\n2. While 1D CNNs look promising, it's impossible to quickly iterate on a script that takes ~2 hours to complete, so in a sense I had to pick the fastest option available. UPD: when I used both CNN and GRU (commit 109), I got the validation error of 0.69; I guess the lesson is don't leave it until the last few hours to do the work :)</p>\n\n<p>One thing that survived intact was the <code>DataGenerator</code> class. I didn't see anything like it in public kernels, and I found it useful for text processing. Maybe I'm wrong and there's a better way of working with text on a time and memory budget, please tell me if that's the case!</p>\n\n<p>I also found it very useful to use Numba on my data augmentation function, it's so much faster this way! Also, using it made me rewrite the whole thing in the array processing style instead of a bunch of generators. <code>yield</code> is slow, and it's sad.</p>\n\n<p>P.S.: My favorite false negative: \"Why didn't the United States intervene to save Jesus from the Romans?\" How <em>do</em> you teach a model what a joke is? :)</p>",
      "rawMarkdown": "It's the first time I dedicated a serious amount of time to a Kaggle competition. I wanted to share a few thoughts about it and ask for comments and suggestions. Since the deadline has passed, I figured it's okay now to publish and discuss my work (and it's not anything near high-scoring anyway).\n\nAt first though, I mistakenly tried an all-CNN approach in a notebook format:\nhttps://www.kaggle.com/asquarek/fork-of-word2vec-cnn\n\nIn the end, I ended up mostly using Cuda-backed GRUs because of the blazing fast iteration speed they allowed me, and also switched to a vanilla Python script format:\nhttps://www.kaggle.com/asquarek/from-scratch  (commit 106 was the last one before the deadline)\n\nLooking at these kernels side by side, I can clearly see two mistakes that I made initially:\n1. Using the Jupyter notebook. It encourages very linear programs a la 1950s Fortran, which makes it incredibly hard to manage changes and strategies, and also negates one advantage I have in a competition, namely being a good software engineer. Python scripts work so much better for me for anything except data exploration. The code in the second notebook just spun out of control and after taking a two month break from the competition, I couldn't salvage or even really understand it so I started over.\n2. While 1D CNNs look promising, it's impossible to quickly iterate on a script that takes ~2 hours to complete, so in a sense I had to pick the fastest option available. UPD: when I used both CNN and GRU (commit 109), I got the validation error of 0.69; I guess the lesson is don't leave it until the last few hours to do the work :)\n\nOne thing that survived intact was the `DataGenerator` class. I didn't see anything like it in public kernels, and I found it useful for text processing. Maybe I'm wrong and there's a better way of working with text on a time and memory budget, please tell me if that's the case!\n\nI also found it very useful to use Numba on my data augmentation function, it's so much faster this way! Also, using it made me rewrite the whole thing in the array processing style instead of a bunch of generators. `yield` is slow, and it's sad.\n\nP.S.: My favorite false negative: \"Why didn't the United States intervene to save Jesus from the Romans?\" How *do* you teach a model what a joke is? :)",
      "votes": null
    },
    {
      "id": "468147",
      "postDate": "02/08/2019 11:10:33",
      "content": "<p>well written</p>",
      "rawMarkdown": "well written",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 468147,
      "author_name": "econdata",
      "author_url": "",
      "post_date": "02/08/2019 11:10:33",
      "content": "<p>well written</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "467314": "It's the first time I dedicated a serious amount of time to a Kaggle competition. I wanted to share a few thoughts about it and ask for comments and suggestions. Since the deadline has passed, I figured it's okay now to publish and discuss my work (and it's not anything near high-scoring anyway).\n\nAt first though, I mistakenly tried an all-CNN approach in a notebook format:\nhttps://www.kaggle.com/asquarek/fork-of-word2vec-cnn\n\nIn the end, I ended up mostly using Cuda-backed GRUs because of the blazing fast iteration speed they allowed me, and also switched to a vanilla Python script format:\nhttps://www.kaggle.com/asquarek/from-scratch  (commit 106 was the last one before the deadline)\n\nLooking at these kernels side by side, I can clearly see two mistakes that I made initially:\n1. Using the Jupyter notebook. It encourages very linear programs a la 1950s Fortran, which makes it incredibly hard to manage changes and strategies, and also negates one advantage I have in a competition, namely being a good software engineer. Python scripts work so much better for me for anything except data exploration. The code in the second notebook just spun out of control and after taking a two month break from the competition, I couldn't salvage or even really understand it so I started over.\n2. While 1D CNNs look promising, it's impossible to quickly iterate on a script that takes ~2 hours to complete, so in a sense I had to pick the fastest option available. UPD: when I used both CNN and GRU (commit 109), I got the validation error of 0.69; I guess the lesson is don't leave it until the last few hours to do the work :)\n\nOne thing that survived intact was the `DataGenerator` class. I didn't see anything like it in public kernels, and I found it useful for text processing. Maybe I'm wrong and there's a better way of working with text on a time and memory budget, please tell me if that's the case!\n\nI also found it very useful to use Numba on my data augmentation function, it's so much faster this way! Also, using it made me rewrite the whole thing in the array processing style instead of a bunch of generators. `yield` is slow, and it's sad.\n\nP.S.: My favorite false negative: \"Why didn't the United States intervene to save Jesus from the Romans?\" How *do* you teach a model what a joke is? :)",
    "468147": "well written"
  },
  "source": "meta"
}