{
  "id": 160853,
  "title": "From zero to two silver medals: some organizational tips",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/160853",
  "author_name": "",
  "post_date": "2020-06-23T00:17:16.518515400Z",
  "votes": 25,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Since this competition is now finished, I'd like to first thank all my great teammates <a href=\"/ibtesama\">@ibtesama</a> <a href=\"/shahules\">@shahules</a> and my boy <a href=\"/tanulsingh077\">@tanulsingh077</a> . It's been great working with you, and I wish you all the best for your future competitions.</p>\n\n<p>Second, congratulations to everybody for the efforts you have put in this competition for the last 3 months and congratulations to all the medal winners.</p>\n\n<p>As there are already many great write-ups about winning solutions for this competition, I wanted to share with you four aspects that I find critical to succeed in a Kaggle competition. <strong>They have made the difference for me between a competition with and without medal</strong>. This is my way to tackle a Kaggle problem. <strong>It is obviously not perfect and I'm open to criticism as long as it is constructive.</strong></p>\n\n<p>Unlike other recommendations, those ones are not <strong>technical but rather organizational.</strong> I've applied this framework with my teammates for this competition and the Tweet sentiment competition where I got two silver medals back to back.</p>\n\n<p>When tackling a Kaggle competition, you should always have your <strong>own framework</strong>. This is mine, which I have built throughout those competitions. There's no doubt it will be different in 6 months, I encourage you to continuously improve it. </p>\n\n<p>First, read an EDA or do your own analysis. After doing your analysis of the data, framing the problem, trying to think of a similar competition you did, throw your ideas on a note. Don't forget to have a look at other winning solutions, to try the same ideas as them. At this stage, you should have an unsorted list of ideas to try: preprocessing, model architecture, blending ideas...</p>\n\n<p>Second, now you have a list of ideas to try, you need to sort them out. My main criterion is the computational time. <strong>You want to iterate as fast as you can to quickly have a fully-running pipeline</strong>. That's why I would recommend starting with the ideas that don't add too much computational time to your pipeline. Hence, you should start first with preprocessing methods (cleaning for text...), sampling techniques (random weighted sampler...) and then data augmentation techniques. You can then add exotic model architectures (from the simplest to the more complex). The key message here is to iterate quickly, try with a simple architecture at first. If using Transformers, use BERT-base or even DistilBERT. For a computer vision problem, use ResNet34 or EfficientNetB0. As you can see in the picture below, I have begun with a simpler model bert-multilingual-uncased before choosing a more complex one like XLM-Roberta. It has saved me time, especially like me when you have no other GPU/TPU resources than the ones offered by Kaggle (thanks to them!).</p>\n\n<p>Third, and <strong>it is probably the most important idea,</strong> *<em>report everything on a table sheet</em>*. You need to track everything from your experiments. Obviously make sure to seed everything as you want your results to be comparable. This is (a part of) the table I used for Jigsaw:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1302247%2F6ecd271b6272b3785d3e5825dad23230%2FScreen_Shot_2020-06-23_at_12.02.50_AM.png?generation=1592871307292789&amp;alt=media\" alt=\"\"></p>\n\n<p>Fourth, before starting any modeling, make sure to build a solid validation strategy. <a href=\"/cdeotte\">@cdeotte</a>  and <a href=\"/cpmpml\">@cpmpml</a> have given some tips on the discussion thread of Tweet sentiment extraction. You must make sure to have reliable CV strategy, otherwise at best the final stages of the competition will be a tense moment for you as you fear shake-up and at worst you are out of the medal zone after having worked multiple weeks on a problem. Furthermore, in SIIM-ISIC melanoma competition, I just started computing my CV statistics: mean and standard deviation. One technique is to add as many folds as you need to reduce the standard deviation of your CV scores above some threshold. Another one pointed out by <a href=\"/cdeotte\">@cdeotte</a> is to use different seeds. That way you'll have a reliable CV strategy that you can trust when you need to choose between submissions.</p>\n\n<p>I hope this quick message helped you. I think it is critical to have a clear-defined structure to follow and is an invaluable perk to have as a data scientist / ML engineer in your daily job. </p>\n\n<p>All the best,</p>\n\n<p>Pierre-Antoine</p>",
  "messages": [
    {
      "id": "897531",
      "postDate": "06/23/2020 00:17:16",
      "content": "<p>Since this competition is now finished, I'd like to first thank all my great teammates <a href=\"/ibtesama\">@ibtesama</a> <a href=\"/shahules\">@shahules</a> and my boy <a href=\"/tanulsingh077\">@tanulsingh077</a> . It's been great working with you, and I wish you all the best for your future competitions.</p>\n\n<p>Second, congratulations to everybody for the efforts you have put in this competition for the last 3 months and congratulations to all the medal winners.</p>\n\n<p>As there are already many great write-ups about winning solutions for this competition, I wanted to share with you four aspects that I find critical to succeed in a Kaggle competition. <strong>They have made the difference for me between a competition with and without medal</strong>. This is my way to tackle a Kaggle problem. <strong>It is obviously not perfect and I'm open to criticism as long as it is constructive.</strong></p>\n\n<p>Unlike other recommendations, those ones are not <strong>technical but rather organizational.</strong> I've applied this framework with my teammates for this competition and the Tweet sentiment competition where I got two silver medals back to back.</p>\n\n<p>When tackling a Kaggle competition, you should always have your <strong>own framework</strong>. This is mine, which I have built throughout those competitions. There's no doubt it will be different in 6 months, I encourage you to continuously improve it. </p>\n\n<p>First, read an EDA or do your own analysis. After doing your analysis of the data, framing the problem, trying to think of a similar competition you did, throw your ideas on a note. Don't forget to have a look at other winning solutions, to try the same ideas as them. At this stage, you should have an unsorted list of ideas to try: preprocessing, model architecture, blending ideas...</p>\n\n<p>Second, now you have a list of ideas to try, you need to sort them out. My main criterion is the computational time. <strong>You want to iterate as fast as you can to quickly have a fully-running pipeline</strong>. That's why I would recommend starting with the ideas that don't add too much computational time to your pipeline. Hence, you should start first with preprocessing methods (cleaning for text...), sampling techniques (random weighted sampler...) and then data augmentation techniques. You can then add exotic model architectures (from the simplest to the more complex). The key message here is to iterate quickly, try with a simple architecture at first. If using Transformers, use BERT-base or even DistilBERT. For a computer vision problem, use ResNet34 or EfficientNetB0. As you can see in the picture below, I have begun with a simpler model bert-multilingual-uncased before choosing a more complex one like XLM-Roberta. It has saved me time, especially like me when you have no other GPU/TPU resources than the ones offered by Kaggle (thanks to them!).</p>\n\n<p>Third, and <strong>it is probably the most important idea,</strong> *<em>report everything on a table sheet</em>*. You need to track everything from your experiments. Obviously make sure to seed everything as you want your results to be comparable. This is (a part of) the table I used for Jigsaw:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1302247%2F6ecd271b6272b3785d3e5825dad23230%2FScreen_Shot_2020-06-23_at_12.02.50_AM.png?generation=1592871307292789&amp;alt=media\" alt=\"\"></p>\n\n<p>Fourth, before starting any modeling, make sure to build a solid validation strategy. <a href=\"/cdeotte\">@cdeotte</a>  and <a href=\"/cpmpml\">@cpmpml</a> have given some tips on the discussion thread of Tweet sentiment extraction. You must make sure to have reliable CV strategy, otherwise at best the final stages of the competition will be a tense moment for you as you fear shake-up and at worst you are out of the medal zone after having worked multiple weeks on a problem. Furthermore, in SIIM-ISIC melanoma competition, I just started computing my CV statistics: mean and standard deviation. One technique is to add as many folds as you need to reduce the standard deviation of your CV scores above some threshold. Another one pointed out by <a href=\"/cdeotte\">@cdeotte</a> is to use different seeds. That way you'll have a reliable CV strategy that you can trust when you need to choose between submissions.</p>\n\n<p>I hope this quick message helped you. I think it is critical to have a clear-defined structure to follow and is an invaluable perk to have as a data scientist / ML engineer in your daily job. </p>\n\n<p>All the best,</p>\n\n<p>Pierre-Antoine</p>",
      "rawMarkdown": "Since this competition is now finished, I'd like to first thank all my great teammates @ibtesama @shahules and my boy @tanulsingh077 . It's been great working with you, and I wish you all the best for your future competitions.\n\nSecond, congratulations to everybody for the efforts you have put in this competition for the last 3 months and congratulations to all the medal winners.\n\nAs there are already many great write-ups about winning solutions for this competition, I wanted to share with you four aspects that I find critical to succeed in a Kaggle competition. **They have made the difference for me between a competition with and without medal**. This is my way to tackle a Kaggle problem. **It is obviously not perfect and I'm open to criticism as long as it is constructive.**\n\nUnlike other recommendations, those ones are not **technical but rather organizational.** I've applied this framework with my teammates for this competition and the Tweet sentiment competition where I got two silver medals back to back.\n\nWhen tackling a Kaggle competition, you should always have your **own framework**. This is mine, which I have built throughout those competitions. There's no doubt it will be different in 6 months, I encourage you to continuously improve it. \n\nFirst, read an EDA or do your own analysis. After doing your analysis of the data, framing the problem, trying to think of a similar competition you did, throw your ideas on a note. Don't forget to have a look at other winning solutions, to try the same ideas as them. At this stage, you should have an unsorted list of ideas to try: preprocessing, model architecture, blending ideas...\n\nSecond, now you have a list of ideas to try, you need to sort them out. My main criterion is the computational time. **You want to iterate as fast as you can to quickly have a fully-running pipeline**. That's why I would recommend starting with the ideas that don't add too much computational time to your pipeline. Hence, you should start first with preprocessing methods (cleaning for text...), sampling techniques (random weighted sampler...) and then data augmentation techniques. You can then add exotic model architectures (from the simplest to the more complex). The key message here is to iterate quickly, try with a simple architecture at first. If using Transformers, use BERT-base or even DistilBERT. For a computer vision problem, use ResNet34 or EfficientNetB0. As you can see in the picture below, I have begun with a simpler model bert-multilingual-uncased before choosing a more complex one like XLM-Roberta. It has saved me time, especially like me when you have no other GPU/TPU resources than the ones offered by Kaggle (thanks to them!).\n\nThird, and **it is probably the most important idea,** **report everything on a table sheet**. You need to track everything from your experiments. Obviously make sure to seed everything as you want your results to be comparable. This is (a part of) the table I used for Jigsaw:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1302247%2F6ecd271b6272b3785d3e5825dad23230%2FScreen_Shot_2020-06-23_at_12.02.50_AM.png?generation=1592871307292789&amp;alt=media)\n\nFourth, before starting any modeling, make sure to build a solid validation strategy. @cdeotte  and @cpmpml have given some tips on the discussion thread of Tweet sentiment extraction. You must make sure to have reliable CV strategy, otherwise at best the final stages of the competition will be a tense moment for you as you fear shake-up and at worst you are out of the medal zone after having worked multiple weeks on a problem. Furthermore, in SIIM-ISIC melanoma competition, I just started computing my CV statistics: mean and standard deviation. One technique is to add as many folds as you need to reduce the standard deviation of your CV scores above some threshold. Another one pointed out by @cdeotte is to use different seeds. That way you'll have a reliable CV strategy that you can trust when you need to choose between submissions.\n\nI hope this quick message helped you. I think it is critical to have a clear-defined structure to follow and is an invaluable perk to have as a data scientist / ML engineer in your daily job. \n\nAll the best,\n\nPierre-Antoine",
      "votes": null
    },
    {
      "id": "897604",
      "postDate": "06/23/2020 01:43:35",
      "content": "<p>Congratulations Pierre</p>",
      "rawMarkdown": "Congratulations Pierre",
      "votes": null
    },
    {
      "id": "897610",
      "postDate": "06/23/2020 01:49:46",
      "content": "<p>Any preferred tool for tracking experiments? There is mention of a tool called mag there but I have not tried it. \n<a href=\"https://www.kaggle.com/c/google-quest-challenge/discussion/129840#742757\">https://www.kaggle.com/c/google-quest-challenge/discussion/129840#742757</a></p>\n\n<p>In my experience tracking in a spreadsheet requires a lot of discipline and lacks flexibility, hence I was wondering about more automated tools. But maybe free text like this best fits the needs.</p>",
      "rawMarkdown": "Any preferred tool for tracking experiments? There is mention of a tool called mag there but I have not tried it. \nhttps://www.kaggle.com/c/google-quest-challenge/discussion/129840#742757\n\nIn my experience tracking in a spreadsheet requires a lot of discipline and lacks flexibility, hence I was wondering about more automated tools. But maybe free text like this best fits the needs.",
      "votes": null
    },
    {
      "id": "897763",
      "postDate": "06/23/2020 04:42:48",
      "content": "<p>Congratulations <a href=\"/rftexas\">@rftexas</a> ! Learned a lot from you!</p>",
      "rawMarkdown": "Congratulations @rftexas ! Learned a lot from you!",
      "votes": null
    },
    {
      "id": "898011",
      "postDate": "06/23/2020 08:22:39",
      "content": "<p>I'm using a spreadsheet either on Excel, I'm trying a note-taking app called Notion. That works quite well.</p>",
      "rawMarkdown": "I'm using a spreadsheet either on Excel, I'm trying a note-taking app called Notion. That works quite well.",
      "votes": null
    },
    {
      "id": "898126",
      "postDate": "06/23/2020 10:23:37",
      "content": "<p>Very nice advise :) </p>",
      "rawMarkdown": "Very nice advise :)",
      "votes": null
    },
    {
      "id": "898314",
      "postDate": "06/23/2020 12:48:11",
      "content": "<p>Much appreciated! Thanks!</p>",
      "rawMarkdown": "Much appreciated! Thanks!",
      "votes": null
    },
    {
      "id": "898808",
      "postDate": "06/23/2020 19:06:02",
      "content": "<p>Need of the hour!! <a href=\"/rftexas\">@rftexas</a> </p>",
      "rawMarkdown": "Need of the hour!! @rftexas",
      "votes": null
    },
    {
      "id": "899373",
      "postDate": "06/24/2020 07:39:42",
      "content": "<p>Noted thanks!</p>",
      "rawMarkdown": "Noted thanks!",
      "votes": null
    },
    {
      "id": "899823",
      "postDate": "06/24/2020 13:16:18",
      "content": "<p>congrats <a href=\"/rftexas\">@rftexas</a>  ^^</p>",
      "rawMarkdown": "congrats @rftexas  ^^",
      "votes": null
    },
    {
      "id": "915186",
      "postDate": "07/04/2020 14:26:44",
      "content": "<p>Hi <a href=\"/rftexas\">@rftexas</a>, could you elaborate a bit a why and how using different seeds makes for a reliable CV strategy ? thanks a lot for the writeup !</p>",
      "rawMarkdown": "Hi @rftexas, could you elaborate a bit a why and how using different seeds makes for a reliable CV strategy ? thanks a lot for the writeup !",
      "votes": null
    },
    {
      "id": "915335",
      "postDate": "07/04/2020 16:45:39",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "915340",
      "postDate": "07/04/2020 16:50:07",
      "content": "<p>Sure! One way to make your CV strategy more reliable is to use the same number of samples, say 100k samples divided into 5 folds. You can evaluate the performance using k-fold. But you can also use the same 100k samples using another seed to create your folds. Do that 3 or 4 times and you'll have validated on 400k samples. </p>\n\n<p>Of course, I anticipate your remark saying that the k-fold scores you'll get using different seeds will be correlated since those are essentially the same samples fed to the model even though they are fed in a different order. But it has been empirically shown that this technique works quite well. I've tried it and you can come up with a CV correlated to LB. <a href=\"/cdeotte\">@cdeotte</a> wrote on it on a previous post, I recommend you checking this post as it was very instructive.</p>",
      "rawMarkdown": "Sure! One way to make your CV strategy more reliable is to use the same number of samples, say 100k samples divided into 5 folds. You can evaluate the performance using k-fold. But you can also use the same 100k samples using another seed to create your folds. Do that 3 or 4 times and you'll have validated on 400k samples. \n\nOf course, I anticipate your remark saying that the k-fold scores you'll get using different seeds will be correlated since those are essentially the same samples fed to the model even though they are fed in a different order. But it has been empirically shown that this technique works quite well. I've tried it and you can come up with a CV correlated to LB. @cdeotte wrote on it on a previous post, I recommend you checking this post as it was very instructive.",
      "votes": null
    },
    {
      "id": "915935",
      "postDate": "07/05/2020 08:11:47",
      "content": "<p>Thanks for taking the time to answer. I never considered it that way ! I'll try to find his post... is it in this competition discussion as well ? </p>",
      "rawMarkdown": "Thanks for taking the time to answer. I never considered it that way ! I'll try to find his post... is it in this competition discussion as well ?",
      "votes": null
    },
    {
      "id": "988903",
      "postDate": "08/28/2020 11:50:57",
      "content": "<p><a href=\"https://www.kaggle.com/rftexas\" target=\"_blank\">@rftexas</a>  Really great advice  . Can you suggest some \" When tackling a Kaggle competition, you should always have your own framework \" . Would you elaborate a bit on this  . <br>\nIf possible Could you share your pipeline for this competition ?</p>",
      "rawMarkdown": "rftexas  Really great advice  . Can you suggest some \" When tackling a Kaggle competition, you should always have your own framework \" . Would you elaborate a bit on this  . \nIf possible Could you share your pipeline for this competition ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 988903,
      "author_name": "swarajshinde",
      "author_url": "",
      "post_date": "08/28/2020 11:50:57",
      "content": "<p><a href=\"https://www.kaggle.com/rftexas\" target=\"_blank\">@rftexas</a>  Really great advice  . Can you suggest some \" When tackling a Kaggle competition, you should always have your own framework \" . Would you elaborate a bit on this  . <br>\nIf possible Could you share your pipeline for this competition ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 897604,
      "author_name": "mehdimabrouki",
      "author_url": "",
      "post_date": "06/23/2020 01:43:35",
      "content": "<p>Congratulations Pierre</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 897610,
      "author_name": "sebastienm",
      "author_url": "",
      "post_date": "06/23/2020 01:49:46",
      "content": "<p>Any preferred tool for tracking experiments? There is mention of a tool called mag there but I have not tried it. \n<a href=\"https://www.kaggle.com/c/google-quest-challenge/discussion/129840#742757\">https://www.kaggle.com/c/google-quest-challenge/discussion/129840#742757</a></p>\n\n<p>In my experience tracking in a spreadsheet requires a lot of discipline and lacks flexibility, hence I was wondering about more automated tools. But maybe free text like this best fits the needs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898011,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "06/23/2020 08:22:39",
          "content": "<p>I'm using a spreadsheet either on Excel, I'm trying a note-taking app called Notion. That works quite well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899373,
          "author_name": "sebastienm",
          "author_url": "",
          "post_date": "06/24/2020 07:39:42",
          "content": "<p>Noted thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897763,
      "author_name": "ibtesama",
      "author_url": "",
      "post_date": "06/23/2020 04:42:48",
      "content": "<p>Congratulations <a href=\"/rftexas\">@rftexas</a> ! Learned a lot from you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898126,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "06/23/2020 10:23:37",
      "content": "<p>Very nice advise :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 898314,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "06/23/2020 12:48:11",
          "content": "<p>Much appreciated! Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898808,
      "author_name": "nandhuelan",
      "author_url": "",
      "post_date": "06/23/2020 19:06:02",
      "content": "<p>Need of the hour!! <a href=\"/rftexas\">@rftexas</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899823,
      "author_name": "mahdhiashraf",
      "author_url": "",
      "post_date": "06/24/2020 13:16:18",
      "content": "<p>congrats <a href=\"/rftexas\">@rftexas</a>  ^^</p>",
      "votes": null,
      "replies": [
        {
          "id": 915335,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "07/04/2020 16:45:39",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 915186,
      "author_name": "bdubreu",
      "author_url": "",
      "post_date": "07/04/2020 14:26:44",
      "content": "<p>Hi <a href=\"/rftexas\">@rftexas</a>, could you elaborate a bit a why and how using different seeds makes for a reliable CV strategy ? thanks a lot for the writeup !</p>",
      "votes": null,
      "replies": [
        {
          "id": 915340,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "07/04/2020 16:50:07",
          "content": "<p>Sure! One way to make your CV strategy more reliable is to use the same number of samples, say 100k samples divided into 5 folds. You can evaluate the performance using k-fold. But you can also use the same 100k samples using another seed to create your folds. Do that 3 or 4 times and you'll have validated on 400k samples. </p>\n\n<p>Of course, I anticipate your remark saying that the k-fold scores you'll get using different seeds will be correlated since those are essentially the same samples fed to the model even though they are fed in a different order. But it has been empirically shown that this technique works quite well. I've tried it and you can come up with a CV correlated to LB. <a href=\"/cdeotte\">@cdeotte</a> wrote on it on a previous post, I recommend you checking this post as it was very instructive.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 915935,
          "author_name": "bdubreu",
          "author_url": "",
          "post_date": "07/05/2020 08:11:47",
          "content": "<p>Thanks for taking the time to answer. I never considered it that way ! I'll try to find his post... is it in this competition discussion as well ? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "897531": "Since this competition is now finished, I'd like to first thank all my great teammates @ibtesama @shahules and my boy @tanulsingh077 . It's been great working with you, and I wish you all the best for your future competitions.\n\nSecond, congratulations to everybody for the efforts you have put in this competition for the last 3 months and congratulations to all the medal winners.\n\nAs there are already many great write-ups about winning solutions for this competition, I wanted to share with you four aspects that I find critical to succeed in a Kaggle competition. **They have made the difference for me between a competition with and without medal**. This is my way to tackle a Kaggle problem. **It is obviously not perfect and I'm open to criticism as long as it is constructive.**\n\nUnlike other recommendations, those ones are not **technical but rather organizational.** I've applied this framework with my teammates for this competition and the Tweet sentiment competition where I got two silver medals back to back.\n\nWhen tackling a Kaggle competition, you should always have your **own framework**. This is mine, which I have built throughout those competitions. There's no doubt it will be different in 6 months, I encourage you to continuously improve it. \n\nFirst, read an EDA or do your own analysis. After doing your analysis of the data, framing the problem, trying to think of a similar competition you did, throw your ideas on a note. Don't forget to have a look at other winning solutions, to try the same ideas as them. At this stage, you should have an unsorted list of ideas to try: preprocessing, model architecture, blending ideas...\n\nSecond, now you have a list of ideas to try, you need to sort them out. My main criterion is the computational time. **You want to iterate as fast as you can to quickly have a fully-running pipeline**. That's why I would recommend starting with the ideas that don't add too much computational time to your pipeline. Hence, you should start first with preprocessing methods (cleaning for text...), sampling techniques (random weighted sampler...) and then data augmentation techniques. You can then add exotic model architectures (from the simplest to the more complex). The key message here is to iterate quickly, try with a simple architecture at first. If using Transformers, use BERT-base or even DistilBERT. For a computer vision problem, use ResNet34 or EfficientNetB0. As you can see in the picture below, I have begun with a simpler model bert-multilingual-uncased before choosing a more complex one like XLM-Roberta. It has saved me time, especially like me when you have no other GPU/TPU resources than the ones offered by Kaggle (thanks to them!).\n\nThird, and **it is probably the most important idea,** **report everything on a table sheet**. You need to track everything from your experiments. Obviously make sure to seed everything as you want your results to be comparable. This is (a part of) the table I used for Jigsaw:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1302247%2F6ecd271b6272b3785d3e5825dad23230%2FScreen_Shot_2020-06-23_at_12.02.50_AM.png?generation=1592871307292789&amp;alt=media)\n\nFourth, before starting any modeling, make sure to build a solid validation strategy. @cdeotte  and @cpmpml have given some tips on the discussion thread of Tweet sentiment extraction. You must make sure to have reliable CV strategy, otherwise at best the final stages of the competition will be a tense moment for you as you fear shake-up and at worst you are out of the medal zone after having worked multiple weeks on a problem. Furthermore, in SIIM-ISIC melanoma competition, I just started computing my CV statistics: mean and standard deviation. One technique is to add as many folds as you need to reduce the standard deviation of your CV scores above some threshold. Another one pointed out by @cdeotte is to use different seeds. That way you'll have a reliable CV strategy that you can trust when you need to choose between submissions.\n\nI hope this quick message helped you. I think it is critical to have a clear-defined structure to follow and is an invaluable perk to have as a data scientist / ML engineer in your daily job. \n\nAll the best,\n\nPierre-Antoine",
    "897604": "Congratulations Pierre",
    "897610": "Any preferred tool for tracking experiments? There is mention of a tool called mag there but I have not tried it. \nhttps://www.kaggle.com/c/google-quest-challenge/discussion/129840#742757\n\nIn my experience tracking in a spreadsheet requires a lot of discipline and lacks flexibility, hence I was wondering about more automated tools. But maybe free text like this best fits the needs.",
    "897763": "Congratulations @rftexas ! Learned a lot from you!",
    "898011": "I'm using a spreadsheet either on Excel, I'm trying a note-taking app called Notion. That works quite well.",
    "898126": "Very nice advise :)",
    "898314": "Much appreciated! Thanks!",
    "898808": "Need of the hour!! @rftexas",
    "899373": "Noted thanks!",
    "899823": "congrats @rftexas  ^^",
    "915186": "Hi @rftexas, could you elaborate a bit a why and how using different seeds makes for a reliable CV strategy ? thanks a lot for the writeup !",
    "915335": "Thanks!",
    "915340": "Sure! One way to make your CV strategy more reliable is to use the same number of samples, say 100k samples divided into 5 folds. You can evaluate the performance using k-fold. But you can also use the same 100k samples using another seed to create your folds. Do that 3 or 4 times and you'll have validated on 400k samples. \n\nOf course, I anticipate your remark saying that the k-fold scores you'll get using different seeds will be correlated since those are essentially the same samples fed to the model even though they are fed in a different order. But it has been empirically shown that this technique works quite well. I've tried it and you can come up with a CV correlated to LB. @cdeotte wrote on it on a previous post, I recommend you checking this post as it was very instructive.",
    "915935": "Thanks for taking the time to answer. I never considered it that way ! I'll try to find his post... is it in this competition discussion as well ?",
    "988903": "rftexas  Really great advice  . Can you suggest some \" When tackling a Kaggle competition, you should always have your own framework \" . Would you elaborate a bit on this  . \nIf possible Could you share your pipeline for this competition ?"
  },
  "source": "meta"
}