{
  "id": 140299,
  "title": "Mistakes",
  "url": "/competitions/deepfake-detection-challenge/discussion/140299",
  "author_name": "",
  "post_date": "2020-04-01T08:23:51.029825800Z",
  "votes": 51,
  "comment_count": 25,
  "views": 0,
  "content": "<p>It is always good and inspiring to read success stories.</p>\n\n<p>But I also value failure stories so I would like to share one:)</p>\n\n<p>Well, this is not exactly a failure(actually I am kind of happy with my current score which is a lot better than I expected at the beginning as a DL noob) story but a list of mistakes I did as a noob(skip to the \"mistakes\" part if you are not interested in the \"story\" part).</p>\n\n<p>Last August I decided to stop being a full time employee and take a funemployement break for 6 months. Then I quit my job at November.</p>\n\n<p>I wanted to spend time with my family, learn new things and do whatever I like for 6 months:)</p>\n\n<p>Machine learning, specifically deep learning was one of the things I wanted/tried to learn in the past but never had enough time to invest in. I built a few basic classic models in the past but it was basically just following some other people' footsteps without really knowing what i am doing.</p>\n\n<p>So during my break I wanted to revisit it and at least learn the basics.. the fastai videos by Jeremy Howard really motivated me as a starter tutorial.\nBut I have never been good at learning from books/tutorials/videos only. I always needed a target and some kind of deadline to motivate me. Luckily I have seen this competition and thought it was a good opportunity for my learning process.</p>\n\n<p>Before this tournament; I did not like/know python much(i still don't know it well but at least now I like it:)); I haven't used pytorch/tensorflor/keras etc before.</p>\n\n<p><strong>So please read the following \"mistakes\"/\"learnings\" knowing that they are not coming from an expert:)</strong></p>\n\n<p>Anyways as of today; when I look through past few months; I see clear mistakes I did as a noob; and I want to share them in this post.</p>\n\n<h3>Mistakes</h3>\n\n<p>My learning below can still be noob conclusions but it is what I think as of today. So if you have comments/advice on my takes; you are very welcome.</p>\n\n<p>tl;dr everything that makes you waste time is a mistake because this loss time is so valuable for making more experiments. And doing more experiments is everything(or 95% of everything).</p>\n\n<h3>1- Choosing the wrong environment/hardware.</h3>\n\n<p>The first environment I have chosen for this tournament was Google colab + google drive. And the google drive part was a huge mistake. I subscribed to some terabyte plan and I thought it would be ok to use google drive because you can access google drive data from colab and what could go wrong? \nWell; never thought I would hit some kind of daily download size limit(I did not even know it existed) on google drive because whenever you mount google drive from colab; and start using data it is counted as a \"download\" although both sides are Google. \nOnce I started to hit the limits; I needed to redo things from scratch on AWS(thanks to the free credits) \nAlso I realized that the io performance in this competition is as important as the GPU performance so I switched to my local machine which has worse GPU than AWS instance but have 4TB NVMe SSD which is a lot faster than AWS disks. And my training times became X3 faster than AWS. I should have foreseen it earlier.</p>\n\n<h3>2- Spending time on choosing which framework too much.</h3>\n\n<p>After spending some significant time on understanding some basics of fastai/keras/tensorflow/pytorch; i have picked pytorch and progressed with it. Now when I look back; I think it was not necessary to spend that much time on framework selection; because now I think what matters most was understanding the problem and the data. </p>\n\n<h3>3- Not keeping a good scientific journal.</h3>\n\n<p>a month ago I realized; the biggest mistake I have done was not keeping a good journal. I took notes, saved notebooks etc but this was not enough. Because after experimenting many many things; you lose track of \"what was working/what was not working/what is next to try\" and sometimes you repeat yourself. So it is really important to keep a detailed journal/logs; it can be an excel table, a list of notebooks whatever; but it should be kept in a structured and very well organized way I think. Later I started to save all parameters and preliminary steps(like everything from how I prepared the datasets to results and whatever there is) together with the model file and also started to keep a table. After I start doing this; it was a lot easier to see the picture.</p>\n\n<h3>4- Not spending more time on data separation.</h3>\n\n<p>I knew it is really important to keep your validation and test sets isolated from the training set but  (I don't know if it is related to this competition) making sure that you are really doing this was really challenging. And if I was redoing this competition I would have spent more time at the beginning to understand the data better(how many different types of fakizations there, similar actors, which folders are close to the other folders in terms of similarity etc) and build the necessary tools to make sure they are isolated in a better way. Towards the end of the competition I built some utility functions to check this and it really helped but i should have spent more time to have more robust utilities to make sure of it.</p>\n\n<h3>5- Waiting.</h3>\n\n<p>This was also one of my biggest mistakes. Fortunately; I fixed it early. At the beginning; several times I believed I found a good methodology; but I tried it on a dataset which is bigger than needed(not the full dataset but still bigger than needed). So I needed to wait to get the results of the experiments. After some time I realized this is a clear mistake. And whenever I found a method to build a dataset I also built a \"small\" and \"medium\" versions of it which contains enough characteristics but give the results earlier. Now I believe if you are waiting too much; you are making a mistake. Finally I had a pipeline with 3 stages; first I try the method in the smallest test(10 minutes max) then medium set(2 hours max) then full set. So I catch mistakes/problems earlier.</p>\n\n<p>And finally; thanks for the competition it was a huge fun and a lot of learning!</p>",
  "messages": [
    {
      "id": "793721",
      "postDate": "04/01/2020 08:23:51",
      "content": "<p>It is always good and inspiring to read success stories.</p>\n\n<p>But I also value failure stories so I would like to share one:)</p>\n\n<p>Well, this is not exactly a failure(actually I am kind of happy with my current score which is a lot better than I expected at the beginning as a DL noob) story but a list of mistakes I did as a noob(skip to the \"mistakes\" part if you are not interested in the \"story\" part).</p>\n\n<p>Last August I decided to stop being a full time employee and take a funemployement break for 6 months. Then I quit my job at November.</p>\n\n<p>I wanted to spend time with my family, learn new things and do whatever I like for 6 months:)</p>\n\n<p>Machine learning, specifically deep learning was one of the things I wanted/tried to learn in the past but never had enough time to invest in. I built a few basic classic models in the past but it was basically just following some other people' footsteps without really knowing what i am doing.</p>\n\n<p>So during my break I wanted to revisit it and at least learn the basics.. the fastai videos by Jeremy Howard really motivated me as a starter tutorial.\nBut I have never been good at learning from books/tutorials/videos only. I always needed a target and some kind of deadline to motivate me. Luckily I have seen this competition and thought it was a good opportunity for my learning process.</p>\n\n<p>Before this tournament; I did not like/know python much(i still don't know it well but at least now I like it:)); I haven't used pytorch/tensorflor/keras etc before.</p>\n\n<p><strong>So please read the following \"mistakes\"/\"learnings\" knowing that they are not coming from an expert:)</strong></p>\n\n<p>Anyways as of today; when I look through past few months; I see clear mistakes I did as a noob; and I want to share them in this post.</p>\n\n<h3>Mistakes</h3>\n\n<p>My learning below can still be noob conclusions but it is what I think as of today. So if you have comments/advice on my takes; you are very welcome.</p>\n\n<p>tl;dr everything that makes you waste time is a mistake because this loss time is so valuable for making more experiments. And doing more experiments is everything(or 95% of everything).</p>\n\n<h3>1- Choosing the wrong environment/hardware.</h3>\n\n<p>The first environment I have chosen for this tournament was Google colab + google drive. And the google drive part was a huge mistake. I subscribed to some terabyte plan and I thought it would be ok to use google drive because you can access google drive data from colab and what could go wrong? \nWell; never thought I would hit some kind of daily download size limit(I did not even know it existed) on google drive because whenever you mount google drive from colab; and start using data it is counted as a \"download\" although both sides are Google. \nOnce I started to hit the limits; I needed to redo things from scratch on AWS(thanks to the free credits) \nAlso I realized that the io performance in this competition is as important as the GPU performance so I switched to my local machine which has worse GPU than AWS instance but have 4TB NVMe SSD which is a lot faster than AWS disks. And my training times became X3 faster than AWS. I should have foreseen it earlier.</p>\n\n<h3>2- Spending time on choosing which framework too much.</h3>\n\n<p>After spending some significant time on understanding some basics of fastai/keras/tensorflow/pytorch; i have picked pytorch and progressed with it. Now when I look back; I think it was not necessary to spend that much time on framework selection; because now I think what matters most was understanding the problem and the data. </p>\n\n<h3>3- Not keeping a good scientific journal.</h3>\n\n<p>a month ago I realized; the biggest mistake I have done was not keeping a good journal. I took notes, saved notebooks etc but this was not enough. Because after experimenting many many things; you lose track of \"what was working/what was not working/what is next to try\" and sometimes you repeat yourself. So it is really important to keep a detailed journal/logs; it can be an excel table, a list of notebooks whatever; but it should be kept in a structured and very well organized way I think. Later I started to save all parameters and preliminary steps(like everything from how I prepared the datasets to results and whatever there is) together with the model file and also started to keep a table. After I start doing this; it was a lot easier to see the picture.</p>\n\n<h3>4- Not spending more time on data separation.</h3>\n\n<p>I knew it is really important to keep your validation and test sets isolated from the training set but  (I don't know if it is related to this competition) making sure that you are really doing this was really challenging. And if I was redoing this competition I would have spent more time at the beginning to understand the data better(how many different types of fakizations there, similar actors, which folders are close to the other folders in terms of similarity etc) and build the necessary tools to make sure they are isolated in a better way. Towards the end of the competition I built some utility functions to check this and it really helped but i should have spent more time to have more robust utilities to make sure of it.</p>\n\n<h3>5- Waiting.</h3>\n\n<p>This was also one of my biggest mistakes. Fortunately; I fixed it early. At the beginning; several times I believed I found a good methodology; but I tried it on a dataset which is bigger than needed(not the full dataset but still bigger than needed). So I needed to wait to get the results of the experiments. After some time I realized this is a clear mistake. And whenever I found a method to build a dataset I also built a \"small\" and \"medium\" versions of it which contains enough characteristics but give the results earlier. Now I believe if you are waiting too much; you are making a mistake. Finally I had a pipeline with 3 stages; first I try the method in the smallest test(10 minutes max) then medium set(2 hours max) then full set. So I catch mistakes/problems earlier.</p>\n\n<p>And finally; thanks for the competition it was a huge fun and a lot of learning!</p>",
      "rawMarkdown": "It is always good and inspiring to read success stories.\n\nBut I also value failure stories so I would like to share one:)\n\nWell, this is not exactly a failure(actually I am kind of happy with my current score which is a lot better than I expected at the beginning as a DL noob) story but a list of mistakes I did as a noob(skip to the \"mistakes\" part if you are not interested in the \"story\" part).\n\nLast August I decided to stop being a full time employee and take a funemployement break for 6 months. Then I quit my job at November.\n\nI wanted to spend time with my family, learn new things and do whatever I like for 6 months:)\n\nMachine learning, specifically deep learning was one of the things I wanted/tried to learn in the past but never had enough time to invest in. I built a few basic classic models in the past but it was basically just following some other people' footsteps without really knowing what i am doing.\n\nSo during my break I wanted to revisit it and at least learn the basics.. the fastai videos by Jeremy Howard really motivated me as a starter tutorial.\nBut I have never been good at learning from books/tutorials/videos only. I always needed a target and some kind of deadline to motivate me. Luckily I have seen this competition and thought it was a good opportunity for my learning process.\n\nBefore this tournament; I did not like/know python much(i still don't know it well but at least now I like it:)); I haven't used pytorch/tensorflor/keras etc before.\n\n\n**So please read the following \"mistakes\"/\"learnings\" knowing that they are not coming from an expert:)**\n\n\nAnyways as of today; when I look through past few months; I see clear mistakes I did as a noob; and I want to share them in this post.\n\n\n\n### Mistakes\n\nMy learning below can still be noob conclusions but it is what I think as of today. So if you have comments/advice on my takes; you are very welcome.\n\ntl;dr everything that makes you waste time is a mistake because this loss time is so valuable for making more experiments. And doing more experiments is everything(or 95% of everything).\n\n\n### 1- Choosing the wrong environment/hardware.\n\nThe first environment I have chosen for this tournament was Google colab + google drive. And the google drive part was a huge mistake. I subscribed to some terabyte plan and I thought it would be ok to use google drive because you can access google drive data from colab and what could go wrong? \nWell; never thought I would hit some kind of daily download size limit(I did not even know it existed) on google drive because whenever you mount google drive from colab; and start using data it is counted as a \"download\" although both sides are Google. \nOnce I started to hit the limits; I needed to redo things from scratch on AWS(thanks to the free credits) \nAlso I realized that the io performance in this competition is as important as the GPU performance so I switched to my local machine which has worse GPU than AWS instance but have 4TB NVMe SSD which is a lot faster than AWS disks. And my training times became X3 faster than AWS. I should have foreseen it earlier.\n\n\n### 2- Spending time on choosing which framework too much.\n\nAfter spending some significant time on understanding some basics of fastai/keras/tensorflow/pytorch; i have picked pytorch and progressed with it. Now when I look back; I think it was not necessary to spend that much time on framework selection; because now I think what matters most was understanding the problem and the data. \n\n\n### 3- Not keeping a good scientific journal.\n\na month ago I realized; the biggest mistake I have done was not keeping a good journal. I took notes, saved notebooks etc but this was not enough. Because after experimenting many many things; you lose track of \"what was working/what was not working/what is next to try\" and sometimes you repeat yourself. So it is really important to keep a detailed journal/logs; it can be an excel table, a list of notebooks whatever; but it should be kept in a structured and very well organized way I think. Later I started to save all parameters and preliminary steps(like everything from how I prepared the datasets to results and whatever there is) together with the model file and also started to keep a table. After I start doing this; it was a lot easier to see the picture.\n\n\n### 4- Not spending more time on data separation.\n\nI knew it is really important to keep your validation and test sets isolated from the training set but  (I don't know if it is related to this competition) making sure that you are really doing this was really challenging. And if I was redoing this competition I would have spent more time at the beginning to understand the data better(how many different types of fakizations there, similar actors, which folders are close to the other folders in terms of similarity etc) and build the necessary tools to make sure they are isolated in a better way. Towards the end of the competition I built some utility functions to check this and it really helped but i should have spent more time to have more robust utilities to make sure of it.\n\n\n### 5- Waiting.\n\nThis was also one of my biggest mistakes. Fortunately; I fixed it early. At the beginning; several times I believed I found a good methodology; but I tried it on a dataset which is bigger than needed(not the full dataset but still bigger than needed). So I needed to wait to get the results of the experiments. After some time I realized this is a clear mistake. And whenever I found a method to build a dataset I also built a \"small\" and \"medium\" versions of it which contains enough characteristics but give the results earlier. Now I believe if you are waiting too much; you are making a mistake. Finally I had a pipeline with 3 stages; first I try the method in the smallest test(10 minutes max) then medium set(2 hours max) then full set. So I catch mistakes/problems earlier.\n\nAnd finally; thanks for the competition it was a huge fun and a lot of learning!",
      "votes": null
    },
    {
      "id": "793730",
      "postDate": "04/01/2020 08:29:34",
      "content": "<p>Very well explained. Learning from mistakes is very important. \nHope you learned a lot from this competition</p>",
      "rawMarkdown": "Very well explained. Learning from mistakes is very important. \nHope you learned a lot from this competition",
      "votes": null
    },
    {
      "id": "793741",
      "postDate": "04/01/2020 08:37:12",
      "content": "<p>Thanks! I definitely learned a lot.</p>",
      "rawMarkdown": "Thanks! I definitely learned a lot.",
      "votes": null
    },
    {
      "id": "793755",
      "postDate": "04/01/2020 08:47:33",
      "content": "<p>Your story is as good as success story. And also you made decent success in this competition.\nThank you.</p>",
      "rawMarkdown": "Your story is as good as success story. And also you made decent success in this competition.\nThank you.",
      "votes": null
    },
    {
      "id": "793789",
      "postDate": "04/01/2020 09:25:59",
      "content": "<p>Very helpful. Honest person. Thx.</p>",
      "rawMarkdown": "Very helpful. Honest person. Thx.",
      "votes": null
    },
    {
      "id": "793829",
      "postDate": "04/01/2020 10:19:05",
      "content": "<p>Thank you for sharing this great story. I can't believe how similar it is to mine. I also quit my job to focus on ML/DL.\nThere are lots of common mistakes: \n1) I also started with colab and I faced issues with its timeout and storage limitations. I haven't downloaded the whole dataset until last week where I started using AWS but without free credits :(!\n2) The mistakes that I didn't fix were the data separation and waiting for 3 hours for a model to finish.\nThere is one more mistake that I was working all alone, I think it would have been better to join a team of 1 or 2 people. This would have helped me avoid some mistakes or at least fix them earlier.</p>\n\n<p>In my case, I consider this was a lack of experience for me which I am sure I gained a lot of it now.</p>\n\n<p>On a side note, there is a \"mistake\" in point 4 because it has text from part 3 (So it is really important to keep a detailed journal/logs...)</p>\n\n<p>Thank you again and good luck with current and future competitions</p>",
      "rawMarkdown": "Thank you for sharing this great story. I can't believe how similar it is to mine. I also quit my job to focus on ML/DL.\nThere are lots of common mistakes: \n1) I also started with colab and I faced issues with its timeout and storage limitations. I haven't downloaded the whole dataset until last week where I started using AWS but without free credits :(!\n2) The mistakes that I didn't fix were the data separation and waiting for 3 hours for a model to finish.\nThere is one more mistake that I was working all alone, I think it would have been better to join a team of 1 or 2 people. This would have helped me avoid some mistakes or at least fix them earlier.\n\nIn my case, I consider this was a lack of experience for me which I am sure I gained a lot of it now.\n\nOn a side note, there is a \"mistake\" in point 4 because it has text from part 3 (So it is really important to keep a detailed journal/logs...)\n\nThank you again and good luck with current and future competitions",
      "votes": null
    },
    {
      "id": "793841",
      "postDate": "04/01/2020 10:35:21",
      "content": "<p>I wish I have sufficient conditions of all aspects to have a 6-month break like you :(. Anyway, thanks for sharing your great experience. It's always beneficial to review the goods and bads of one's self after the competition.</p>",
      "rawMarkdown": "I wish I have sufficient conditions of all aspects to have a 6-month break like you :(. Anyway, thanks for sharing your great experience. It's always beneficial to review the goods and bads of one's self after the competition.",
      "votes": null
    },
    {
      "id": "793859",
      "postDate": "04/01/2020 11:06:54",
      "content": "<p>yes I have fixed it thanks!</p>",
      "rawMarkdown": "yes I have fixed it thanks!",
      "votes": null
    },
    {
      "id": "794107",
      "postDate": "04/01/2020 15:08:02",
      "content": "<p>Welcome and I think this is a great success story you should be proud of :) In terms of data, I couldn't agree more I spent 3 weeks on data and 1 week on modeling. In my case, I simply used OneNote to keep all my ideas and conceptualize them with pros/cons before even starting coding anything, of course discussions and kernels helped a lot to get this going. I also recommend <a href=\"https://nbdev.fast.ai/\">nbdev</a> if you like notebooks. I used it for a Kaggle competition for the first time but it improved my productivity x2-x3. It has the power of iterating fast with notebooks, keeping scientific journals in the same repo and also modularizing your code by converting notebooks into py files similar to a regular python library. Good luck on your journey!</p>",
      "rawMarkdown": "Welcome and I think this is a great success story you should be proud of :) In terms of data, I couldn't agree more I spent 3 weeks on data and 1 week on modeling. In my case, I simply used OneNote to keep all my ideas and conceptualize them with pros/cons before even starting coding anything, of course discussions and kernels helped a lot to get this going. I also recommend [nbdev](https://nbdev.fast.ai/) if you like notebooks. I used it for a Kaggle competition for the first time but it improved my productivity x2-x3. It has the power of iterating fast with notebooks, keeping scientific journals in the same repo and also modularizing your code by converting notebooks into py files similar to a regular python library. Good luck on your journey!",
      "votes": null
    },
    {
      "id": "794123",
      "postDate": "04/01/2020 15:26:41",
      "content": "<p>Very inspiring! All the best.</p>",
      "rawMarkdown": "Very inspiring! All the best.",
      "votes": null
    },
    {
      "id": "794160",
      "postDate": "04/01/2020 16:05:07",
      "content": "<p>Excellent Job!  When I work on my car and need to watch a video, I watch a professionally edited video from a pro so I make sure I know all the details of what parts and tolls I need and the procedure for the fix.  Then I watch an average guy like myself so it, so that I can see some reality.  Oh, it is going to actually take this long!  Watch out for this!  Ooops don't do that!</p>\n\n<p>Your post was excellent.</p>\n\n<p>So some questions, now.  Under \"Choosing the wrong environment/hardware\" you said you switch to your local machine.  You mean like your personal laptop/desktop?</p>\n\n<p>Under \"Not keeping a good scientific journal\".  Do you have a suggested format or a link to a journal you like?</p>\n\n<p>Keep-up the good work.</p>",
      "rawMarkdown": "Excellent Job!  When I work on my car and need to watch a video, I watch a professionally edited video from a pro so I make sure I know all the details of what parts and tolls I need and the procedure for the fix.  Then I watch an average guy like myself so it, so that I can see some reality.  Oh, it is going to actually take this long!  Watch out for this!  Ooops don't do that!\n\nYour post was excellent.\n\nSo some questions, now.  Under \"Choosing the wrong environment/hardware\" you said you switch to your local machine.  You mean like your personal laptop/desktop?\n\nUnder \"Not keeping a good scientific journal\".  Do you have a suggested format or a link to a journal you like?\n\nKeep-up the good work.",
      "votes": null
    },
    {
      "id": "794227",
      "postDate": "04/01/2020 17:10:22",
      "content": "<p>Wow nbdev looks/sounds great I will definitely give it a try!</p>",
      "rawMarkdown": "Wow nbdev looks/sounds great I will definitely give it a try!",
      "votes": null
    },
    {
      "id": "794274",
      "postDate": "04/01/2020 17:45:18",
      "content": "<p>Thanks a lot!\nYes I used my personal desktop(GPU: gtx 1080 ti, SSD: 4TB NVMe, OS: ubuntu). \nMy observation is; even the cheapest AWS EC2 GPU instances(like g3s) have better GPUs(especially in terms of memory) than my desktop but their EBS disks are utilized for many small IO requests(very low latency) but not as good as NVMe SSDs with m.2 interface in terms of continuous read/write speed. \nTo have an EC2 instance with NVMe SSD disks with enough storage size you need to pay around $5-$20 / hour and it would be impossible for me to fit into free credit budget if I used that. (you can see the EC2 instance lists here: <a href=\"https://aws.amazon.com/ec2/pricing/on-demand/\">https://aws.amazon.com/ec2/pricing/on-demand/</a> , I realized using an instance with NVMe SSD vs EBS storage has great performance difference if you are dealing with a lot of medium sized files. and instances with NVMe SSDs with enough size are expensive.\nBut using AWS was not a mistake(I just wanted to point the importance of IO speed in this competition); it was fast enough. \nTrying to handle this competition with colab+google drive on the other hand; was a clear mistake for me. </p>\n\n<p>About the scientific journal; I don't know a good tool. I used a very basic structure; \n- I created a dictionary with all the parameters(lr,wd,freeze_layer,batch_size, augmentation probabilities&amp;parameters, pretrained model pointer, train_data_path, val_data_path, train part nos, val_part_nos, train loss, val loss, epoch no,  etc etc.) and whenever I save the model I also saved that dictionary together with it. \n- I also created a spreadsheet on excel to record all these with my comments.</p>\n\n<p>But <a href=\"/keremt\">@keremt</a> shared in his comment below a link to a library(nbdev) I will give it a try.</p>",
      "rawMarkdown": "Thanks a lot!\nYes I used my personal desktop(GPU: gtx 1080 ti, SSD: 4TB NVMe, OS: ubuntu). \nMy observation is; even the cheapest AWS EC2 GPU instances(like g3s) have better GPUs(especially in terms of memory) than my desktop but their EBS disks are utilized for many small IO requests(very low latency) but not as good as NVMe SSDs with m.2 interface in terms of continuous read/write speed. \nTo have an EC2 instance with NVMe SSD disks with enough storage size you need to pay around $5-$20 / hour and it would be impossible for me to fit into free credit budget if I used that. (you can see the EC2 instance lists here: https://aws.amazon.com/ec2/pricing/on-demand/ , I realized using an instance with NVMe SSD vs EBS storage has great performance difference if you are dealing with a lot of medium sized files. and instances with NVMe SSDs with enough size are expensive.\nBut using AWS was not a mistake(I just wanted to point the importance of IO speed in this competition); it was fast enough. \nTrying to handle this competition with colab+google drive on the other hand; was a clear mistake for me. \n\nAbout the scientific journal; I don't know a good tool. I used a very basic structure; \n- I created a dictionary with all the parameters(lr,wd,freeze_layer,batch_size, augmentation probabilities¶meters, pretrained model pointer, train_data_path, val_data_path, train part nos, val_part_nos, train loss, val loss, epoch no,  etc etc.) and whenever I save the model I also saved that dictionary together with it. \n- I also created a spreadsheet on excel to record all these with my comments.\n\nBut @keremt shared in his comment below a link to a library(nbdev) I will give it a try.",
      "votes": null
    },
    {
      "id": "794404",
      "postDate": "04/01/2020 19:36:58",
      "content": "<p>Thanks for sharing your journey. As for the \"journaling\" part, I have used <a href=\"https://mlflow.org/\"><strong>mlflow</strong></a> (mostly at work) and it is very useful to record metadata about your experiments. </p>",
      "rawMarkdown": "Thanks for sharing your journey. As for the \"journaling\" part, I have used [**mlflow**](https://mlflow.org/) (mostly at work) and it is very useful to record metadata about your experiments.",
      "votes": null
    },
    {
      "id": "794652",
      "postDate": "04/02/2020 01:35:02",
      "content": "<p>Thank you for sharing this. I can relate with you when it comes to #3 and #5....especially #3 (I'm a disaster when it comes to organization). About framework, I feel that keras is the easiest to get into compared to plain tensorflow or pytorch (i'm also a newbie in this field trying to learn. I code in C and assembly for a living)</p>",
      "rawMarkdown": "Thank you for sharing this. I can relate with you when it comes to #3 and #5....especially #3 (I'm a disaster when it comes to organization). About framework, I feel that keras is the easiest to get into compared to plain tensorflow or pytorch (i'm also a newbie in this field trying to learn. I code in C and assembly for a living)",
      "votes": null
    },
    {
      "id": "796486",
      "postDate": "04/03/2020 16:13:44",
      "content": "<p>Thanks <a href=\"/emrebayram\">@emrebayram</a> for sharing this inspiring journal of yours. I fully resonate with \"..doing more experiments is everything...\". Outcomes of the experiment gives so many ideas.</p>\n\n<p>I have a much broader question regarding keeping a journal (sorry to pull on this).  Keeping journal, though sounds simple, is not simple at all. For example, storing one line per model in excel seems like a nice idea, but when I try something, I make so many small-small changes and also tinker with architecture etc. At this point, excel is no longer a good option (Still I believe the best. I specced a whole application to do this….someday I'll build it as well.). </p>\n\n<p>A month back, I created some variance plots on fake and real data. Now I can’t even find where it is stored!! (Feeling frustrated, I came back and thought of asking... whats the harm anyway)</p>\n\n<p>So could you elaborate (more) on how you journal stuff? I believe you also keep track of your ideas somewhere. Do you have any sorting mechanism etc as well? </p>",
      "rawMarkdown": "Thanks @emrebayram for sharing this inspiring journal of yours. I fully resonate with \"..doing more experiments is everything...\". Outcomes of the experiment gives so many ideas.\n\nI have a much broader question regarding keeping a journal (sorry to pull on this).  Keeping journal, though sounds simple, is not simple at all. For example, storing one line per model in excel seems like a nice idea, but when I try something, I make so many small-small changes and also tinker with architecture etc. At this point, excel is no longer a good option (Still I believe the best. I specced a whole application to do this….someday I'll build it as well.). \n\nA month back, I created some variance plots on fake and real data. Now I can’t even find where it is stored!! (Feeling frustrated, I came back and thought of asking... whats the harm anyway)\n\nSo could you elaborate (more) on how you journal stuff? I believe you also keep track of your ideas somewhere. Do you have any sorting mechanism etc as well?",
      "votes": null
    },
    {
      "id": "796563",
      "postDate": "04/03/2020 17:32:32",
      "content": "<p>Thanks <a href=\"/aknirala\">@aknirala</a> , I don't have a perfect solution for keeping a journal. \" I make so many small-small changes and also tinker with architecture etc. \" was definitely a big problem for me too. But what worked for me best is; I created a global dictionary(something called global_params); and whenever i am playing with something (training files, augmentation, architecture, epoch numbers, random probabilities, pre-processing variations, pretrain file, frozen layers everything); i carried it to that dictionary and modified there and read from there. For example if i realized i am playing with some code part; i made it configurable and carried these configuration params to the global params dictionary. \nThis might not be a good practice in terms of software engineering but here worked for me. And I had a simple model save method and whenever I am saving the model state dict and optimizer state dict; I saved this global params with it together with my loss value. so whatever i played is stored together with the model. So for every point; i saved 2 files like this:\n 92258335 Mar 31 21:03 0.32895452830785804_model.pth\n183316294 Mar 31 21:05 0.32895452830785804_rsm_pt.pth\nfile names contains the loss value; it was easier that way to see for me. model.pth is the model state dict. rsm_pt stands for resume point and contains:\noptimizer state dict, all params; basically everything.\nAnd I pushed these files to git together with the current code. \nThat way; if i wanted to see what i have tried I loaded and printed the global params variable and observed it(it is basically a summary of the current code/all params/methodologies). and if i wanted to resume to a point to move from there; i pulled that version of code together with model file and global params.\nThis is not a perfect solution; i am evaluating existing tools/methods for doing this in a better way. but this was my quick solution while doing this competition.</p>\n\n<p>I used excel to store high level ideas/comments/params for significant milestones.</p>",
      "rawMarkdown": "Thanks @aknirala , I don't have a perfect solution for keeping a journal. \" I make so many small-small changes and also tinker with architecture etc. \" was definitely a big problem for me too. But what worked for me best is; I created a global dictionary(something called global_params); and whenever i am playing with something (training files, augmentation, architecture, epoch numbers, random probabilities, pre-processing variations, pretrain file, frozen layers everything); i carried it to that dictionary and modified there and read from there. For example if i realized i am playing with some code part; i made it configurable and carried these configuration params to the global params dictionary. \nThis might not be a good practice in terms of software engineering but here worked for me. And I had a simple model save method and whenever I am saving the model state dict and optimizer state dict; I saved this global params with it together with my loss value. so whatever i played is stored together with the model. So for every point; i saved 2 files like this:\n 92258335 Mar 31 21:03 0.32895452830785804_model.pth\n183316294 Mar 31 21:05 0.32895452830785804_rsm_pt.pth\nfile names contains the loss value; it was easier that way to see for me. model.pth is the model state dict. rsm_pt stands for resume point and contains:\noptimizer state dict, all params; basically everything.\nAnd I pushed these files to git together with the current code. \nThat way; if i wanted to see what i have tried I loaded and printed the global params variable and observed it(it is basically a summary of the current code/all params/methodologies). and if i wanted to resume to a point to move from there; i pulled that version of code together with model file and global params.\nThis is not a perfect solution; i am evaluating existing tools/methods for doing this in a better way. but this was my quick solution while doing this competition.\n\nI used excel to store high level ideas/comments/params for significant milestones.",
      "votes": null
    },
    {
      "id": "796567",
      "postDate": "04/03/2020 17:39:05",
      "content": "<p>Thank you for your reply.</p>",
      "rawMarkdown": "Thank you for your reply.",
      "votes": null
    },
    {
      "id": "797525",
      "postDate": "04/04/2020 16:30:13",
      "content": "<p><a href=\"/emrebayram\">@emrebayram</a> Well Explained....!! </p>",
      "rawMarkdown": "emrebayram Well Explained....!!",
      "votes": null
    },
    {
      "id": "797655",
      "postDate": "04/04/2020 19:04:35",
      "content": "<p>very helpful !</p>",
      "rawMarkdown": "<p>very helpful !</p>",
      "votes": null
    },
    {
      "id": "798597",
      "postDate": "04/05/2020 17:10:23",
      "content": "<p>Helpful 👍 </p>",
      "rawMarkdown": "Helpful 👍",
      "votes": null
    },
    {
      "id": "799200",
      "postDate": "04/06/2020 09:09:48",
      "content": "<p>Clear suggestions!</p>",
      "rawMarkdown": "Clear suggestions!",
      "votes": null
    },
    {
      "id": "799319",
      "postDate": "04/06/2020 11:14:47",
      "content": "<p>Great post, what I liked most is even though you didn't know much, you took the competition head on and accelerated the learning process. Good job!</p>",
      "rawMarkdown": "Great post, what I liked most is even though you didn't know much, you took the competition head on and accelerated the learning process. Good job!",
      "votes": null
    },
    {
      "id": "802643",
      "postDate": "04/09/2020 17:15:12",
      "content": "<p>That is very inspiring!  I was always wondering what should I do when I'm waiting and now I know I'm wrong and I should make a more specific pipeline for small data.</p>",
      "rawMarkdown": "That is very inspiring!  I was always wondering what should I do when I'm waiting and now I know I'm wrong and I should make a more specific pipeline for small data.",
      "votes": null
    },
    {
      "id": "806648",
      "postDate": "04/13/2020 23:12:24",
      "content": "<p>Worth a read, definitely!</p>",
      "rawMarkdown": "Worth a read, definitely!",
      "votes": null
    },
    {
      "id": "817847",
      "postDate": "04/23/2020 13:10:35",
      "content": "<p>Very helpful 👍 </p>",
      "rawMarkdown": "Very helpful 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 793730,
      "author_name": "eswarchandt",
      "author_url": "",
      "post_date": "04/01/2020 08:29:34",
      "content": "<p>Very well explained. Learning from mistakes is very important. \nHope you learned a lot from this competition</p>",
      "votes": null,
      "replies": [
        {
          "id": 793741,
          "author_name": "emrebayram",
          "author_url": "",
          "post_date": "04/01/2020 08:37:12",
          "content": "<p>Thanks! I definitely learned a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 793755,
      "author_name": "gwsong",
      "author_url": "",
      "post_date": "04/01/2020 08:47:33",
      "content": "<p>Your story is as good as success story. And also you made decent success in this competition.\nThank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 793789,
      "author_name": "matsuryu",
      "author_url": "",
      "post_date": "04/01/2020 09:25:59",
      "content": "<p>Very helpful. Honest person. Thx.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 793829,
      "author_name": "mohammadhatoum",
      "author_url": "",
      "post_date": "04/01/2020 10:19:05",
      "content": "<p>Thank you for sharing this great story. I can't believe how similar it is to mine. I also quit my job to focus on ML/DL.\nThere are lots of common mistakes: \n1) I also started with colab and I faced issues with its timeout and storage limitations. I haven't downloaded the whole dataset until last week where I started using AWS but without free credits :(!\n2) The mistakes that I didn't fix were the data separation and waiting for 3 hours for a model to finish.\nThere is one more mistake that I was working all alone, I think it would have been better to join a team of 1 or 2 people. This would have helped me avoid some mistakes or at least fix them earlier.</p>\n\n<p>In my case, I consider this was a lack of experience for me which I am sure I gained a lot of it now.</p>\n\n<p>On a side note, there is a \"mistake\" in point 4 because it has text from part 3 (So it is really important to keep a detailed journal/logs...)</p>\n\n<p>Thank you again and good luck with current and future competitions</p>",
      "votes": null,
      "replies": [
        {
          "id": 793859,
          "author_name": "emrebayram",
          "author_url": "",
          "post_date": "04/01/2020 11:06:54",
          "content": "<p>yes I have fixed it thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 793841,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "04/01/2020 10:35:21",
      "content": "<p>I wish I have sufficient conditions of all aspects to have a 6-month break like you :(. Anyway, thanks for sharing your great experience. It's always beneficial to review the goods and bads of one's self after the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 794107,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "04/01/2020 15:08:02",
      "content": "<p>Welcome and I think this is a great success story you should be proud of :) In terms of data, I couldn't agree more I spent 3 weeks on data and 1 week on modeling. In my case, I simply used OneNote to keep all my ideas and conceptualize them with pros/cons before even starting coding anything, of course discussions and kernels helped a lot to get this going. I also recommend <a href=\"https://nbdev.fast.ai/\">nbdev</a> if you like notebooks. I used it for a Kaggle competition for the first time but it improved my productivity x2-x3. It has the power of iterating fast with notebooks, keeping scientific journals in the same repo and also modularizing your code by converting notebooks into py files similar to a regular python library. Good luck on your journey!</p>",
      "votes": null,
      "replies": [
        {
          "id": 794227,
          "author_name": "emrebayram",
          "author_url": "",
          "post_date": "04/01/2020 17:10:22",
          "content": "<p>Wow nbdev looks/sounds great I will definitely give it a try!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 794123,
      "author_name": "debanga",
      "author_url": "",
      "post_date": "04/01/2020 15:26:41",
      "content": "<p>Very inspiring! All the best.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 794160,
      "author_name": "larsen0966",
      "author_url": "",
      "post_date": "04/01/2020 16:05:07",
      "content": "<p>Excellent Job!  When I work on my car and need to watch a video, I watch a professionally edited video from a pro so I make sure I know all the details of what parts and tolls I need and the procedure for the fix.  Then I watch an average guy like myself so it, so that I can see some reality.  Oh, it is going to actually take this long!  Watch out for this!  Ooops don't do that!</p>\n\n<p>Your post was excellent.</p>\n\n<p>So some questions, now.  Under \"Choosing the wrong environment/hardware\" you said you switch to your local machine.  You mean like your personal laptop/desktop?</p>\n\n<p>Under \"Not keeping a good scientific journal\".  Do you have a suggested format or a link to a journal you like?</p>\n\n<p>Keep-up the good work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 794274,
          "author_name": "emrebayram",
          "author_url": "",
          "post_date": "04/01/2020 17:45:18",
          "content": "<p>Thanks a lot!\nYes I used my personal desktop(GPU: gtx 1080 ti, SSD: 4TB NVMe, OS: ubuntu). \nMy observation is; even the cheapest AWS EC2 GPU instances(like g3s) have better GPUs(especially in terms of memory) than my desktop but their EBS disks are utilized for many small IO requests(very low latency) but not as good as NVMe SSDs with m.2 interface in terms of continuous read/write speed. \nTo have an EC2 instance with NVMe SSD disks with enough storage size you need to pay around $5-$20 / hour and it would be impossible for me to fit into free credit budget if I used that. (you can see the EC2 instance lists here: <a href=\"https://aws.amazon.com/ec2/pricing/on-demand/\">https://aws.amazon.com/ec2/pricing/on-demand/</a> , I realized using an instance with NVMe SSD vs EBS storage has great performance difference if you are dealing with a lot of medium sized files. and instances with NVMe SSDs with enough size are expensive.\nBut using AWS was not a mistake(I just wanted to point the importance of IO speed in this competition); it was fast enough. \nTrying to handle this competition with colab+google drive on the other hand; was a clear mistake for me. </p>\n\n<p>About the scientific journal; I don't know a good tool. I used a very basic structure; \n- I created a dictionary with all the parameters(lr,wd,freeze_layer,batch_size, augmentation probabilities&amp;parameters, pretrained model pointer, train_data_path, val_data_path, train part nos, val_part_nos, train loss, val loss, epoch no,  etc etc.) and whenever I save the model I also saved that dictionary together with it. \n- I also created a spreadsheet on excel to record all these with my comments.</p>\n\n<p>But <a href=\"/keremt\">@keremt</a> shared in his comment below a link to a library(nbdev) I will give it a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 794404,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/01/2020 19:36:58",
          "content": "<p>Thanks for sharing your journey. As for the \"journaling\" part, I have used <a href=\"https://mlflow.org/\"><strong>mlflow</strong></a> (mostly at work) and it is very useful to record metadata about your experiments. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 796486,
          "author_name": "aknirala",
          "author_url": "",
          "post_date": "04/03/2020 16:13:44",
          "content": "<p>Thanks <a href=\"/emrebayram\">@emrebayram</a> for sharing this inspiring journal of yours. I fully resonate with \"..doing more experiments is everything...\". Outcomes of the experiment gives so many ideas.</p>\n\n<p>I have a much broader question regarding keeping a journal (sorry to pull on this).  Keeping journal, though sounds simple, is not simple at all. For example, storing one line per model in excel seems like a nice idea, but when I try something, I make so many small-small changes and also tinker with architecture etc. At this point, excel is no longer a good option (Still I believe the best. I specced a whole application to do this….someday I'll build it as well.). </p>\n\n<p>A month back, I created some variance plots on fake and real data. Now I can’t even find where it is stored!! (Feeling frustrated, I came back and thought of asking... whats the harm anyway)</p>\n\n<p>So could you elaborate (more) on how you journal stuff? I believe you also keep track of your ideas somewhere. Do you have any sorting mechanism etc as well? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 796563,
          "author_name": "emrebayram",
          "author_url": "",
          "post_date": "04/03/2020 17:32:32",
          "content": "<p>Thanks <a href=\"/aknirala\">@aknirala</a> , I don't have a perfect solution for keeping a journal. \" I make so many small-small changes and also tinker with architecture etc. \" was definitely a big problem for me too. But what worked for me best is; I created a global dictionary(something called global_params); and whenever i am playing with something (training files, augmentation, architecture, epoch numbers, random probabilities, pre-processing variations, pretrain file, frozen layers everything); i carried it to that dictionary and modified there and read from there. For example if i realized i am playing with some code part; i made it configurable and carried these configuration params to the global params dictionary. \nThis might not be a good practice in terms of software engineering but here worked for me. And I had a simple model save method and whenever I am saving the model state dict and optimizer state dict; I saved this global params with it together with my loss value. so whatever i played is stored together with the model. So for every point; i saved 2 files like this:\n 92258335 Mar 31 21:03 0.32895452830785804_model.pth\n183316294 Mar 31 21:05 0.32895452830785804_rsm_pt.pth\nfile names contains the loss value; it was easier that way to see for me. model.pth is the model state dict. rsm_pt stands for resume point and contains:\noptimizer state dict, all params; basically everything.\nAnd I pushed these files to git together with the current code. \nThat way; if i wanted to see what i have tried I loaded and printed the global params variable and observed it(it is basically a summary of the current code/all params/methodologies). and if i wanted to resume to a point to move from there; i pulled that version of code together with model file and global params.\nThis is not a perfect solution; i am evaluating existing tools/methods for doing this in a better way. but this was my quick solution while doing this competition.</p>\n\n<p>I used excel to store high level ideas/comments/params for significant milestones.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 796567,
          "author_name": "aknirala",
          "author_url": "",
          "post_date": "04/03/2020 17:39:05",
          "content": "<p>Thank you for your reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 794652,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "04/02/2020 01:35:02",
      "content": "<p>Thank you for sharing this. I can relate with you when it comes to #3 and #5....especially #3 (I'm a disaster when it comes to organization). About framework, I feel that keras is the easiest to get into compared to plain tensorflow or pytorch (i'm also a newbie in this field trying to learn. I code in C and assembly for a living)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 797525,
      "author_name": "saurav9786",
      "author_url": "",
      "post_date": "04/04/2020 16:30:13",
      "content": "<p><a href=\"/emrebayram\">@emrebayram</a> Well Explained....!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 797655,
      "author_name": "saadas96",
      "author_url": "",
      "post_date": "04/04/2020 19:04:35",
      "content": "<p>very helpful !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 798597,
      "author_name": "prasadpatil99",
      "author_url": "",
      "post_date": "04/05/2020 17:10:23",
      "content": "<p>Helpful 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 799200,
      "author_name": "tiurii",
      "author_url": "",
      "post_date": "04/06/2020 09:09:48",
      "content": "<p>Clear suggestions!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 799319,
      "author_name": "amoghjrules",
      "author_url": "",
      "post_date": "04/06/2020 11:14:47",
      "content": "<p>Great post, what I liked most is even though you didn't know much, you took the competition head on and accelerated the learning process. Good job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 802643,
      "author_name": "zowlex",
      "author_url": "",
      "post_date": "04/09/2020 17:15:12",
      "content": "<p>That is very inspiring!  I was always wondering what should I do when I'm waiting and now I know I'm wrong and I should make a more specific pipeline for small data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 806648,
      "author_name": "drabhinav",
      "author_url": "",
      "post_date": "04/13/2020 23:12:24",
      "content": "<p>Worth a read, definitely!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 817847,
      "author_name": "kayadagli",
      "author_url": "",
      "post_date": "04/23/2020 13:10:35",
      "content": "<p>Very helpful 👍 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "793721": "It is always good and inspiring to read success stories.\n\nBut I also value failure stories so I would like to share one:)\n\nWell, this is not exactly a failure(actually I am kind of happy with my current score which is a lot better than I expected at the beginning as a DL noob) story but a list of mistakes I did as a noob(skip to the \"mistakes\" part if you are not interested in the \"story\" part).\n\nLast August I decided to stop being a full time employee and take a funemployement break for 6 months. Then I quit my job at November.\n\nI wanted to spend time with my family, learn new things and do whatever I like for 6 months:)\n\nMachine learning, specifically deep learning was one of the things I wanted/tried to learn in the past but never had enough time to invest in. I built a few basic classic models in the past but it was basically just following some other people' footsteps without really knowing what i am doing.\n\nSo during my break I wanted to revisit it and at least learn the basics.. the fastai videos by Jeremy Howard really motivated me as a starter tutorial.\nBut I have never been good at learning from books/tutorials/videos only. I always needed a target and some kind of deadline to motivate me. Luckily I have seen this competition and thought it was a good opportunity for my learning process.\n\nBefore this tournament; I did not like/know python much(i still don't know it well but at least now I like it:)); I haven't used pytorch/tensorflor/keras etc before.\n\n\n**So please read the following \"mistakes\"/\"learnings\" knowing that they are not coming from an expert:)**\n\n\nAnyways as of today; when I look through past few months; I see clear mistakes I did as a noob; and I want to share them in this post.\n\n\n\n### Mistakes\n\nMy learning below can still be noob conclusions but it is what I think as of today. So if you have comments/advice on my takes; you are very welcome.\n\ntl;dr everything that makes you waste time is a mistake because this loss time is so valuable for making more experiments. And doing more experiments is everything(or 95% of everything).\n\n\n### 1- Choosing the wrong environment/hardware.\n\nThe first environment I have chosen for this tournament was Google colab + google drive. And the google drive part was a huge mistake. I subscribed to some terabyte plan and I thought it would be ok to use google drive because you can access google drive data from colab and what could go wrong? \nWell; never thought I would hit some kind of daily download size limit(I did not even know it existed) on google drive because whenever you mount google drive from colab; and start using data it is counted as a \"download\" although both sides are Google. \nOnce I started to hit the limits; I needed to redo things from scratch on AWS(thanks to the free credits) \nAlso I realized that the io performance in this competition is as important as the GPU performance so I switched to my local machine which has worse GPU than AWS instance but have 4TB NVMe SSD which is a lot faster than AWS disks. And my training times became X3 faster than AWS. I should have foreseen it earlier.\n\n\n### 2- Spending time on choosing which framework too much.\n\nAfter spending some significant time on understanding some basics of fastai/keras/tensorflow/pytorch; i have picked pytorch and progressed with it. Now when I look back; I think it was not necessary to spend that much time on framework selection; because now I think what matters most was understanding the problem and the data. \n\n\n### 3- Not keeping a good scientific journal.\n\na month ago I realized; the biggest mistake I have done was not keeping a good journal. I took notes, saved notebooks etc but this was not enough. Because after experimenting many many things; you lose track of \"what was working/what was not working/what is next to try\" and sometimes you repeat yourself. So it is really important to keep a detailed journal/logs; it can be an excel table, a list of notebooks whatever; but it should be kept in a structured and very well organized way I think. Later I started to save all parameters and preliminary steps(like everything from how I prepared the datasets to results and whatever there is) together with the model file and also started to keep a table. After I start doing this; it was a lot easier to see the picture.\n\n\n### 4- Not spending more time on data separation.\n\nI knew it is really important to keep your validation and test sets isolated from the training set but  (I don't know if it is related to this competition) making sure that you are really doing this was really challenging. And if I was redoing this competition I would have spent more time at the beginning to understand the data better(how many different types of fakizations there, similar actors, which folders are close to the other folders in terms of similarity etc) and build the necessary tools to make sure they are isolated in a better way. Towards the end of the competition I built some utility functions to check this and it really helped but i should have spent more time to have more robust utilities to make sure of it.\n\n\n### 5- Waiting.\n\nThis was also one of my biggest mistakes. Fortunately; I fixed it early. At the beginning; several times I believed I found a good methodology; but I tried it on a dataset which is bigger than needed(not the full dataset but still bigger than needed). So I needed to wait to get the results of the experiments. After some time I realized this is a clear mistake. And whenever I found a method to build a dataset I also built a \"small\" and \"medium\" versions of it which contains enough characteristics but give the results earlier. Now I believe if you are waiting too much; you are making a mistake. Finally I had a pipeline with 3 stages; first I try the method in the smallest test(10 minutes max) then medium set(2 hours max) then full set. So I catch mistakes/problems earlier.\n\nAnd finally; thanks for the competition it was a huge fun and a lot of learning!",
    "793730": "Very well explained. Learning from mistakes is very important. \nHope you learned a lot from this competition",
    "793741": "Thanks! I definitely learned a lot.",
    "793755": "Your story is as good as success story. And also you made decent success in this competition.\nThank you.",
    "793789": "Very helpful. Honest person. Thx.",
    "793829": "Thank you for sharing this great story. I can't believe how similar it is to mine. I also quit my job to focus on ML/DL.\nThere are lots of common mistakes: \n1) I also started with colab and I faced issues with its timeout and storage limitations. I haven't downloaded the whole dataset until last week where I started using AWS but without free credits :(!\n2) The mistakes that I didn't fix were the data separation and waiting for 3 hours for a model to finish.\nThere is one more mistake that I was working all alone, I think it would have been better to join a team of 1 or 2 people. This would have helped me avoid some mistakes or at least fix them earlier.\n\nIn my case, I consider this was a lack of experience for me which I am sure I gained a lot of it now.\n\nOn a side note, there is a \"mistake\" in point 4 because it has text from part 3 (So it is really important to keep a detailed journal/logs...)\n\nThank you again and good luck with current and future competitions",
    "793841": "I wish I have sufficient conditions of all aspects to have a 6-month break like you :(. Anyway, thanks for sharing your great experience. It's always beneficial to review the goods and bads of one's self after the competition.",
    "793859": "yes I have fixed it thanks!",
    "794107": "Welcome and I think this is a great success story you should be proud of :) In terms of data, I couldn't agree more I spent 3 weeks on data and 1 week on modeling. In my case, I simply used OneNote to keep all my ideas and conceptualize them with pros/cons before even starting coding anything, of course discussions and kernels helped a lot to get this going. I also recommend [nbdev](https://nbdev.fast.ai/) if you like notebooks. I used it for a Kaggle competition for the first time but it improved my productivity x2-x3. It has the power of iterating fast with notebooks, keeping scientific journals in the same repo and also modularizing your code by converting notebooks into py files similar to a regular python library. Good luck on your journey!",
    "794123": "Very inspiring! All the best.",
    "794160": "Excellent Job!  When I work on my car and need to watch a video, I watch a professionally edited video from a pro so I make sure I know all the details of what parts and tolls I need and the procedure for the fix.  Then I watch an average guy like myself so it, so that I can see some reality.  Oh, it is going to actually take this long!  Watch out for this!  Ooops don't do that!\n\nYour post was excellent.\n\nSo some questions, now.  Under \"Choosing the wrong environment/hardware\" you said you switch to your local machine.  You mean like your personal laptop/desktop?\n\nUnder \"Not keeping a good scientific journal\".  Do you have a suggested format or a link to a journal you like?\n\nKeep-up the good work.",
    "794227": "Wow nbdev looks/sounds great I will definitely give it a try!",
    "794274": "Thanks a lot!\nYes I used my personal desktop(GPU: gtx 1080 ti, SSD: 4TB NVMe, OS: ubuntu). \nMy observation is; even the cheapest AWS EC2 GPU instances(like g3s) have better GPUs(especially in terms of memory) than my desktop but their EBS disks are utilized for many small IO requests(very low latency) but not as good as NVMe SSDs with m.2 interface in terms of continuous read/write speed. \nTo have an EC2 instance with NVMe SSD disks with enough storage size you need to pay around $5-$20 / hour and it would be impossible for me to fit into free credit budget if I used that. (you can see the EC2 instance lists here: https://aws.amazon.com/ec2/pricing/on-demand/ , I realized using an instance with NVMe SSD vs EBS storage has great performance difference if you are dealing with a lot of medium sized files. and instances with NVMe SSDs with enough size are expensive.\nBut using AWS was not a mistake(I just wanted to point the importance of IO speed in this competition); it was fast enough. \nTrying to handle this competition with colab+google drive on the other hand; was a clear mistake for me. \n\nAbout the scientific journal; I don't know a good tool. I used a very basic structure; \n- I created a dictionary with all the parameters(lr,wd,freeze_layer,batch_size, augmentation probabilities¶meters, pretrained model pointer, train_data_path, val_data_path, train part nos, val_part_nos, train loss, val loss, epoch no,  etc etc.) and whenever I save the model I also saved that dictionary together with it. \n- I also created a spreadsheet on excel to record all these with my comments.\n\nBut @keremt shared in his comment below a link to a library(nbdev) I will give it a try.",
    "794404": "Thanks for sharing your journey. As for the \"journaling\" part, I have used [**mlflow**](https://mlflow.org/) (mostly at work) and it is very useful to record metadata about your experiments.",
    "794652": "Thank you for sharing this. I can relate with you when it comes to #3 and #5....especially #3 (I'm a disaster when it comes to organization). About framework, I feel that keras is the easiest to get into compared to plain tensorflow or pytorch (i'm also a newbie in this field trying to learn. I code in C and assembly for a living)",
    "796486": "Thanks @emrebayram for sharing this inspiring journal of yours. I fully resonate with \"..doing more experiments is everything...\". Outcomes of the experiment gives so many ideas.\n\nI have a much broader question regarding keeping a journal (sorry to pull on this).  Keeping journal, though sounds simple, is not simple at all. For example, storing one line per model in excel seems like a nice idea, but when I try something, I make so many small-small changes and also tinker with architecture etc. At this point, excel is no longer a good option (Still I believe the best. I specced a whole application to do this….someday I'll build it as well.). \n\nA month back, I created some variance plots on fake and real data. Now I can’t even find where it is stored!! (Feeling frustrated, I came back and thought of asking... whats the harm anyway)\n\nSo could you elaborate (more) on how you journal stuff? I believe you also keep track of your ideas somewhere. Do you have any sorting mechanism etc as well?",
    "796563": "Thanks @aknirala , I don't have a perfect solution for keeping a journal. \" I make so many small-small changes and also tinker with architecture etc. \" was definitely a big problem for me too. But what worked for me best is; I created a global dictionary(something called global_params); and whenever i am playing with something (training files, augmentation, architecture, epoch numbers, random probabilities, pre-processing variations, pretrain file, frozen layers everything); i carried it to that dictionary and modified there and read from there. For example if i realized i am playing with some code part; i made it configurable and carried these configuration params to the global params dictionary. \nThis might not be a good practice in terms of software engineering but here worked for me. And I had a simple model save method and whenever I am saving the model state dict and optimizer state dict; I saved this global params with it together with my loss value. so whatever i played is stored together with the model. So for every point; i saved 2 files like this:\n 92258335 Mar 31 21:03 0.32895452830785804_model.pth\n183316294 Mar 31 21:05 0.32895452830785804_rsm_pt.pth\nfile names contains the loss value; it was easier that way to see for me. model.pth is the model state dict. rsm_pt stands for resume point and contains:\noptimizer state dict, all params; basically everything.\nAnd I pushed these files to git together with the current code. \nThat way; if i wanted to see what i have tried I loaded and printed the global params variable and observed it(it is basically a summary of the current code/all params/methodologies). and if i wanted to resume to a point to move from there; i pulled that version of code together with model file and global params.\nThis is not a perfect solution; i am evaluating existing tools/methods for doing this in a better way. but this was my quick solution while doing this competition.\n\nI used excel to store high level ideas/comments/params for significant milestones.",
    "796567": "Thank you for your reply.",
    "797525": "emrebayram Well Explained....!!",
    "797655": "<p>very helpful !</p>",
    "798597": "Helpful 👍",
    "799200": "Clear suggestions!",
    "799319": "Great post, what I liked most is even though you didn't know much, you took the competition head on and accelerated the learning process. Good job!",
    "802643": "That is very inspiring!  I was always wondering what should I do when I'm waiting and now I know I'm wrong and I should make a more specific pipeline for small data.",
    "806648": "Worth a read, definitely!",
    "817847": "Very helpful 👍"
  },
  "source": "meta"
}