{
  "id": 226288,
  "title": "What could you have done better?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/226288",
  "author_name": "",
  "post_date": "2021-03-16T00:01:00.027404200Z",
  "votes": 19,
  "comment_count": 14,
  "views": 0,
  "content": "<p>After every competition, I try to review a bit and consider what I could have done better in terms of process to have more success the next time. How would I have found the insight or technique to create a better solution or come to the same solution faster? Would be interested in hearing what other people feel they could have improved on process-wise. </p>\n<p>For me this competition, I think one of my biggest misses was my poor logging setup. I did a poor job of keeping track of training and validation results across all of the metrics of interest. Sometimes accidentally overwriting results or leaving more information in notebooks than in the actual logs of the training procedure. For example, only showing the summary validation loss and auc rather than the loss and auc of individual columns. </p>\n<p>A few times I would iterate on notebooks accidentally in place rather than moving to a new notebook when doing an experiment and would no longer have the detailed information from training something or even worse have a good result, but no longer have the code that generated it. This is one of the big downfalls of notebooks IMO. I typically can operate fine with them just copying the notebook before I make changes, but have to be disciplined with it. </p>\n<p>I've seen some others do a few things to prevent this from happening:</p>\n<ol>\n<li><p>store models, logs, validation indexes and validation/test predictions in their own folders</p>\n<ul>\n<li>This keeps everything together so when you are selecting models to use for submission and analyzing what worked and what didn't you have a direct path back to the original source. For me, I started building a hierarchy where I had model/name_of_experiment/fold so I could look at all of my resnet18s and the variations I tried with them and then also compare specific folds to each other. I tried to keep each individual file specifically named as well so if I ever transferred to kaggle or other places it was clear what file was what. </li></ul></li>\n<li><p>New folders instead of overwriting</p>\n<ul>\n<li>one of the issues I ran into was overwriting previous logs or outputs from notebooks. There are two different things I think could be useful here:<ul>\n<li>at the end of every epoch or some other unit push a copy of the notebook to the same folder with all of the other results. This bundles the code and the outputs so even if you later rewrite code in the existing notebook you have a version to roll back to that even has the recordings of all of the cells. sort of a poor mans version control that happens automatically during training</li>\n<li>every new rerun of some code, rather than overwriting anything just iterate a counter and make a new folder to put everything in. Storage is very cheap, save everything because it will save time down the line. </li></ul></li></ul></li>\n</ol>\n<p>One of the things that I ran into that kind of clobbered some of my original logging plans was when I wanted to recontinue training I would lose some aspect of my logs. hindsight 20/20 I should have looked at my various procedures and planned them out more thoughtfully early on. Maybe I finally invest the time in setting up a proper solution like neptune.ai or sacred for the next one. </p>",
  "messages": [
    {
      "id": "1239708",
      "postDate": "03/16/2021 00:01:00",
      "content": "<p>After every competition, I try to review a bit and consider what I could have done better in terms of process to have more success the next time. How would I have found the insight or technique to create a better solution or come to the same solution faster? Would be interested in hearing what other people feel they could have improved on process-wise. </p>\n<p>For me this competition, I think one of my biggest misses was my poor logging setup. I did a poor job of keeping track of training and validation results across all of the metrics of interest. Sometimes accidentally overwriting results or leaving more information in notebooks than in the actual logs of the training procedure. For example, only showing the summary validation loss and auc rather than the loss and auc of individual columns. </p>\n<p>A few times I would iterate on notebooks accidentally in place rather than moving to a new notebook when doing an experiment and would no longer have the detailed information from training something or even worse have a good result, but no longer have the code that generated it. This is one of the big downfalls of notebooks IMO. I typically can operate fine with them just copying the notebook before I make changes, but have to be disciplined with it. </p>\n<p>I've seen some others do a few things to prevent this from happening:</p>\n<ol>\n<li><p>store models, logs, validation indexes and validation/test predictions in their own folders</p>\n<ul>\n<li>This keeps everything together so when you are selecting models to use for submission and analyzing what worked and what didn't you have a direct path back to the original source. For me, I started building a hierarchy where I had model/name_of_experiment/fold so I could look at all of my resnet18s and the variations I tried with them and then also compare specific folds to each other. I tried to keep each individual file specifically named as well so if I ever transferred to kaggle or other places it was clear what file was what. </li></ul></li>\n<li><p>New folders instead of overwriting</p>\n<ul>\n<li>one of the issues I ran into was overwriting previous logs or outputs from notebooks. There are two different things I think could be useful here:<ul>\n<li>at the end of every epoch or some other unit push a copy of the notebook to the same folder with all of the other results. This bundles the code and the outputs so even if you later rewrite code in the existing notebook you have a version to roll back to that even has the recordings of all of the cells. sort of a poor mans version control that happens automatically during training</li>\n<li>every new rerun of some code, rather than overwriting anything just iterate a counter and make a new folder to put everything in. Storage is very cheap, save everything because it will save time down the line. </li></ul></li></ul></li>\n</ol>\n<p>One of the things that I ran into that kind of clobbered some of my original logging plans was when I wanted to recontinue training I would lose some aspect of my logs. hindsight 20/20 I should have looked at my various procedures and planned them out more thoughtfully early on. Maybe I finally invest the time in setting up a proper solution like neptune.ai or sacred for the next one. </p>",
      "rawMarkdown": "After every competition, I try to review a bit and consider what I could have done better in terms of process to have more success the next time. How would I have found the insight or technique to create a better solution or come to the same solution faster? Would be interested in hearing what other people feel they could have improved on process-wise. \n\nFor me this competition, I think one of my biggest misses was my poor logging setup. I did a poor job of keeping track of training and validation results across all of the metrics of interest. Sometimes accidentally overwriting results or leaving more information in notebooks than in the actual logs of the training procedure. For example, only showing the summary validation loss and auc rather than the loss and auc of individual columns. \n\nA few times I would iterate on notebooks accidentally in place rather than moving to a new notebook when doing an experiment and would no longer have the detailed information from training something or even worse have a good result, but no longer have the code that generated it. This is one of the big downfalls of notebooks IMO. I typically can operate fine with them just copying the notebook before I make changes, but have to be disciplined with it. \n\nI've seen some others do a few things to prevent this from happening:\n1. store models, logs, validation indexes and validation/test predictions in their own folders\n\n  - This keeps everything together so when you are selecting models to use for submission and analyzing what worked and what didn't you have a direct path back to the original source. For me, I started building a hierarchy where I had model/name_of_experiment/fold so I could look at all of my resnet18s and the variations I tried with them and then also compare specific folds to each other. I tried to keep each individual file specifically named as well so if I ever transferred to kaggle or other places it was clear what file was what. \n\n2. New folders instead of overwriting\n\n  - one of the issues I ran into was overwriting previous logs or outputs from notebooks. There are two different things I think could be useful here:\n     - at the end of every epoch or some other unit push a copy of the notebook to the same folder with all of the other results. This bundles the code and the outputs so even if you later rewrite code in the existing notebook you have a version to roll back to that even has the recordings of all of the cells. sort of a poor mans version control that happens automatically during training\n     -  every new rerun of some code, rather than overwriting anything just iterate a counter and make a new folder to put everything in. Storage is very cheap, save everything because it will save time down the line. \n\nOne of the things that I ran into that kind of clobbered some of my original logging plans was when I wanted to recontinue training I would lose some aspect of my logs. hindsight 20/20 I should have looked at my various procedures and planned them out more thoughtfully early on. Maybe I finally invest the time in setting up a proper solution like neptune.ai or sacred for the next one.",
      "votes": null
    },
    {
      "id": "1239710",
      "postDate": "03/16/2021 00:07:22",
      "content": "<p>Neptune.ai is the best :D. Takes 5 minutes to learn and use.</p>",
      "rawMarkdown": "Neptune.ai is the best :D. Takes 5 minutes to learn and use.",
      "votes": null
    },
    {
      "id": "1239720",
      "postDate": "03/16/2021 00:32:10",
      "content": "<p>I'll have to try it out. I looked at solutions a couple years ago but never committed to them because it just seemed like too much overhead just to do some logging and analysis</p>",
      "rawMarkdown": "I'll have to try it out. I looked at solutions a couple years ago but never committed to them because it just seemed like too much overhead just to do some logging and analysis",
      "votes": null
    },
    {
      "id": "1239722",
      "postDate": "03/16/2021 00:35:46",
      "content": "<blockquote>\n  <p><strong>After</strong> every competition. </p>\n</blockquote>\n<p>Yet you post before its over…</p>",
      "rawMarkdown": "> **After** every competition. \n\nYet you post before its over...",
      "votes": null
    },
    {
      "id": "1239724",
      "postDate": "03/16/2021 00:39:38",
      "content": "<p>Over for <strong>me</strong>. Given up on this one</p>",
      "rawMarkdown": "Over for **me**. Given up on this one",
      "votes": null
    },
    {
      "id": "1239775",
      "postDate": "03/16/2021 02:19:33",
      "content": "<p>Welcome to the world of </p>\n<h1>DevOps</h1>\n<p>interesting to see that kaggle competitions have become difficult that numerous experiments are done and proper logging become a must.</p>\n<p>in short, even kaggling needs devops</p>\n<hr>\n<p>last time, i was just developing messy code at the beginning. then i would clean up my code and experiments later if i have good solution.</p>\n<hr>\n<p>now i find that this approach is not efficient nor sufficient. I have to keep good habits of developing good code, good log and good experiments report #right at the start#. I don't have time to go back and re-organize.</p>\n<hr>\n<p>you can try this:</p>\n<ul>\n<li>separate code from input data and output result.</li>\n<li>first line in your code run it runs<ul>\n<li>zip the whole project code (for repeatability) and save with timestamp</li>\n<li>create a new output folder </li>\n<li>start a logfile<br>\nsince i save my code every time it runs, i can always revert it back</li></ul></li>\n</ul>\n<p>you need to have good diskspace. for each kaggle competition, i typically endup with 200 to 500gb of rubbish</p>",
      "rawMarkdown": "Welcome to the world of \n\n# DevOps\n\n\ninteresting to see that kaggle competitions have become difficult that numerous experiments are done and proper logging become a must.\n\nin short, even kaggling needs devops\n\n---\n\nlast time, i was just developing messy code at the beginning. then i would clean up my code and experiments later if i have good solution.\n\n---\nnow i find that this approach is not efficient nor sufficient. I have to keep good habits of developing good code, good log and good experiments report #right at the start#. I don't have time to go back and re-organize.\n\n---\n\nyou can try this:\n- separate code from input data and output result.\n- first line in your code run it runs\n  - zip the whole project code (for repeatability) and save with timestamp\n - create a new output folder \n  - start a logfile\nsince i save my code every time it runs, i can always revert it back\n\nyou need to have good diskspace. for each kaggle competition, i typically endup with 200 to 500gb of rubbish",
      "votes": null
    },
    {
      "id": "1239776",
      "postDate": "03/16/2021 02:20:49",
      "content": "<p>Yeah, this is really the part of ML that I don't like but I realize it is becoming more and more necessary. </p>",
      "rawMarkdown": "Yeah, this is really the part of ML that I don't like but I realize it is becoming more and more necessary.",
      "votes": null
    },
    {
      "id": "1239807",
      "postDate": "03/16/2021 03:17:02",
      "content": "<p>In terms of version control and folders etc.  <a href=\"https://www.kaggle.com/product-feedback/221448\" target=\"_blank\">Feature Launch Open Notebooks from Github</a> may be useful.  Have not tried this yet, but private repository and export to Github are coming soon. </p>",
      "rawMarkdown": "In terms of version control and folders etc.  [Feature Launch Open Notebooks from Github](https://www.kaggle.com/product-feedback/221448) may be useful.  Have not tried this yet, but private repository and export to Github are coming soon.",
      "votes": null
    },
    {
      "id": "1239813",
      "postDate": "03/16/2021 03:25:13",
      "content": "<p>Interesting. I know I have heard Philipp and Christof mention github actions as some way to push code and have it trigger new training. I'm sure there is some good workflow that could be created with all the various tools but I haven't spent the time to figure it out. </p>",
      "rawMarkdown": "Interesting. I know I have heard Philipp and Christof mention github actions as some way to push code and have it trigger new training. I'm sure there is some good workflow that could be created with all the various tools but I haven't spent the time to figure it out.",
      "votes": null
    },
    {
      "id": "1239837",
      "postDate": "03/16/2021 04:09:59",
      "content": "<p>Yea, MLOps is a thing…People are being paid top dollars to do it lol 😂</p>",
      "rawMarkdown": "Yea, MLOps is a thing...People are being paid top dollars to do it lol 😂",
      "votes": null
    },
    {
      "id": "1239920",
      "postDate": "03/16/2021 05:55:21",
      "content": "<p>A week before when I started submissions I was just submitting notebooks with some modifications and taking look at LB if it is improved or not. After 2 days of doing the same, I got confused at a point, not able to decide which 96.8 gave me a boost in rank, which modification(like ensemble ratio, parameters, weights) worked best, and opening notebooks, again and again, was time wasted. So to solve this issue:</p>\n<ul>\n<li>I started naming notebooks properly with modifications I have done in this version.</li>\n<li>Giving version names properly like including date of submission and timing(approx) it ran for, rank before and after(if changed).</li>\n<li>Below there's a very good feature provided by Kaggle for writing descriptions which I noticed and started using from this competition only(might be old for other people). I used to write like what specific hyperparameter impacted LB. Also used to write if there was an improvement or not.<br>\n<img src=\"https://raw.githubusercontent.com/sanchitvj/Kaggle-Competitions/main/Ranzcr%20CLIP/kaggle%20sub_li.jpg\" alt=\"sub img\"></li>\n</ul>\n<p>I hope this is helpful.</p>",
      "rawMarkdown": "A week before when I started submissions I was just submitting notebooks with some modifications and taking look at LB if it is improved or not. After 2 days of doing the same, I got confused at a point, not able to decide which 96.8 gave me a boost in rank, which modification(like ensemble ratio, parameters, weights) worked best, and opening notebooks, again and again, was time wasted. So to solve this issue:\n- I started naming notebooks properly with modifications I have done in this version.\n- Giving version names properly like including date of submission and timing(approx) it ran for, rank before and after(if changed).\n- Below there's a very good feature provided by Kaggle for writing descriptions which I noticed and started using from this competition only(might be old for other people). I used to write like what specific hyperparameter impacted LB. Also used to write if there was an improvement or not.\n![sub img](https://raw.githubusercontent.com/sanchitvj/Kaggle-Competitions/main/Ranzcr%20CLIP/kaggle%20sub_li.jpg)\n\nI hope this is helpful.",
      "votes": null
    },
    {
      "id": "1240209",
      "postDate": "03/16/2021 09:36:21",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> had the same logic before. Now I realized that investing a few hours in learning how to use some infrastructure / tracking tool will save many more hours later :) I have also started using Neptune.ai this year and it works great.</p>",
      "rawMarkdown": "ryches had the same logic before. Now I realized that investing a few hours in learning how to use some infrastructure / tracking tool will save many more hours later :) I have also started using Neptune.ai this year and it works great.",
      "votes": null
    },
    {
      "id": "1240230",
      "postDate": "03/16/2021 10:01:23",
      "content": "<p><a href=\"https://github.com/marketplace/actions/push-kaggle-dataset\" target=\"_blank\">https://github.com/marketplace/actions/push-kaggle-dataset</a></p>",
      "rawMarkdown": "https://github.com/marketplace/actions/push-kaggle-dataset",
      "votes": null
    },
    {
      "id": "1261354",
      "postDate": "04/03/2021 01:37:38",
      "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> :D I took 30 minutes and still fidgeting… Do you have any nice notebook of yours to share for tracking/storing metrics/processes using neptune.ai? I even wanted to store augmentations params…</p>",
      "rawMarkdown": "underwearfitting :D I took 30 minutes and still fidgeting... Do you have any nice notebook of yours to share for tracking/storing metrics/processes using neptune.ai? I even wanted to store augmentations params...",
      "votes": null
    },
    {
      "id": "1261369",
      "postDate": "04/03/2021 02:23:24",
      "content": "<p>I used neptune with pytorchlightning. Here's an example: <a href=\"https://github.com/kagglesintracking/kaggle-Cassava-Leaf-Disease-Classification/blob/main/src/main.py\" target=\"_blank\">https://github.com/kagglesintracking/kaggle-Cassava-Leaf-Disease-Classification/blob/main/src/main.py</a></p>",
      "rawMarkdown": "I used neptune with pytorchlightning. Here's an example: https://github.com/kagglesintracking/kaggle-Cassava-Leaf-Disease-Classification/blob/main/src/main.py",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1239710,
      "author_name": "underwearfitting",
      "author_url": "",
      "post_date": "03/16/2021 00:07:22",
      "content": "<p>Neptune.ai is the best :D. Takes 5 minutes to learn and use.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239720,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/16/2021 00:32:10",
          "content": "<p>I'll have to try it out. I looked at solutions a couple years ago but never committed to them because it just seemed like too much overhead just to do some logging and analysis</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240209,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "03/16/2021 09:36:21",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> had the same logic before. Now I realized that investing a few hours in learning how to use some infrastructure / tracking tool will save many more hours later :) I have also started using Neptune.ai this year and it works great.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261354,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "04/03/2021 01:37:38",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> :D I took 30 minutes and still fidgeting… Do you have any nice notebook of yours to share for tracking/storing metrics/processes using neptune.ai? I even wanted to store augmentations params…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261369,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "04/03/2021 02:23:24",
          "content": "<p>I used neptune with pytorchlightning. Here's an example: <a href=\"https://github.com/kagglesintracking/kaggle-Cassava-Leaf-Disease-Classification/blob/main/src/main.py\" target=\"_blank\">https://github.com/kagglesintracking/kaggle-Cassava-Leaf-Disease-Classification/blob/main/src/main.py</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239722,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "03/16/2021 00:35:46",
      "content": "<blockquote>\n  <p><strong>After</strong> every competition. </p>\n</blockquote>\n<p>Yet you post before its over…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239724,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/16/2021 00:39:38",
          "content": "<p>Over for <strong>me</strong>. Given up on this one</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239775,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/16/2021 02:19:33",
      "content": "<p>Welcome to the world of </p>\n<h1>DevOps</h1>\n<p>interesting to see that kaggle competitions have become difficult that numerous experiments are done and proper logging become a must.</p>\n<p>in short, even kaggling needs devops</p>\n<hr>\n<p>last time, i was just developing messy code at the beginning. then i would clean up my code and experiments later if i have good solution.</p>\n<hr>\n<p>now i find that this approach is not efficient nor sufficient. I have to keep good habits of developing good code, good log and good experiments report #right at the start#. I don't have time to go back and re-organize.</p>\n<hr>\n<p>you can try this:</p>\n<ul>\n<li>separate code from input data and output result.</li>\n<li>first line in your code run it runs<ul>\n<li>zip the whole project code (for repeatability) and save with timestamp</li>\n<li>create a new output folder </li>\n<li>start a logfile<br>\nsince i save my code every time it runs, i can always revert it back</li></ul></li>\n</ul>\n<p>you need to have good diskspace. for each kaggle competition, i typically endup with 200 to 500gb of rubbish</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239776,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/16/2021 02:20:49",
          "content": "<p>Yeah, this is really the part of ML that I don't like but I realize it is becoming more and more necessary. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239837,
          "author_name": "jy2tong",
          "author_url": "",
          "post_date": "03/16/2021 04:09:59",
          "content": "<p>Yea, MLOps is a thing…People are being paid top dollars to do it lol 😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239807,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "03/16/2021 03:17:02",
      "content": "<p>In terms of version control and folders etc.  <a href=\"https://www.kaggle.com/product-feedback/221448\" target=\"_blank\">Feature Launch Open Notebooks from Github</a> may be useful.  Have not tried this yet, but private repository and export to Github are coming soon. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1239813,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/16/2021 03:25:13",
          "content": "<p>Interesting. I know I have heard Philipp and Christof mention github actions as some way to push code and have it trigger new training. I'm sure there is some good workflow that could be created with all the various tools but I haven't spent the time to figure it out. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240230,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/16/2021 10:01:23",
          "content": "<p><a href=\"https://github.com/marketplace/actions/push-kaggle-dataset\" target=\"_blank\">https://github.com/marketplace/actions/push-kaggle-dataset</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239920,
      "author_name": "sanchitvj",
      "author_url": "",
      "post_date": "03/16/2021 05:55:21",
      "content": "<p>A week before when I started submissions I was just submitting notebooks with some modifications and taking look at LB if it is improved or not. After 2 days of doing the same, I got confused at a point, not able to decide which 96.8 gave me a boost in rank, which modification(like ensemble ratio, parameters, weights) worked best, and opening notebooks, again and again, was time wasted. So to solve this issue:</p>\n<ul>\n<li>I started naming notebooks properly with modifications I have done in this version.</li>\n<li>Giving version names properly like including date of submission and timing(approx) it ran for, rank before and after(if changed).</li>\n<li>Below there's a very good feature provided by Kaggle for writing descriptions which I noticed and started using from this competition only(might be old for other people). I used to write like what specific hyperparameter impacted LB. Also used to write if there was an improvement or not.<br>\n<img src=\"https://raw.githubusercontent.com/sanchitvj/Kaggle-Competitions/main/Ranzcr%20CLIP/kaggle%20sub_li.jpg\" alt=\"sub img\"></li>\n</ul>\n<p>I hope this is helpful.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1239708": "After every competition, I try to review a bit and consider what I could have done better in terms of process to have more success the next time. How would I have found the insight or technique to create a better solution or come to the same solution faster? Would be interested in hearing what other people feel they could have improved on process-wise. \n\nFor me this competition, I think one of my biggest misses was my poor logging setup. I did a poor job of keeping track of training and validation results across all of the metrics of interest. Sometimes accidentally overwriting results or leaving more information in notebooks than in the actual logs of the training procedure. For example, only showing the summary validation loss and auc rather than the loss and auc of individual columns. \n\nA few times I would iterate on notebooks accidentally in place rather than moving to a new notebook when doing an experiment and would no longer have the detailed information from training something or even worse have a good result, but no longer have the code that generated it. This is one of the big downfalls of notebooks IMO. I typically can operate fine with them just copying the notebook before I make changes, but have to be disciplined with it. \n\nI've seen some others do a few things to prevent this from happening:\n1. store models, logs, validation indexes and validation/test predictions in their own folders\n\n  - This keeps everything together so when you are selecting models to use for submission and analyzing what worked and what didn't you have a direct path back to the original source. For me, I started building a hierarchy where I had model/name_of_experiment/fold so I could look at all of my resnet18s and the variations I tried with them and then also compare specific folds to each other. I tried to keep each individual file specifically named as well so if I ever transferred to kaggle or other places it was clear what file was what. \n\n2. New folders instead of overwriting\n\n  - one of the issues I ran into was overwriting previous logs or outputs from notebooks. There are two different things I think could be useful here:\n     - at the end of every epoch or some other unit push a copy of the notebook to the same folder with all of the other results. This bundles the code and the outputs so even if you later rewrite code in the existing notebook you have a version to roll back to that even has the recordings of all of the cells. sort of a poor mans version control that happens automatically during training\n     -  every new rerun of some code, rather than overwriting anything just iterate a counter and make a new folder to put everything in. Storage is very cheap, save everything because it will save time down the line. \n\nOne of the things that I ran into that kind of clobbered some of my original logging plans was when I wanted to recontinue training I would lose some aspect of my logs. hindsight 20/20 I should have looked at my various procedures and planned them out more thoughtfully early on. Maybe I finally invest the time in setting up a proper solution like neptune.ai or sacred for the next one.",
    "1239710": "Neptune.ai is the best :D. Takes 5 minutes to learn and use.",
    "1239720": "I'll have to try it out. I looked at solutions a couple years ago but never committed to them because it just seemed like too much overhead just to do some logging and analysis",
    "1239722": "> **After** every competition. \n\nYet you post before its over...",
    "1239724": "Over for **me**. Given up on this one",
    "1239775": "Welcome to the world of \n\n# DevOps\n\n\ninteresting to see that kaggle competitions have become difficult that numerous experiments are done and proper logging become a must.\n\nin short, even kaggling needs devops\n\n---\n\nlast time, i was just developing messy code at the beginning. then i would clean up my code and experiments later if i have good solution.\n\n---\nnow i find that this approach is not efficient nor sufficient. I have to keep good habits of developing good code, good log and good experiments report #right at the start#. I don't have time to go back and re-organize.\n\n---\n\nyou can try this:\n- separate code from input data and output result.\n- first line in your code run it runs\n  - zip the whole project code (for repeatability) and save with timestamp\n - create a new output folder \n  - start a logfile\nsince i save my code every time it runs, i can always revert it back\n\nyou need to have good diskspace. for each kaggle competition, i typically endup with 200 to 500gb of rubbish",
    "1239776": "Yeah, this is really the part of ML that I don't like but I realize it is becoming more and more necessary.",
    "1239807": "In terms of version control and folders etc.  [Feature Launch Open Notebooks from Github](https://www.kaggle.com/product-feedback/221448) may be useful.  Have not tried this yet, but private repository and export to Github are coming soon.",
    "1239813": "Interesting. I know I have heard Philipp and Christof mention github actions as some way to push code and have it trigger new training. I'm sure there is some good workflow that could be created with all the various tools but I haven't spent the time to figure it out.",
    "1239837": "Yea, MLOps is a thing...People are being paid top dollars to do it lol 😂",
    "1239920": "A week before when I started submissions I was just submitting notebooks with some modifications and taking look at LB if it is improved or not. After 2 days of doing the same, I got confused at a point, not able to decide which 96.8 gave me a boost in rank, which modification(like ensemble ratio, parameters, weights) worked best, and opening notebooks, again and again, was time wasted. So to solve this issue:\n- I started naming notebooks properly with modifications I have done in this version.\n- Giving version names properly like including date of submission and timing(approx) it ran for, rank before and after(if changed).\n- Below there's a very good feature provided by Kaggle for writing descriptions which I noticed and started using from this competition only(might be old for other people). I used to write like what specific hyperparameter impacted LB. Also used to write if there was an improvement or not.\n![sub img](https://raw.githubusercontent.com/sanchitvj/Kaggle-Competitions/main/Ranzcr%20CLIP/kaggle%20sub_li.jpg)\n\nI hope this is helpful.",
    "1240209": "ryches had the same logic before. Now I realized that investing a few hours in learning how to use some infrastructure / tracking tool will save many more hours later :) I have also started using Neptune.ai this year and it works great.",
    "1240230": "https://github.com/marketplace/actions/push-kaggle-dataset",
    "1261354": "underwearfitting :D I took 30 minutes and still fidgeting... Do you have any nice notebook of yours to share for tracking/storing metrics/processes using neptune.ai? I even wanted to store augmentations params...",
    "1261369": "I used neptune with pytorchlightning. Here's an example: https://github.com/kagglesintracking/kaggle-Cassava-Leaf-Disease-Classification/blob/main/src/main.py"
  },
  "source": "meta"
}