{
  "id": 53881,
  "title": "(Manageable) Code Structure/Environment for Competitions Code",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53881",
  "author_name": "",
  "post_date": "2018-04-06T10:37:50.883051600Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi All, \nI want to know if you guys have a specific code structure that you guys follow? For me, it is getting increasingly unmanageable to plug in features or creating orthogonal models. Most of the kernels I see have a single python file with all the code dumped. </p>\n\n<ol>\n<li>Do you guys do EDA in a different Notebooks and then dump code to single notebook and then publish?</li>\n<li>Do you use a notebook importer like nbimporter or something else?</li>\n<li>Do you use notebooks in PyCharm or some other IDE?</li>\n<li>Has anyone tried an object oriented approach to these solutions, to try and manage better? on the lines of helper classes, or like wise?</li>\n<li>Do you have a CI pipeline or something with the EC2 Instance?</li>\n<li>Do you have tests or use a test framework as a safety net?</li>\n</ol>",
  "messages": [
    {
      "id": "309976",
      "postDate": "04/06/2018 10:37:50",
      "content": "<p>Hi All, \nI want to know if you guys have a specific code structure that you guys follow? For me, it is getting increasingly unmanageable to plug in features or creating orthogonal models. Most of the kernels I see have a single python file with all the code dumped. </p>\n\n<ol>\n<li>Do you guys do EDA in a different Notebooks and then dump code to single notebook and then publish?</li>\n<li>Do you use a notebook importer like nbimporter or something else?</li>\n<li>Do you use notebooks in PyCharm or some other IDE?</li>\n<li>Has anyone tried an object oriented approach to these solutions, to try and manage better? on the lines of helper classes, or like wise?</li>\n<li>Do you have a CI pipeline or something with the EC2 Instance?</li>\n<li>Do you have tests or use a test framework as a safety net?</li>\n</ol>",
      "rawMarkdown": "Hi All, \nI want to know if you guys have a specific code structure that you guys follow? For me, it is getting increasingly unmanageable to plug in features or creating orthogonal models. Most of the kernels I see have a single python file with all the code dumped. \n\n 1. Do you guys do EDA in a different Notebooks and then dump code to single notebook and then publish?\n 2. Do you use a notebook importer like nbimporter or something else?\n 3. Do you use notebooks in PyCharm or some other IDE?\n 4. Has anyone tried an object oriented approach to these solutions, to try and manage better? on the lines of helper classes, or like wise?\n 5. Do you have a CI pipeline or something with the EC2 Instance?\n 6. Do you have tests or use a test framework as a safety net?",
      "votes": null
    },
    {
      "id": "310088",
      "postDate": "04/06/2018 15:25:26",
      "content": "<p>This is also a problem for me, right now I basically separate codes into different structure, namely, train validation separate, feature generation, model tuning, test and prediction outcome, each in different folder  and import as class object.however, I often need to rename all the files (transformed data, models and corresponding outputs) within a manageable style which is quite headache.\nwaiting for top players sharing.</p>",
      "rawMarkdown": "This is also a problem for me, right now I basically separate codes into different structure, namely, train validation separate, feature generation, model tuning, test and prediction outcome, each in different folder  and import as class object.however, I often need to rename all the files (transformed data, models and corresponding outputs) within a manageable style which is quite headache.\nwaiting for top players sharing.",
      "votes": null
    },
    {
      "id": "310092",
      "postDate": "04/06/2018 15:34:22",
      "content": "<p>As far as one goes on refactoring stuff, I tired an approach of using IntelliJ as the IDE(this surprisingly supports python notebooks very well.), where refactoring variables with py files was easy, but used to get messy with imports overtime and changes not getting reflected in real time. Had to restart kernels.</p>",
      "rawMarkdown": "As far as one goes on refactoring stuff, I tired an approach of using IntelliJ as the IDE(this surprisingly supports python notebooks very well.), where refactoring variables with py files was easy, but used to get messy with imports overtime and changes not getting reflected in real time. Had to restart kernels.",
      "votes": null
    },
    {
      "id": "310241",
      "postDate": "04/06/2018 22:05:37",
      "content": "<p>You don't need anything too fancy, I separate my code like so:</p>\n\n<ol>\n<li>Separate all feature engineering code from your modeling code</li>\n<li>Save your dataframe as pkl/hdf5/ etc in between the FE/ preprocessing steps, sharding if your df gets too big.</li>\n<li>Load finished df when you want to start modeling</li>\n<li>Repeat 2-3 as many times as you want.</li>\n</ol>\n\n<p>I don't use notebooks/ides at all, but being able to load your finished df whenever you want will probably save you a lot of time.</p>",
      "rawMarkdown": "You don't need anything too fancy, I separate my code like so:\n\n 1. Separate all feature engineering code from your modeling code\n 2. Save your dataframe as pkl/hdf5/ etc in between the FE/ preprocessing steps, sharding if your df gets too big.\n 3. Load finished df when you want to start modeling\n 4. Repeat 2-3 as many times as you want.\n\nI don't use notebooks/ides at all, but being able to load your finished df whenever you want will probably save you a lot of time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 310088,
      "author_name": "daqixu",
      "author_url": "",
      "post_date": "04/06/2018 15:25:26",
      "content": "<p>This is also a problem for me, right now I basically separate codes into different structure, namely, train validation separate, feature generation, model tuning, test and prediction outcome, each in different folder  and import as class object.however, I often need to rename all the files (transformed data, models and corresponding outputs) within a manageable style which is quite headache.\nwaiting for top players sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 310092,
          "author_name": "ronakagrawal",
          "author_url": "",
          "post_date": "04/06/2018 15:34:22",
          "content": "<p>As far as one goes on refactoring stuff, I tired an approach of using IntelliJ as the IDE(this surprisingly supports python notebooks very well.), where refactoring variables with py files was easy, but used to get messy with imports overtime and changes not getting reflected in real time. Had to restart kernels.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310241,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "04/06/2018 22:05:37",
      "content": "<p>You don't need anything too fancy, I separate my code like so:</p>\n\n<ol>\n<li>Separate all feature engineering code from your modeling code</li>\n<li>Save your dataframe as pkl/hdf5/ etc in between the FE/ preprocessing steps, sharding if your df gets too big.</li>\n<li>Load finished df when you want to start modeling</li>\n<li>Repeat 2-3 as many times as you want.</li>\n</ol>\n\n<p>I don't use notebooks/ides at all, but being able to load your finished df whenever you want will probably save you a lot of time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "309976": "Hi All, \nI want to know if you guys have a specific code structure that you guys follow? For me, it is getting increasingly unmanageable to plug in features or creating orthogonal models. Most of the kernels I see have a single python file with all the code dumped. \n\n 1. Do you guys do EDA in a different Notebooks and then dump code to single notebook and then publish?\n 2. Do you use a notebook importer like nbimporter or something else?\n 3. Do you use notebooks in PyCharm or some other IDE?\n 4. Has anyone tried an object oriented approach to these solutions, to try and manage better? on the lines of helper classes, or like wise?\n 5. Do you have a CI pipeline or something with the EC2 Instance?\n 6. Do you have tests or use a test framework as a safety net?",
    "310088": "This is also a problem for me, right now I basically separate codes into different structure, namely, train validation separate, feature generation, model tuning, test and prediction outcome, each in different folder  and import as class object.however, I often need to rename all the files (transformed data, models and corresponding outputs) within a manageable style which is quite headache.\nwaiting for top players sharing.",
    "310092": "As far as one goes on refactoring stuff, I tired an approach of using IntelliJ as the IDE(this surprisingly supports python notebooks very well.), where refactoring variables with py files was easy, but used to get messy with imports overtime and changes not getting reflected in real time. Had to restart kernels.",
    "310241": "You don't need anything too fancy, I separate my code like so:\n\n 1. Separate all feature engineering code from your modeling code\n 2. Save your dataframe as pkl/hdf5/ etc in between the FE/ preprocessing steps, sharding if your df gets too big.\n 3. Load finished df when you want to start modeling\n 4. Repeat 2-3 as many times as you want.\n\nI don't use notebooks/ides at all, but being able to load your finished df whenever you want will probably save you a lot of time."
  },
  "source": "meta"
}