{
  "id": 71352,
  "title": "galactic and extra galactic",
  "url": "/competitions/PLAsTiCC-2018/discussion/71352",
  "author_name": "",
  "post_date": "2018-11-12T23:42:41.051114300Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am trying to divide test set on galactic and extra-galactic. Is it worth doing it (the classifier learns the difference anyway, but it could be better to divide for features engineering)?</p>\n\n<p>I tried on cloud with 104 G RAM and I still got a memory error mistake... I publish my code here, maybe someone decides to reuse it and help me to sort the memory issue</p>\n\n<p>```\n\"\"\"\nCreated on Sun Nov 11 01:51:00 2018\ndivide test data on galactic and extra galactic\n\"\"\"\nimport numpy as np\nimport pandas as pd\nimport os\nimport gc # garbage collector\nimport time\nimport tqdm\nimport logging</p>\n\n<h1>pandas display options</h1>\n\n<p>pd.set_option('display.max_columns', None)  </p>\n\n<p>DATA_DIR = '../input'</p>\n\n<p>meta_test = pd.read_csv(DATA_DIR +'test_set_metadata.csv')</p>\n\n<p>test = pd.read_csv(DATA_DIR + 'test_set.csv')</p>\n\n<p>test_full = test.merge(meta_test, how='left', on='object_id')</p>\n\n<p>del test\ngc.collect()</p>\n\n<p>print('test_full', test_full.head())</p>\n\n<h1>galactic</h1>\n\n<p>gal_test = test_full[test_full.distmod.isnull()]\nobj = gal_test['object_id'].unique()\nprint('obj gal', len(obj))</p>\n\n<p>print('gal_test', gal_test.head())\ngal_test.to_csv('gal_test.csv', index=False)</p>\n\n<p>del gal_test\ngc.collect()\n```</p>",
  "messages": [
    {
      "id": "420022",
      "postDate": "11/12/2018 23:42:41",
      "content": "<p>I am trying to divide test set on galactic and extra-galactic. Is it worth doing it (the classifier learns the difference anyway, but it could be better to divide for features engineering)?</p>\n\n<p>I tried on cloud with 104 G RAM and I still got a memory error mistake... I publish my code here, maybe someone decides to reuse it and help me to sort the memory issue</p>\n\n<p>```\n\"\"\"\nCreated on Sun Nov 11 01:51:00 2018\ndivide test data on galactic and extra galactic\n\"\"\"\nimport numpy as np\nimport pandas as pd\nimport os\nimport gc # garbage collector\nimport time\nimport tqdm\nimport logging</p>\n\n<h1>pandas display options</h1>\n\n<p>pd.set_option('display.max_columns', None)  </p>\n\n<p>DATA_DIR = '../input'</p>\n\n<p>meta_test = pd.read_csv(DATA_DIR +'test_set_metadata.csv')</p>\n\n<p>test = pd.read_csv(DATA_DIR + 'test_set.csv')</p>\n\n<p>test_full = test.merge(meta_test, how='left', on='object_id')</p>\n\n<p>del test\ngc.collect()</p>\n\n<p>print('test_full', test_full.head())</p>\n\n<h1>galactic</h1>\n\n<p>gal_test = test_full[test_full.distmod.isnull()]\nobj = gal_test['object_id'].unique()\nprint('obj gal', len(obj))</p>\n\n<p>print('gal_test', gal_test.head())\ngal_test.to_csv('gal_test.csv', index=False)</p>\n\n<p>del gal_test\ngc.collect()\n```</p>",
      "rawMarkdown": "I am trying to divide test set on galactic and extra-galactic. Is it worth doing it (the classifier learns the difference anyway, but it could be better to divide for features engineering)?\n \nI tried on cloud with 104 G RAM and I still got a memory error mistake... I publish my code here, maybe someone decides to reuse it and help me to sort the memory issue\n\n```\n\"\"\"\nCreated on Sun Nov 11 01:51:00 2018\ndivide test data on galactic and extra galactic\n\"\"\"\nimport numpy as np\nimport pandas as pd\nimport os\nimport gc # garbage collector\nimport time\nimport tqdm\nimport logging\n\n# pandas display options\npd.set_option('display.max_columns', None)  \n\nDATA_DIR = '../input'\n\nmeta_test = pd.read_csv(DATA_DIR +'test_set_metadata.csv')\n\ntest = pd.read_csv(DATA_DIR + 'test_set.csv')\n\ntest_full = test.merge(meta_test, how='left', on='object_id')\n\ndel test\ngc.collect()\n\nprint('test_full', test_full.head())\n\n# galactic\ngal_test = test_full[test_full.distmod.isnull()]\nobj = gal_test['object_id'].unique()\nprint('obj gal', len(obj))\n\nprint('gal_test', gal_test.head())\ngal_test.to_csv('gal_test.csv', index=False)\n\ndel gal_test\ngc.collect()\n```",
      "votes": null
    },
    {
      "id": "420073",
      "postDate": "11/13/2018 02:30:13",
      "content": "<p>I never load the test data in full, I always work in chunks.  Your issue is that you get almost 3 copies of test data at a point.  Moreover, in order to get the object_ids for galactic and extra galactic you don't need test, you only need meta_test.</p>",
      "rawMarkdown": "I never load the test data in full, I always work in chunks.  Your issue is that you get almost 3 copies of test data at a point.  Moreover, in order to get the object_ids for galactic and extra galactic you don't need test, you only need meta_test.",
      "votes": null
    },
    {
      "id": "420252",
      "postDate": "11/13/2018 11:00:40",
      "content": "<p>I did try in chunks, but my program got a bug... therefore tried to increase memory. It is really worth it? (we learn the diff from photoz feature and separate those classes on confusion matrix pretty well anyway) </p>",
      "rawMarkdown": "I did try in chunks, but my program got a bug... therefore tried to increase memory. It is really worth it? (we learn the diff from photoz feature and separate those classes on confusion matrix pretty well anyway)",
      "votes": null
    },
    {
      "id": "420253",
      "postDate": "11/13/2018 11:05:52",
      "content": "<p>I run in about 6GB.</p>",
      "rawMarkdown": "I run in about 6GB.",
      "votes": null
    },
    {
      "id": "420265",
      "postDate": "11/13/2018 11:24:55",
      "content": "<p>I have 7.5 G, will try to team up and fix my programming issues. What I meant:  is it really worth it to separate those models for gal and extragal</p>",
      "rawMarkdown": "I have 7.5 G, will try to team up and fix my programming issues. What I meant:  is it really worth it to separate those models for gal and extragal",
      "votes": null
    },
    {
      "id": "420270",
      "postDate": "11/13/2018 11:28:54",
      "content": "<p>I have a single model for both.  I once tried to have one for each, but score was worse.  I didn't checked recently.  But you should definitely try as what is true for one way to approach the model may not be true with another way.</p>",
      "rawMarkdown": "I have a single model for both.  I once tried to have one for each, but score was worse.  I didn't checked recently.  But you should definitely try as what is true for one way to approach the model may not be true with another way.",
      "votes": null
    },
    {
      "id": "420276",
      "postDate": "11/13/2018 11:35:00",
      "content": "<p>Thank you. <code>what is true for one way to approach the model may not be true with another way</code> -- speaking of which another questions wonders me: when I design nice features for lgbm classifier (it's faster), but want to use them for nn classifier (it works better for me), can I be confident that set of features good for one classifier will still be good for another? (I am not very experienced in ML, it looks logical that if features capture light curve behavior well, it should not matter much what classifier to use)</p>",
      "rawMarkdown": "Thank you. ```what is true for one way to approach the model may not be true with another way ``` -- speaking of which another questions wonders me: when I design nice features for lgbm classifier (it's faster), but want to use them for nn classifier (it works better for me), can I be confident that set of features good for one classifier will still be good for another? (I am not very experienced in ML, it looks logical that if features capture light curve behavior well, it should not matter much what classifier to use)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 420073,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/13/2018 02:30:13",
      "content": "<p>I never load the test data in full, I always work in chunks.  Your issue is that you get almost 3 copies of test data at a point.  Moreover, in order to get the object_ids for galactic and extra galactic you don't need test, you only need meta_test.</p>",
      "votes": null,
      "replies": [
        {
          "id": 420252,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/13/2018 11:00:40",
          "content": "<p>I did try in chunks, but my program got a bug... therefore tried to increase memory. It is really worth it? (we learn the diff from photoz feature and separate those classes on confusion matrix pretty well anyway) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420253,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/13/2018 11:05:52",
          "content": "<p>I run in about 6GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420265,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/13/2018 11:24:55",
          "content": "<p>I have 7.5 G, will try to team up and fix my programming issues. What I meant:  is it really worth it to separate those models for gal and extragal</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420270,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/13/2018 11:28:54",
          "content": "<p>I have a single model for both.  I once tried to have one for each, but score was worse.  I didn't checked recently.  But you should definitely try as what is true for one way to approach the model may not be true with another way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420276,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/13/2018 11:35:00",
          "content": "<p>Thank you. <code>what is true for one way to approach the model may not be true with another way</code> -- speaking of which another questions wonders me: when I design nice features for lgbm classifier (it's faster), but want to use them for nn classifier (it works better for me), can I be confident that set of features good for one classifier will still be good for another? (I am not very experienced in ML, it looks logical that if features capture light curve behavior well, it should not matter much what classifier to use)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "420022": "I am trying to divide test set on galactic and extra-galactic. Is it worth doing it (the classifier learns the difference anyway, but it could be better to divide for features engineering)?\n \nI tried on cloud with 104 G RAM and I still got a memory error mistake... I publish my code here, maybe someone decides to reuse it and help me to sort the memory issue\n\n```\n\"\"\"\nCreated on Sun Nov 11 01:51:00 2018\ndivide test data on galactic and extra galactic\n\"\"\"\nimport numpy as np\nimport pandas as pd\nimport os\nimport gc # garbage collector\nimport time\nimport tqdm\nimport logging\n\n# pandas display options\npd.set_option('display.max_columns', None)  \n\nDATA_DIR = '../input'\n\nmeta_test = pd.read_csv(DATA_DIR +'test_set_metadata.csv')\n\ntest = pd.read_csv(DATA_DIR + 'test_set.csv')\n\ntest_full = test.merge(meta_test, how='left', on='object_id')\n\ndel test\ngc.collect()\n\nprint('test_full', test_full.head())\n\n# galactic\ngal_test = test_full[test_full.distmod.isnull()]\nobj = gal_test['object_id'].unique()\nprint('obj gal', len(obj))\n\nprint('gal_test', gal_test.head())\ngal_test.to_csv('gal_test.csv', index=False)\n\ndel gal_test\ngc.collect()\n```",
    "420073": "I never load the test data in full, I always work in chunks.  Your issue is that you get almost 3 copies of test data at a point.  Moreover, in order to get the object_ids for galactic and extra galactic you don't need test, you only need meta_test.",
    "420252": "I did try in chunks, but my program got a bug... therefore tried to increase memory. It is really worth it? (we learn the diff from photoz feature and separate those classes on confusion matrix pretty well anyway)",
    "420253": "I run in about 6GB.",
    "420265": "I have 7.5 G, will try to team up and fix my programming issues. What I meant:  is it really worth it to separate those models for gal and extragal",
    "420270": "I have a single model for both.  I once tried to have one for each, but score was worse.  I didn't checked recently.  But you should definitely try as what is true for one way to approach the model may not be true with another way.",
    "420276": "Thank you. ```what is true for one way to approach the model may not be true with another way ``` -- speaking of which another questions wonders me: when I design nice features for lgbm classifier (it's faster), but want to use them for nn classifier (it works better for me), can I be confident that set of features good for one classifier will still be good for another? (I am not very experienced in ML, it looks logical that if features capture light curve behavior well, it should not matter much what classifier to use)"
  },
  "source": "meta"
}