{
  "id": 55475,
  "title": "TalkingData LB0.9786",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55475",
  "author_name": "Md Asraful Kabir",
  "post_date": "2018-04-26T23:30:11.174000",
  "votes": 15,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Hi,\nI have used this kernel and gained 0.9786 in with 25mil rows. Have a look, If you are interested.  The feature second and minute was omitted during local PC running. Any suggestion for further improvement will be warmly welcomed. FYI, I do not have that much RAM to run whole dataset.\n <a href=\"https://www.kaggle.com/asraful70/notebook-version-of-talkingdata-lightgbm\">https://www.kaggle.com/asraful70/notebook-version-of-talkingdata-lightgbm</a></p>",
  "messages": [
    {
      "id": 319835,
      "postDate": "2018-04-26T23:30:11.173Z",
      "content": "<p>Hi,\nI have used this kernel and gained 0.9786 in with 25mil rows. Have a look, If you are interested.  The feature second and minute was omitted during local PC running. Any suggestion for further improvement will be warmly welcomed. FYI, I do not have that much RAM to run whole dataset.\n <a href=\"https://www.kaggle.com/asraful70/notebook-version-of-talkingdata-lightgbm\">https://www.kaggle.com/asraful70/notebook-version-of-talkingdata-lightgbm</a></p>",
      "rawMarkdown": "Hi,\nI have used this kernel and gained 0.9786 in with 25mil rows. Have a look, If you are interested.  The feature second and minute was omitted during local PC running. Any suggestion for further improvement will be warmly welcomed. FYI, I do not have that much RAM to run whole dataset.\n https://www.kaggle.com/asraful70/notebook-version-of-talkingdata-lightgbm",
      "votes": 15
    },
    {
      "id": 320206,
      "postDate": "2018-04-27T21:18:14.487Z",
      "content": "<p>Thanks. I got 0.9786 and 0.9791 by using 60M and 80M dataset, respectively. I wonder if we really need all the features in the kernel after I noticed a significant gain by having 33% more train samples.  Unfortunately, that's all I can fit on my laptop. Will it be possible to improve even more by having fewer features (eliminating those least significant features in importance), but training more samples?</p>",
      "rawMarkdown": "Thanks. I got 0.9786 and 0.9791 by using 60M and 80M dataset, respectively. I wonder if we really need all the features in the kernel after I noticed a significant gain by having 33% more train samples.  Unfortunately, that's all I can fit on my laptop. Will it be possible to improve even more by having fewer features (eliminating those least significant features in importance), but training more samples?",
      "votes": 1,
      "replies": [
        {
          "id": 320208,
          "postDate": "2018-04-27T21:26:53.400Z",
          "content": "<p>Yes possible. I have used 90 mil with same amount of feature and now at 0.9794. This is the best my laptop can afford.</p>",
          "rawMarkdown": "Yes possible. I have used 90 mil with same amount of feature and now at 0.9794. This is the best my laptop can afford."
        },
        {
          "id": 320236,
          "postDate": "2018-04-28T00:58:17.393Z",
          "content": "<p>No, you don't need all those features, I've tested almost all and some features do harm your model</p>",
          "rawMarkdown": "No, you don't need all those features, I've tested almost all and some features do harm your model",
          "votes": 2
        },
        {
          "id": 320366,
          "postDate": "2018-04-28T13:00:04.833Z",
          "content": "<p>How have you selected best features or  not harmful features?  </p>",
          "rawMarkdown": "How have you selected best features or  not harmful features?  "
        },
        {
          "id": 320707,
          "postDate": "2018-04-29T15:00:39.920Z",
          "content": "<p>I got 0.9796 with 80M data,it's so cool.Thanks</p>",
          "rawMarkdown": "I got 0.9796 with 80M data,it's so cool.Thanks"
        }
      ]
    },
    {
      "id": 323392,
      "postDate": "2018-05-05T02:50:28.317Z",
      "content": "<p>Thanks for the kernel. I got 0.9799 by using 120M data and test_supplement. I guess using full dataset may increase a little bit more.</p>",
      "rawMarkdown": "Thanks for the kernel. I got 0.9799 by using 120M data and test_supplement. I guess using full dataset may increase a little bit more.",
      "replies": [
        {
          "id": 323397,
          "postDate": "2018-05-05T03:07:17.753Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 323414,
          "postDate": "2018-05-05T04:47:38.700Z",
          "content": "<p>I followed the method in this discussion: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378</a>\nBasically get the predictions for test_supplement.csv first and then map predictions back to test.csv</p>",
          "rawMarkdown": "I followed the method in this discussion: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378\nBasically get the predictions for test_supplement.csv first and then map predictions back to test.csv"
        },
        {
          "id": 323416,
          "postDate": "2018-05-05T04:56:11.850Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 323657,
          "postDate": "2018-05-05T19:47:23.743Z",
          "content": "<p>I tried to do linear stacking of different kernels posted, usually it doesn't increase LB much though..</p>",
          "rawMarkdown": "I tried to do linear stacking of different kernels posted, usually it doesn't increase LB much though.."
        },
        {
          "id": 323790,
          "postDate": "2018-05-06T08:14:18.387Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 321125,
      "postDate": "2018-04-30T17:06:57.647Z",
      "content": "<p>Nice kernel, but you can improve the quality of it if you add all positive objects to the \"head\" of your training part with chunk reading.</p>",
      "rawMarkdown": "Nice kernel, but you can improve the quality of it if you add all positive objects to the \"head\" of your training part with chunk reading.",
      "replies": [
        {
          "id": 321186,
          "postDate": "2018-04-30T19:25:52.213Z",
          "content": "<p>Hi,\nI did not understand your suggestion.  Can you please explain a bit further with one small example please. </p>",
          "rawMarkdown": "Hi,\nI did not understand your suggestion.  Can you please explain a bit further with one small example please. "
        },
        {
          "id": 321188,
          "postDate": "2018-04-30T19:32:57.257Z",
          "content": "<p>for chunk in pd.read_csv('data/train.csv', chunksize=chunksize,  nrows=frm,dtype=dtypes,parse_dates=['click_time']):\n    filtered = (chunk[(np.where(chunk['is_attributed']==1, True, False))])\n    target= pd.concat([target, filtered], ignore_index=False, )</p>\n\n<p>you can add it (\"target\") to the head of your training data or go further and add all related objects of \"target\" to your data, it will improve your kernel if u cant load the whole dataset in memory.</p>",
          "rawMarkdown": "for chunk in pd.read_csv('data/train.csv', chunksize=chunksize,  nrows=frm,dtype=dtypes,parse_dates=['click_time']):\n    filtered = (chunk[(np.where(chunk['is_attributed']==1, True, False))])\n    target= pd.concat([target, filtered], ignore_index=False, )\n\nyou can add it (\"target\") to the head of your training data or go further and add all related objects of \"target\" to your data, it will improve your kernel if u cant load the whole dataset in memory.",
          "votes": 1
        }
      ]
    },
    {
      "id": 320279,
      "postDate": "2018-04-28T06:10:27.657Z",
      "content": "<p>How did you train the 40million rows data, Kaggle kernal or your local PC(CPU) or local PC(GPU) or AWS clound and so on? How much time will cost?</p>",
      "rawMarkdown": "How did you train the 40million rows data, Kaggle kernal or your local PC(CPU) or local PC(GPU) or AWS clound and so on? How much time will cost?",
      "replies": [
        {
          "id": 320364,
          "postDate": "2018-04-28T12:58:18.373Z",
          "content": "<p>In my local PC core i7, 32 GB RAM. It took  approximately 2 hrs </p>",
          "rawMarkdown": "In my local PC core i7, 32 GB RAM. It took  approximately 2 hrs ",
          "votes": 3
        },
        {
          "id": 320424,
          "postDate": "2018-04-28T16:29:48.207Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 319945,
      "postDate": "2018-04-27T06:39:27.500Z",
      "content": "<p>Have you try all train dataset？</p>",
      "rawMarkdown": "Have you try all train dataset？",
      "replies": [
        {
          "id": 319968,
          "postDate": "2018-04-27T07:50:28.877Z",
          "content": "<p>No.  not really. 40mil rows for train and 10% for validate </p>",
          "rawMarkdown": "No.  not really. 40mil rows for train and 10% for validate "
        },
        {
          "id": 320110,
          "postDate": "2018-04-27T14:37:02.120Z",
          "content": "<p>More train data is useful, you should try it.</p>",
          "rawMarkdown": "More train data is useful, you should try it."
        },
        {
          "id": 320115,
          "postDate": "2018-04-27T14:44:40.437Z",
          "content": "<p>Yes. i know. but there is always a RAM limitation</p>",
          "rawMarkdown": "Yes. i know. but there is always a RAM limitation"
        },
        {
          "id": 320119,
          "postDate": "2018-04-27T14:48:56.973Z",
          "content": "<p>What a pity!\nI think you can reach 0.98X with all train dataset, if you now can get lb0.9794 now.</p>",
          "rawMarkdown": "What a pity!\nI think you can reach 0.98X with all train dataset, if you now can get lb0.9794 now."
        },
        {
          "id": 321065,
          "postDate": "2018-04-30T14:15:12.773Z",
          "content": "<p>i get 0.9798 with all data</p>",
          "rawMarkdown": "i get 0.9798 with all data"
        },
        {
          "id": 321069,
          "postDate": "2018-04-30T14:21:36.177Z",
          "content": "<p>Do you only use the last 2500000 rows as the val data?</p>",
          "rawMarkdown": "Do you only use the last 2500000 rows as the val data?"
        },
        {
          "id": 321079,
          "postDate": "2018-04-30T14:39:11.450Z",
          "content": "<p>@Liu congratulation. Its good to know that by using my script someone get such a high accuracy. I am helpless as my laptop has not that much capacity. How much validation data you have used?</p>",
          "rawMarkdown": "@Liu congratulation. Its good to know that by using my script someone get such a high accuracy. I am helpless as my laptop has not that much capacity. How much validation data you have used?"
        },
        {
          "id": 321306,
          "postDate": "2018-05-01T01:46:24.810Z",
          "content": "<p>@Kabir I use last 2.5M as val data, run 600 rounds, I try to use 10M in another script,but LB score decrease.</p>",
          "rawMarkdown": "@Kabir I use last 2.5M as val data, run 600 rounds, I try to use 10M in another script,but LB score decrease."
        },
        {
          "id": 321308,
          "postDate": "2018-05-01T01:50:06.383Z",
          "content": "<p>@Kabir I  think you can try use GCP to run your script,It's free for one year or you can get 300 credits,and all you need is just a credit card to verify your identity. I run script in GCP by using 300 credits too,<strong>You must try!!!!</strong> </p>",
          "rawMarkdown": "@Kabir I  think you can try use GCP to run your script,It's free for one year or you can get 300 credits,and all you need is just a credit card to verify your identity. I run script in GCP by using 300 credits too,**You must try!!!!** "
        },
        {
          "id": 321550,
          "postDate": "2018-05-01T14:32:55.747Z",
          "content": "<p>@Liu I would suggest to use atleast 10% for validation for 1000 rounds. Lets see.</p>",
          "rawMarkdown": "@Liu I would suggest to use atleast 10% for validation for 1000 rounds. Lets see."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 320206,
      "author_name": "Tapioca",
      "author_url": "",
      "post_date": "2018-04-27T21:18:14.487000",
      "content": "<p>Thanks. I got 0.9786 and 0.9791 by using 60M and 80M dataset, respectively. I wonder if we really need all the features in the kernel after I noticed a significant gain by having 33% more train samples.  Unfortunately, that's all I can fit on my laptop. Will it be possible to improve even more by having fewer features (eliminating those least significant features in importance), but training more samples?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 320208,
          "author_name": "Md Asraful Kabir",
          "author_url": "",
          "post_date": "2018-04-27T21:26:53.400000",
          "content": "<p>Yes possible. I have used 90 mil with same amount of feature and now at 0.9794. This is the best my laptop can afford.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320236,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-28T00:58:17.393000",
          "content": "<p>No, you don't need all those features, I've tested almost all and some features do harm your model</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 320366,
          "author_name": "Md Asraful Kabir",
          "author_url": "",
          "post_date": "2018-04-28T13:00:04.833000",
          "content": "<p>How have you selected best features or  not harmful features?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320707,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-29T15:00:39.920000",
          "content": "<p>I got 0.9796 with 80M data,it's so cool.Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 323392,
      "author_name": "FrLi",
      "author_url": "",
      "post_date": "2018-05-05T02:50:28.317000",
      "content": "<p>Thanks for the kernel. I got 0.9799 by using 120M data and test_supplement. I guess using full dataset may increase a little bit more.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 323397,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-05T03:07:17.753000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323414,
          "author_name": "FrLi",
          "author_url": "",
          "post_date": "2018-05-05T04:47:38.700000",
          "content": "<p>I followed the method in this discussion: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378</a>\nBasically get the predictions for test_supplement.csv first and then map predictions back to test.csv</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323416,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-05T04:56:11.850000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323657,
          "author_name": "FrLi",
          "author_url": "",
          "post_date": "2018-05-05T19:47:23.743000",
          "content": "<p>I tried to do linear stacking of different kernels posted, usually it doesn't increase LB much though..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323790,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-06T08:14:18.387000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 321125,
      "author_name": "Izmaylov Konstantin",
      "author_url": "",
      "post_date": "2018-04-30T17:06:57.647000",
      "content": "<p>Nice kernel, but you can improve the quality of it if you add all positive objects to the \"head\" of your training part with chunk reading.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 321186,
          "author_name": "Md Asraful Kabir",
          "author_url": "",
          "post_date": "2018-04-30T19:25:52.213000",
          "content": "<p>Hi,\nI did not understand your suggestion.  Can you please explain a bit further with one small example please. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321188,
          "author_name": "Izmaylov Konstantin",
          "author_url": "",
          "post_date": "2018-04-30T19:32:57.257000",
          "content": "<p>for chunk in pd.read_csv('data/train.csv', chunksize=chunksize,  nrows=frm,dtype=dtypes,parse_dates=['click_time']):\n    filtered = (chunk[(np.where(chunk['is_attributed']==1, True, False))])\n    target= pd.concat([target, filtered], ignore_index=False, )</p>\n\n<p>you can add it (\"target\") to the head of your training data or go further and add all related objects of \"target\" to your data, it will improve your kernel if u cant load the whole dataset in memory.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 320279,
      "author_name": "StudyExchange",
      "author_url": "",
      "post_date": "2018-04-28T06:10:27.657000",
      "content": "<p>How did you train the 40million rows data, Kaggle kernal or your local PC(CPU) or local PC(GPU) or AWS clound and so on? How much time will cost?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 320364,
          "author_name": "Md Asraful Kabir",
          "author_url": "",
          "post_date": "2018-04-28T12:58:18.373000",
          "content": "<p>In my local PC core i7, 32 GB RAM. It took  approximately 2 hrs </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 320424,
          "author_name": "StudyExchange",
          "author_url": "",
          "post_date": "2018-04-28T16:29:48.207000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319945,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "2018-04-27T06:39:27.500000",
      "content": "<p>Have you try all train dataset？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 319968,
          "author_name": "Md Asraful Kabir",
          "author_url": "",
          "post_date": "2018-04-27T07:50:28.877000",
          "content": "<p>No.  not really. 40mil rows for train and 10% for validate </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320110,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "2018-04-27T14:37:02.120000",
          "content": "<p>More train data is useful, you should try it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320115,
          "author_name": "Md Asraful Kabir",
          "author_url": "",
          "post_date": "2018-04-27T14:44:40.437000",
          "content": "<p>Yes. i know. but there is always a RAM limitation</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320119,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "2018-04-27T14:48:56.973000",
          "content": "<p>What a pity!\nI think you can reach 0.98X with all train dataset, if you now can get lb0.9794 now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321065,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-30T14:15:12.773000",
          "content": "<p>i get 0.9798 with all data</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321069,
          "author_name": "yyll008",
          "author_url": "",
          "post_date": "2018-04-30T14:21:36.177000",
          "content": "<p>Do you only use the last 2500000 rows as the val data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321079,
          "author_name": "Md Asraful Kabir",
          "author_url": "",
          "post_date": "2018-04-30T14:39:11.450000",
          "content": "<p>@Liu congratulation. Its good to know that by using my script someone get such a high accuracy. I am helpless as my laptop has not that much capacity. How much validation data you have used?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321306,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-05-01T01:46:24.810000",
          "content": "<p>@Kabir I use last 2.5M as val data, run 600 rounds, I try to use 10M in another script,but LB score decrease.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321308,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-05-01T01:50:06.383000",
          "content": "<p>@Kabir I  think you can try use GCP to run your script,It's free for one year or you can get 300 credits,and all you need is just a credit card to verify your identity. I run script in GCP by using 300 credits too,<strong>You must try!!!!</strong> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321550,
          "author_name": "Md Asraful Kabir",
          "author_url": "",
          "post_date": "2018-05-01T14:32:55.747000",
          "content": "<p>@Liu I would suggest to use atleast 10% for validation for 1000 rounds. Lets see.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "319835": "Hi,\nI have used this kernel and gained 0.9786 in with 25mil rows. Have a look, If you are interested.  The feature second and minute was omitted during local PC running. Any suggestion for further improvement will be warmly welcomed. FYI, I do not have that much RAM to run whole dataset.\n https://www.kaggle.com/asraful70/notebook-version-of-talkingdata-lightgbm",
    "320206": "Thanks. I got 0.9786 and 0.9791 by using 60M and 80M dataset, respectively. I wonder if we really need all the features in the kernel after I noticed a significant gain by having 33% more train samples.  Unfortunately, that's all I can fit on my laptop. Will it be possible to improve even more by having fewer features (eliminating those least significant features in importance), but training more samples?",
    "323392": "Thanks for the kernel. I got 0.9799 by using 120M data and test_supplement. I guess using full dataset may increase a little bit more.",
    "321125": "Nice kernel, but you can improve the quality of it if you add all positive objects to the \"head\" of your training part with chunk reading.",
    "320279": "How did you train the 40million rows data, Kaggle kernal or your local PC(CPU) or local PC(GPU) or AWS clound and so on? How much time will cost?",
    "319945": "Have you try all train dataset？"
  }
}