{
  "id": 56571,
  "title": "22nd Place Overview",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/james-trotman-22nd-place-overview",
  "author_name": "",
  "post_date": "2020-02-20T20:14:02.813Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to Kaggle &amp; TalkingData for another entertaining diversion :) Congrats to all winners and thanks to all who shared Kernels, tips and solution write-ups, they’re great to read.</p>\n\n<p>This is a really interesting data set and as I related in my <a href=\"https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots/code#310213\">story here</a>, click fraud is a really insidious problem that drives people to do really <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54765\">odd things</a> (thanks <a href=\"/yifanxie\">@yifanxie</a> - amazing to actually see it!)</p>\n\n<h2>Overview</h2>\n\n<p>I used broadly similar features to everyone else, next click times and count features, and stuck to the route recommended by early leaders – using lightgbm and training on the full training set – it surprised me that my 32Gb machine could. After reading the great <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325\">shared tips</a>, switching to creating a NumPy array upfront then populating that by loading feature columns was very efficient, with 33 features it barely even needed any swap space.</p>\n\n<p>I validated on the 2nd half of the last day of training for simplicity, just one validation set, then I’d save predictions in a dataframe and do groupby(‘hour’) and check AUC that way. I did pay more attention to the hours that matched the test set hours, but I thought other hours might lead to useful insights (but nothing noteworthy to report).</p>\n\n<p>My general process for one iteration was quite lightweight:</p>\n\n<ol>\n<li>an hour of feature engineering – save columns individually</li>\n<li>a 1.5-2 hour validation run</li>\n<li>a 3 hour full model build</li>\n</ol>\n\n<p>After setting this up, the only real manual effort is step 1, the validation run just determines # of trees (which was quite stable over feature sets, but sometimes repeated if the results were poor). The full model build is a one click process that builds a model, makes test set predictions with varying numbers of trees (to test assumptions about increased data set size &amp; scaling # of trees), then uploads the submissions via the highly recommended <a href=\"https://github.com/Kaggle/kaggle-api\">Kaggle API</a>.</p>\n\n<p>I managed 13 iterations, never more than one in a day, so it was a bit of a slow waiting game, and a long way from optimal.  Getting results in fewer iterations is down to some intuition and luck in selecting what works. There are a lot of features I tried that didn’t work, so my remaining features are mostly fairly baseline/simple things… (<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">Matrix factorization of IP vs app log click counts</a> is the one thing I really wish I’d tried, it will summarize an IP’s affinity for apps and interact very well with the app and IP count features… Well done to <a href=\"/cpmpml\">@cpmpml</a> &amp; other top teams :)</p>\n\n<h2>Entropy features</h2>\n\n<p>Whereas most people used nunique, I thought that might be noisy. With a click count histogram like e.g. counting os’s for an app = [12345, 5, 1, 1] it’s technically 4 unique os’s, but very low entropy (it’s more like one os).</p>\n\n<p>The original feature encodings made this easier: they were all 0..n, so it was simple to do it all in NumPy using <code>np.add.at()</code> to fill a 2D array with counts, then summarize like this:</p>\n\n<pre><code>from scipy.stats import entropy\n\n# e.g. a=’ip’ ; b=’app’\ncc = np.zeros((int(df[a].max()+1), int(df[b].max()+1)), dtype=np.int32)\nnp.add.at(cc, (df[a].values, df[b].values), 1)\n\n(cc&amp;gt;0).sum(0) # this is nunique for a\n(cc&amp;gt;0).sum(1) # this is nunique for b\nmap(entropy, cc)    # entropy of a over b\nmap(entropy, cc.T)  # entropy of b over a\n</code></pre>\n\n<p>All two way interactions were possible with 32Gb of RAM, and took only minutes to compute this way. I also combined device &amp; os into one new categorical (only 6561 unique values), and app &amp; channel (1518 unique values), which enabled some 3-way interactions like entropy of IP over (device, os).</p>\n\n<h2>Sessions</h2>\n\n<p>Four new features:</p>\n\n<ul>\n<li>sort by ip, dev, os</li>\n<li>define max gap e.g. 60 seconds</li>\n<li>every time there is a change in ip/device/os, or a gap in the clicks, assign categorical session ID</li>\n<li>for sessions, summarize:\n<ul><li>duration</li>\n<li>count</li>\n<li>click rate (duration/count)</li>\n<li>app entropy</li></ul></li>\n</ul>\n\n<p>I intended to add entropy over gaps, e.g. lots of similar gaps of 5-6 secs = low entropy = bots. Thought maybe click rate was enough.</p>\n\n<h2>Naive Count Features</h2>\n\n<p>Normally, count features group on one or more columns and count the actual occurrences. Inspired by Naive Bayes classifiers: another view on the data is to assume the fields are independent, and do:</p>\n\n<ul>\n<li>univariate count</li>\n<li>divide counts by the number of rows to get a probability of seeing the value</li>\n<li>log() the probability</li>\n<li>sum different combinations of these columns similarly to normal feature interactions (i.e. -, +, /, *).</li>\n</ul>\n\n<p>Mathematically, this is the same as a geometric mean probability of seeing the record, under the assumption the columns are independent. Again, extremely fast to compute using NumPy alone.</p>\n\n<p>The combinations &amp; loading code I used:</p>\n\n<pre><code>def loadraw(name, dtype):\n    return np.fromfile(name, dtype=dtype)\n\ndef log_p_feat(t, col):\n    # t is ‘train’ or ‘test’\n    return loadraw('../feats/%s_log_p_%s'%(t,col), np.float32)\n\nx[:,a] = log_p_feat(t, 'ip') + log_p_feat(t, 'dev') + log_p_feat(t, 'os')\nx[:,a+1] = x[:,a] + log_p_feat(t, 'app')\nx[:,a+2] = x[:,a] + log_p_feat(t, 'chan')\nx[:,a+3] = x[:,a] + log_p_feat(t, 'app') + log_p_feat(t, 'chan')\n</code></pre>\n\n<p>Adding these features in with normal 2D/3D counts might pick up on interesting things. With these ‘naive’ counts: rare values overlap - a low (log) probability row might be rare IP with common OS <strong><em>or</em></strong> a rare OS with common IP. (Although IP was much higher cardinality, so probably dominates the log sum. A weighted mix might be better.)</p>\n\n<p>This can produce a very large number of unique values, so lightgbm value binning might come to the rescue here and avoid over-fitting. I’m not 100% convinced this was a good idea… but it did lead to a good single model (0.9811 public, 0.9825 private) where these naive count features had high feature importances.</p>\n\n<h2>Bit Shifting</h2>\n\n<p>Small tip: by looking at the maximum values for each column, you can work out how many bits are required to store each, and happily, the five categoricals all fit in a single 64 bit integer. You can pack them like this:</p>\n\n<pre><code>def do_pack(df):\n    a = 0\n    a += df.app.values.astype(np.uint64)\n    a += df.os.values.astype(np.uint64)&amp;lt;&amp;lt;10\n    a += df.channel.values.astype(np.uint64)&amp;lt;&amp;lt;20\n    a += df.ip.values.astype(np.uint64)&amp;lt;&amp;lt;30\n    a += df.device.values.astype(np.uint64)&amp;lt;&amp;lt;50\n    return a\n</code></pre>\n\n<p>They can be unpacked (these work on values or NumPy arrays):</p>\n\n<pre><code>def fapp(v): return v&amp;amp;0x3ff\ndef fos(v): return (v&amp;gt;&amp;gt;10)&amp;amp;0x3ff\ndef fchan(v): return (v&amp;gt;&amp;gt;20)&amp;amp;0x3ff\ndef fip(v): return (v&amp;gt;&amp;gt;30)&amp;amp;0xfffff\ndef fdev(v): return (v&amp;gt;&amp;gt;50)&amp;amp;0xfffff\n</code></pre>\n\n<p>This is vectorized, so runs extremely quickly, and the result is like a hash, but without collisions. You can use the result to do groupby() operations - more efficiently in both time &amp; RAM. You can also ignore fields by setting all their bits to one, e.g. to ignore device <code>a |= 0xfffffl&amp;lt;&amp;lt;50</code>. (Because 0 was a valid value for each field.) Using <code>np.save()</code> (or <code>a.tofile</code> and <code>np.fromfile</code>) on the 64 bit array you could save the entire 4 days of train &amp; full test into 1.9Gb on disk. Times could be similarly packed into 32 bits, about 900Mb… With these copies of the data I actually managed to do some feature exploration in idle moments (in Java) on an old 8Gb MacBook!</p>\n\n<h2>Categoricals</h2>\n\n<p>Grouping on all five combined fields, there were many different series, some with 10k-20k appearances, but a long tail of shorter series. Instead of separate count &amp; cumcount features I tried a categorical feature to mark where the row is in the rarer/shorter series. Using max_bin of 255, there are 8 bits to play with: 4 bits to mark the length of the series (1..15) and 4 bits to mark where in the series the record appeared (0..14).</p>\n\n<p>Where v is all five fields combined as above, use groupby(v).size() to get df.series_size and groupby(v).cumcount() to get df.series_cumcount, then:</p>\n\n<pre><code>idx = df.series_size.values&amp;lt;=15\ncat = np.where(idx, df.series_size.values&amp;lt;&amp;lt;4, 0)\ncat |= np.where(idx, df.series_cumcount.values, 0)\n</code></pre>\n\n<p>So this is one feature that encodes the length of the ‘series’ of [ip,dev,os,app,chan] <strong>and</strong> the position within the series. Tree splits can then address subsets like:</p>\n\n<pre><code>if cat == 0x32\nif 3_element_series and this_record_is_last_in_series\n\nif cat == 0x20\nif 2_element_series and this_record_is_first_in_series\n</code></pre>\n\n<p>To make the latest versions of lightgbm do these kind of splits you need to set the <code>max_cat_to_onehot</code> parameter correctly, in this case to 256 or more – but beware this applies to all categorical features, so for example channel may be used for one hot style splits too – it depends on how many values are seen in each column at dataset construction time (lightgbm samples data to construct the bins). All model parameters/regularizations are trade-offs of some kind.</p>\n\n<p>Limited time and iterations mean I can’t be sure this didn’t overfit, particularly for series that span train &amp; test times, and things like “length 14 series, element 4” which is not likely to mean anything (the hope is, this could uncover some really quirky leak!) Perhaps addressing only shorter series would be better, or chunking the series by time. I relied on the model to tell me which were more important - a quick inspection of models shows low values like 0x20, 0x21 were used the most.</p>\n\n<h2>Leaderboard</h2>\n\n<p>From <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56524\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56524</a></p>\n\n<p><img src=\"https://kaggle2.blob.core.windows.net/forum-message-attachments/326995/9398/jt_public_lb.png\" alt=\"JT public LB\"></p>\n\n<p>For me, validation scores stayed very consistent, but leaderboard scores were mysteriously lower. It was only eventually removing some features that fixed this: all of the count features that used the corrected Chinese timezone day. (These were: <code>ip_day_count</code>, <code>app_channel_day_count</code>, <code>device_os_day_count</code>.) I’m still not sure what caused that, and it’s puzzling that other’s made these work, but removing them on the penultimate day finally achieved validation / leaderboard correlation, getting me to public #105.</p>\n\n<p>As #105 slid down into the #200’s (I need finer resolution on that graph!) I gave in and as a backup plan rank averaged my main sub with Dirk’s which was worth 2e-5 of AUC and one spot on the private LB!</p>\n\n<pre><code>              private      public    model\nsingle lgbm 0.9825639    0.9811706   v12\nsingle lgbm 0.9825211    0.9810671   v13\nselection 1 0.9828068    0.9813324   rank average v12+v13 (Private #23)\nselection 2 0.9828279    0.9816611   rank average v12+v13 + Dirk (Private #22)\n</code></pre>",
  "messages": [
    {
      "id": "327432",
      "postDate": "05/11/2018 14:43:46",
      "content": "<p>Thanks to Kaggle &amp; TalkingData for another entertaining diversion :) Congrats to all winners and thanks to all who shared Kernels, tips and solution write-ups, they’re great to read.</p>\n\n<p>This is a really interesting data set and as I related in my <a href=\"https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots/code#310213\">story here</a>, click fraud is a really insidious problem that drives people to do really <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54765\">odd things</a> (thanks <a href=\"/yifanxie\">@yifanxie</a> - amazing to actually see it!)</p>\n\n<h2>Overview</h2>\n\n<p>I used broadly similar features to everyone else, next click times and count features, and stuck to the route recommended by early leaders – using lightgbm and training on the full training set – it surprised me that my 32Gb machine could. After reading the great <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325\">shared tips</a>, switching to creating a NumPy array upfront then populating that by loading feature columns was very efficient, with 33 features it barely even needed any swap space.</p>\n\n<p>I validated on the 2nd half of the last day of training for simplicity, just one validation set, then I’d save predictions in a dataframe and do groupby(‘hour’) and check AUC that way. I did pay more attention to the hours that matched the test set hours, but I thought other hours might lead to useful insights (but nothing noteworthy to report).</p>\n\n<p>My general process for one iteration was quite lightweight:</p>\n\n<ol>\n<li>an hour of feature engineering – save columns individually</li>\n<li>a 1.5-2 hour validation run</li>\n<li>a 3 hour full model build</li>\n</ol>\n\n<p>After setting this up, the only real manual effort is step 1, the validation run just determines # of trees (which was quite stable over feature sets, but sometimes repeated if the results were poor). The full model build is a one click process that builds a model, makes test set predictions with varying numbers of trees (to test assumptions about increased data set size &amp; scaling # of trees), then uploads the submissions via the highly recommended <a href=\"https://github.com/Kaggle/kaggle-api\">Kaggle API</a>.</p>\n\n<p>I managed 13 iterations, never more than one in a day, so it was a bit of a slow waiting game, and a long way from optimal.  Getting results in fewer iterations is down to some intuition and luck in selecting what works. There are a lot of features I tried that didn’t work, so my remaining features are mostly fairly baseline/simple things… (<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">Matrix factorization of IP vs app log click counts</a> is the one thing I really wish I’d tried, it will summarize an IP’s affinity for apps and interact very well with the app and IP count features… Well done to <a href=\"/cpmpml\">@cpmpml</a> &amp; other top teams :)</p>\n\n<h2>Entropy features</h2>\n\n<p>Whereas most people used nunique, I thought that might be noisy. With a click count histogram like e.g. counting os’s for an app = [12345, 5, 1, 1] it’s technically 4 unique os’s, but very low entropy (it’s more like one os).</p>\n\n<p>The original feature encodings made this easier: they were all 0..n, so it was simple to do it all in NumPy using <code>np.add.at()</code> to fill a 2D array with counts, then summarize like this:</p>\n\n<pre><code>from scipy.stats import entropy\n\n# e.g. a=’ip’ ; b=’app’\ncc = np.zeros((int(df[a].max()+1), int(df[b].max()+1)), dtype=np.int32)\nnp.add.at(cc, (df[a].values, df[b].values), 1)\n\n(cc&amp;gt;0).sum(0) # this is nunique for a\n(cc&amp;gt;0).sum(1) # this is nunique for b\nmap(entropy, cc)    # entropy of a over b\nmap(entropy, cc.T)  # entropy of b over a\n</code></pre>\n\n<p>All two way interactions were possible with 32Gb of RAM, and took only minutes to compute this way. I also combined device &amp; os into one new categorical (only 6561 unique values), and app &amp; channel (1518 unique values), which enabled some 3-way interactions like entropy of IP over (device, os).</p>\n\n<h2>Sessions</h2>\n\n<p>Four new features:</p>\n\n<ul>\n<li>sort by ip, dev, os</li>\n<li>define max gap e.g. 60 seconds</li>\n<li>every time there is a change in ip/device/os, or a gap in the clicks, assign categorical session ID</li>\n<li>for sessions, summarize:\n<ul><li>duration</li>\n<li>count</li>\n<li>click rate (duration/count)</li>\n<li>app entropy</li></ul></li>\n</ul>\n\n<p>I intended to add entropy over gaps, e.g. lots of similar gaps of 5-6 secs = low entropy = bots. Thought maybe click rate was enough.</p>\n\n<h2>Naive Count Features</h2>\n\n<p>Normally, count features group on one or more columns and count the actual occurrences. Inspired by Naive Bayes classifiers: another view on the data is to assume the fields are independent, and do:</p>\n\n<ul>\n<li>univariate count</li>\n<li>divide counts by the number of rows to get a probability of seeing the value</li>\n<li>log() the probability</li>\n<li>sum different combinations of these columns similarly to normal feature interactions (i.e. -, +, /, *).</li>\n</ul>\n\n<p>Mathematically, this is the same as a geometric mean probability of seeing the record, under the assumption the columns are independent. Again, extremely fast to compute using NumPy alone.</p>\n\n<p>The combinations &amp; loading code I used:</p>\n\n<pre><code>def loadraw(name, dtype):\n    return np.fromfile(name, dtype=dtype)\n\ndef log_p_feat(t, col):\n    # t is ‘train’ or ‘test’\n    return loadraw('../feats/%s_log_p_%s'%(t,col), np.float32)\n\nx[:,a] = log_p_feat(t, 'ip') + log_p_feat(t, 'dev') + log_p_feat(t, 'os')\nx[:,a+1] = x[:,a] + log_p_feat(t, 'app')\nx[:,a+2] = x[:,a] + log_p_feat(t, 'chan')\nx[:,a+3] = x[:,a] + log_p_feat(t, 'app') + log_p_feat(t, 'chan')\n</code></pre>\n\n<p>Adding these features in with normal 2D/3D counts might pick up on interesting things. With these ‘naive’ counts: rare values overlap - a low (log) probability row might be rare IP with common OS <strong><em>or</em></strong> a rare OS with common IP. (Although IP was much higher cardinality, so probably dominates the log sum. A weighted mix might be better.)</p>\n\n<p>This can produce a very large number of unique values, so lightgbm value binning might come to the rescue here and avoid over-fitting. I’m not 100% convinced this was a good idea… but it did lead to a good single model (0.9811 public, 0.9825 private) where these naive count features had high feature importances.</p>\n\n<h2>Bit Shifting</h2>\n\n<p>Small tip: by looking at the maximum values for each column, you can work out how many bits are required to store each, and happily, the five categoricals all fit in a single 64 bit integer. You can pack them like this:</p>\n\n<pre><code>def do_pack(df):\n    a = 0\n    a += df.app.values.astype(np.uint64)\n    a += df.os.values.astype(np.uint64)&amp;lt;&amp;lt;10\n    a += df.channel.values.astype(np.uint64)&amp;lt;&amp;lt;20\n    a += df.ip.values.astype(np.uint64)&amp;lt;&amp;lt;30\n    a += df.device.values.astype(np.uint64)&amp;lt;&amp;lt;50\n    return a\n</code></pre>\n\n<p>They can be unpacked (these work on values or NumPy arrays):</p>\n\n<pre><code>def fapp(v): return v&amp;amp;0x3ff\ndef fos(v): return (v&amp;gt;&amp;gt;10)&amp;amp;0x3ff\ndef fchan(v): return (v&amp;gt;&amp;gt;20)&amp;amp;0x3ff\ndef fip(v): return (v&amp;gt;&amp;gt;30)&amp;amp;0xfffff\ndef fdev(v): return (v&amp;gt;&amp;gt;50)&amp;amp;0xfffff\n</code></pre>\n\n<p>This is vectorized, so runs extremely quickly, and the result is like a hash, but without collisions. You can use the result to do groupby() operations - more efficiently in both time &amp; RAM. You can also ignore fields by setting all their bits to one, e.g. to ignore device <code>a |= 0xfffffl&amp;lt;&amp;lt;50</code>. (Because 0 was a valid value for each field.) Using <code>np.save()</code> (or <code>a.tofile</code> and <code>np.fromfile</code>) on the 64 bit array you could save the entire 4 days of train &amp; full test into 1.9Gb on disk. Times could be similarly packed into 32 bits, about 900Mb… With these copies of the data I actually managed to do some feature exploration in idle moments (in Java) on an old 8Gb MacBook!</p>\n\n<h2>Categoricals</h2>\n\n<p>Grouping on all five combined fields, there were many different series, some with 10k-20k appearances, but a long tail of shorter series. Instead of separate count &amp; cumcount features I tried a categorical feature to mark where the row is in the rarer/shorter series. Using max_bin of 255, there are 8 bits to play with: 4 bits to mark the length of the series (1..15) and 4 bits to mark where in the series the record appeared (0..14).</p>\n\n<p>Where v is all five fields combined as above, use groupby(v).size() to get df.series_size and groupby(v).cumcount() to get df.series_cumcount, then:</p>\n\n<pre><code>idx = df.series_size.values&amp;lt;=15\ncat = np.where(idx, df.series_size.values&amp;lt;&amp;lt;4, 0)\ncat |= np.where(idx, df.series_cumcount.values, 0)\n</code></pre>\n\n<p>So this is one feature that encodes the length of the ‘series’ of [ip,dev,os,app,chan] <strong>and</strong> the position within the series. Tree splits can then address subsets like:</p>\n\n<pre><code>if cat == 0x32\nif 3_element_series and this_record_is_last_in_series\n\nif cat == 0x20\nif 2_element_series and this_record_is_first_in_series\n</code></pre>\n\n<p>To make the latest versions of lightgbm do these kind of splits you need to set the <code>max_cat_to_onehot</code> parameter correctly, in this case to 256 or more – but beware this applies to all categorical features, so for example channel may be used for one hot style splits too – it depends on how many values are seen in each column at dataset construction time (lightgbm samples data to construct the bins). All model parameters/regularizations are trade-offs of some kind.</p>\n\n<p>Limited time and iterations mean I can’t be sure this didn’t overfit, particularly for series that span train &amp; test times, and things like “length 14 series, element 4” which is not likely to mean anything (the hope is, this could uncover some really quirky leak!) Perhaps addressing only shorter series would be better, or chunking the series by time. I relied on the model to tell me which were more important - a quick inspection of models shows low values like 0x20, 0x21 were used the most.</p>\n\n<h2>Leaderboard</h2>\n\n<p>From <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56524\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56524</a></p>\n\n<p><img src=\"https://kaggle2.blob.core.windows.net/forum-message-attachments/326995/9398/jt_public_lb.png\" alt=\"JT public LB\"></p>\n\n<p>For me, validation scores stayed very consistent, but leaderboard scores were mysteriously lower. It was only eventually removing some features that fixed this: all of the count features that used the corrected Chinese timezone day. (These were: <code>ip_day_count</code>, <code>app_channel_day_count</code>, <code>device_os_day_count</code>.) I’m still not sure what caused that, and it’s puzzling that other’s made these work, but removing them on the penultimate day finally achieved validation / leaderboard correlation, getting me to public #105.</p>\n\n<p>As #105 slid down into the #200’s (I need finer resolution on that graph!) I gave in and as a backup plan rank averaged my main sub with Dirk’s which was worth 2e-5 of AUC and one spot on the private LB!</p>\n\n<pre><code>              private      public    model\nsingle lgbm 0.9825639    0.9811706   v12\nsingle lgbm 0.9825211    0.9810671   v13\nselection 1 0.9828068    0.9813324   rank average v12+v13 (Private #23)\nselection 2 0.9828279    0.9816611   rank average v12+v13 + Dirk (Private #22)\n</code></pre>",
      "rawMarkdown": "Thanks to Kaggle &amp; TalkingData for another entertaining diversion :) Congrats to all winners and thanks to all who shared Kernels, tips and solution write-ups, they’re great to read.\n\nThis is a really interesting data set and as I related in my [story here][1], click fraud is a really insidious problem that drives people to do really [odd things][2] (thanks @yifanxie - amazing to actually see it!)\n\n## Overview\n\nI used broadly similar features to everyone else, next click times and count features, and stuck to the route recommended by early leaders – using lightgbm and training on the full training set – it surprised me that my 32Gb machine could. After reading the great [shared tips][3], switching to creating a NumPy array upfront then populating that by loading feature columns was very efficient, with 33 features it barely even needed any swap space.\n\nI validated on the 2nd half of the last day of training for simplicity, just one validation set, then I’d save predictions in a dataframe and do groupby(‘hour’) and check AUC that way. I did pay more attention to the hours that matched the test set hours, but I thought other hours might lead to useful insights (but nothing noteworthy to report).\n\nMy general process for one iteration was quite lightweight:\n\n1. an hour of feature engineering – save columns individually\n2. a 1.5-2 hour validation run\n3. a 3 hour full model build\n\nAfter setting this up, the only real manual effort is step 1, the validation run just determines # of trees (which was quite stable over feature sets, but sometimes repeated if the results were poor). The full model build is a one click process that builds a model, makes test set predictions with varying numbers of trees (to test assumptions about increased data set size &amp; scaling # of trees), then uploads the submissions via the highly recommended [Kaggle API][4].\n\nI managed 13 iterations, never more than one in a day, so it was a bit of a slow waiting game, and a long way from optimal.  Getting results in fewer iterations is down to some intuition and luck in selecting what works. There are a lot of features I tried that didn’t work, so my remaining features are mostly fairly baseline/simple things… ([Matrix factorization of IP vs app log click counts][5] is the one thing I really wish I’d tried, it will summarize an IP’s affinity for apps and interact very well with the app and IP count features… Well done to @cpmpml &amp; other top teams :)\n\n## Entropy features\n\nWhereas most people used nunique, I thought that might be noisy. With a click count histogram like e.g. counting os’s for an app = [12345, 5, 1, 1] it’s technically 4 unique os’s, but very low entropy (it’s more like one os).\n\nThe original feature encodings made this easier: they were all 0..n, so it was simple to do it all in NumPy using `np.add.at()` to fill a 2D array with counts, then summarize like this:\n\n    from scipy.stats import entropy\n    \n    # e.g. a=’ip’ ; b=’app’\n    cc = np.zeros((int(df[a].max()+1), int(df[b].max()+1)), dtype=np.int32)\n    np.add.at(cc, (df[a].values, df[b].values), 1)\n    \n    (cc&gt;0).sum(0) # this is nunique for a\n    (cc&gt;0).sum(1) # this is nunique for b\n    map(entropy, cc)    # entropy of a over b\n    map(entropy, cc.T)  # entropy of b over a\n\nAll two way interactions were possible with 32Gb of RAM, and took only minutes to compute this way. I also combined device &amp; os into one new categorical (only 6561 unique values), and app &amp; channel (1518 unique values), which enabled some 3-way interactions like entropy of IP over (device, os).\n\n## Sessions\n\nFour new features:\n\n - sort by ip, dev, os\n - define max gap e.g. 60 seconds\n - every time there is a change in ip/device/os, or a gap in the clicks, assign categorical session ID\n - for sessions, summarize:\n    - duration\n    - count\n    - click rate (duration/count)\n    - app entropy\n\nI intended to add entropy over gaps, e.g. lots of similar gaps of 5-6 secs = low entropy = bots. Thought maybe click rate was enough.\n\n## Naive Count Features\n\nNormally, count features group on one or more columns and count the actual occurrences. Inspired by Naive Bayes classifiers: another view on the data is to assume the fields are independent, and do:\n\n - univariate count\n - divide counts by the number of rows to get a probability of seeing the value\n - log() the probability\n - sum different combinations of these columns similarly to normal feature interactions (i.e. -, +, /, *).\n\nMathematically, this is the same as a geometric mean probability of seeing the record, under the assumption the columns are independent. Again, extremely fast to compute using NumPy alone.\n\nThe combinations &amp; loading code I used:\n\n\n    def loadraw(name, dtype):\n        return np.fromfile(name, dtype=dtype)\n    \n    def log_p_feat(t, col):\n        # t is ‘train’ or ‘test’\n        return loadraw('../feats/%s_log_p_%s'%(t,col), np.float32)\n    \n    x[:,a] = log_p_feat(t, 'ip') + log_p_feat(t, 'dev') + log_p_feat(t, 'os')\n    x[:,a+1] = x[:,a] + log_p_feat(t, 'app')\n    x[:,a+2] = x[:,a] + log_p_feat(t, 'chan')\n    x[:,a+3] = x[:,a] + log_p_feat(t, 'app') + log_p_feat(t, 'chan')\n\n\nAdding these features in with normal 2D/3D counts might pick up on interesting things. With these ‘naive’ counts: rare values overlap - a low (log) probability row might be rare IP with common OS ***or*** a rare OS with common IP. (Although IP was much higher cardinality, so probably dominates the log sum. A weighted mix might be better.)\n\nThis can produce a very large number of unique values, so lightgbm value binning might come to the rescue here and avoid over-fitting. I’m not 100% convinced this was a good idea… but it did lead to a good single model (0.9811 public, 0.9825 private) where these naive count features had high feature importances.\n\n## Bit Shifting\n\nSmall tip: by looking at the maximum values for each column, you can work out how many bits are required to store each, and happily, the five categoricals all fit in a single 64 bit integer. You can pack them like this:\n\n    def do_pack(df):\n        a = 0\n        a += df.app.values.astype(np.uint64)\n        a += df.os.values.astype(np.uint64)&lt;&lt;10\n        a += df.channel.values.astype(np.uint64)&lt;&lt;20\n        a += df.ip.values.astype(np.uint64)&lt;&lt;30\n        a += df.device.values.astype(np.uint64)&lt;&lt;50\n        return a\n\nThey can be unpacked (these work on values or NumPy arrays):\n\n    def fapp(v): return v&amp;0x3ff\n    def fos(v): return (v&gt;&gt;10)&amp;0x3ff\n    def fchan(v): return (v&gt;&gt;20)&amp;0x3ff\n    def fip(v): return (v&gt;&gt;30)&amp;0xfffff\n    def fdev(v): return (v&gt;&gt;50)&amp;0xfffff\n\nThis is vectorized, so runs extremely quickly, and the result is like a hash, but without collisions. You can use the result to do groupby() operations - more efficiently in both time &amp; RAM. You can also ignore fields by setting all their bits to one, e.g. to ignore device `a |= 0xfffffl&lt;&lt;50`. (Because 0 was a valid value for each field.) Using `np.save()` (or `a.tofile` and `np.fromfile`) on the 64 bit array you could save the entire 4 days of train &amp; full test into 1.9Gb on disk. Times could be similarly packed into 32 bits, about 900Mb… With these copies of the data I actually managed to do some feature exploration in idle moments (in Java) on an old 8Gb MacBook!\n\n## Categoricals\n\nGrouping on all five combined fields, there were many different series, some with 10k-20k appearances, but a long tail of shorter series. Instead of separate count &amp; cumcount features I tried a categorical feature to mark where the row is in the rarer/shorter series. Using max_bin of 255, there are 8 bits to play with: 4 bits to mark the length of the series (1..15) and 4 bits to mark where in the series the record appeared (0..14).\n\nWhere v is all five fields combined as above, use groupby(v).size() to get df.series_size and groupby(v).cumcount() to get df.series_cumcount, then:\n\n    idx = df.series_size.values&lt;=15\n    cat = np.where(idx, df.series_size.values&lt;&lt;4, 0)\n    cat |= np.where(idx, df.series_cumcount.values, 0)\n\nSo this is one feature that encodes the length of the ‘series’ of [ip,dev,os,app,chan] **and** the position within the series. Tree splits can then address subsets like:\n\n    if cat == 0x32\n    if 3_element_series and this_record_is_last_in_series\n    \n    if cat == 0x20\n    if 2_element_series and this_record_is_first_in_series\n\nTo make the latest versions of lightgbm do these kind of splits you need to set the `max_cat_to_onehot` parameter correctly, in this case to 256 or more – but beware this applies to all categorical features, so for example channel may be used for one hot style splits too – it depends on how many values are seen in each column at dataset construction time (lightgbm samples data to construct the bins). All model parameters/regularizations are trade-offs of some kind.\n\nLimited time and iterations mean I can’t be sure this didn’t overfit, particularly for series that span train &amp; test times, and things like “length 14 series, element 4” which is not likely to mean anything (the hope is, this could uncover some really quirky leak!) Perhaps addressing only shorter series would be better, or chunking the series by time. I relied on the model to tell me which were more important - a quick inspection of models shows low values like 0x20, 0x21 were used the most.\n\n## Leaderboard\n\nFrom https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56524\n\n![JT public LB][6]\n\nFor me, validation scores stayed very consistent, but leaderboard scores were mysteriously lower. It was only eventually removing some features that fixed this: all of the count features that used the corrected Chinese timezone day. (These were: `ip_day_count`, `app_channel_day_count`, `device_os_day_count`.) I’m still not sure what caused that, and it’s puzzling that other’s made these work, but removing them on the penultimate day finally achieved validation / leaderboard correlation, getting me to public #105.\n\nAs #105 slid down into the #200’s (I need finer resolution on that graph!) I gave in and as a backup plan rank averaged my main sub with Dirk’s which was worth 2e-5 of AUC and one spot on the private LB!\n\n\n                  private      public    model\n    single lgbm 0.9825639    0.9811706   v12\n    single lgbm 0.9825211    0.9810671   v13\n    selection 1 0.9828068    0.9813324   rank average v12+v13 (Private #23)\n    selection 2 0.9828279    0.9816611   rank average v12+v13 + Dirk (Private #22)\n\n\n  [1]: https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots/code#310213\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54765\n  [3]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325\n  [4]: https://github.com/Kaggle/kaggle-api\n  [5]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\n  [6]: https://kaggle2.blob.core.windows.net/forum-message-attachments/326995/9398/jt_public_lb.png",
      "votes": null
    },
    {
      "id": "327437",
      "postDate": "05/11/2018 14:56:58",
      "content": "<p>Great writeup, thanks, and congrats on the good result.  Thanks for citing me as well ;)</p>\n\n<p>I like your entropy approach and you naive count.  Will try to reuse them in the future!</p>\n\n<p>When I think you wrote you have not much to say ;)</p>",
      "rawMarkdown": "Great writeup, thanks, and congrats on the good result.  Thanks for citing me as well ;)\n\nI like your entropy approach and you naive count.  Will try to reuse them in the future!\n\nWhen I think you wrote you have not much to say ;)",
      "votes": null
    },
    {
      "id": "327452",
      "postDate": "05/11/2018 15:24:43",
      "content": "<p>Thanks... sorry about the wait - these are some pretty lightweight tips, most of which I could probably have safely shared long before the end :)</p>",
      "rawMarkdown": "Thanks... sorry about the wait - these are some pretty lightweight tips, most of which I could probably have safely shared long before the end :)",
      "votes": null
    },
    {
      "id": "327584",
      "postDate": "05/11/2018 23:40:23",
      "content": "<p><a href=\"/jtrotman\">@jtrotman</a> wow, some advanced coding there - much more efficient than the pandas stuff I do\nthanks for sharing, things like the entropy features and bit shifting are stuff that I could never come up on my own\nvery useful for all the \"big data\" competitions we have at the moment :)</p>",
      "rawMarkdown": "jtrotman wow, some advanced coding there - much more efficient than the pandas stuff I do\nthanks for sharing, things like the entropy features and bit shifting are stuff that I could never come up on my own\nvery useful for all the \"big data\" competitions we have at the moment :)",
      "votes": null
    },
    {
      "id": "327746",
      "postDate": "05/12/2018 10:51:13",
      "content": "<p>Thanks Yifan - my background is in embedded software engineering, so I have a habit of trying to cram things into as few bits as possible :)</p>\n\n<p>The use of entropy is something I discovered in my first Kaggle competition, on a similar theme to this one: <a href=\"https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot\">Facebook Recruiting IV: Human or Robot?</a> in the <a href=\"https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot/discussion/14628\">secret sauce thread</a> post by <a href=\"/smallyellowduck\">@smallyellowduck</a> - I remember being blown away by that solution at the time - and thinking that was a lot of data to deal with... in another few years this dataset might start to look small, and if Kaggle keep updating Kernel capability roughly in line with Moore's Law, tips like this could have a short half life!</p>",
      "rawMarkdown": "Thanks Yifan - my background is in embedded software engineering, so I have a habit of trying to cram things into as few bits as possible :)\n\nThe use of entropy is something I discovered in my first Kaggle competition, on a similar theme to this one: [Facebook Recruiting IV: Human or Robot?][1] in the [secret sauce thread][2] post by @smallyellowduck - I remember being blown away by that solution at the time - and thinking that was a lot of data to deal with... in another few years this dataset might start to look small, and if Kaggle keep updating Kernel capability roughly in line with Moore's Law, tips like this could have a short half life!\n\n  [1]: https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot\n  [2]: https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot/discussion/14628",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 327437,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/11/2018 14:56:58",
      "content": "<p>Great writeup, thanks, and congrats on the good result.  Thanks for citing me as well ;)</p>\n\n<p>I like your entropy approach and you naive count.  Will try to reuse them in the future!</p>\n\n<p>When I think you wrote you have not much to say ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 327452,
          "author_name": "jtrotman",
          "author_url": "",
          "post_date": "05/11/2018 15:24:43",
          "content": "<p>Thanks... sorry about the wait - these are some pretty lightweight tips, most of which I could probably have safely shared long before the end :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327584,
      "author_name": "yifanxie",
      "author_url": "",
      "post_date": "05/11/2018 23:40:23",
      "content": "<p><a href=\"/jtrotman\">@jtrotman</a> wow, some advanced coding there - much more efficient than the pandas stuff I do\nthanks for sharing, things like the entropy features and bit shifting are stuff that I could never come up on my own\nvery useful for all the \"big data\" competitions we have at the moment :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 327746,
          "author_name": "jtrotman",
          "author_url": "",
          "post_date": "05/12/2018 10:51:13",
          "content": "<p>Thanks Yifan - my background is in embedded software engineering, so I have a habit of trying to cram things into as few bits as possible :)</p>\n\n<p>The use of entropy is something I discovered in my first Kaggle competition, on a similar theme to this one: <a href=\"https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot\">Facebook Recruiting IV: Human or Robot?</a> in the <a href=\"https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot/discussion/14628\">secret sauce thread</a> post by <a href=\"/smallyellowduck\">@smallyellowduck</a> - I remember being blown away by that solution at the time - and thinking that was a lot of data to deal with... in another few years this dataset might start to look small, and if Kaggle keep updating Kernel capability roughly in line with Moore's Law, tips like this could have a short half life!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "327432": "Thanks to Kaggle &amp; TalkingData for another entertaining diversion :) Congrats to all winners and thanks to all who shared Kernels, tips and solution write-ups, they’re great to read.\n\nThis is a really interesting data set and as I related in my [story here][1], click fraud is a really insidious problem that drives people to do really [odd things][2] (thanks @yifanxie - amazing to actually see it!)\n\n## Overview\n\nI used broadly similar features to everyone else, next click times and count features, and stuck to the route recommended by early leaders – using lightgbm and training on the full training set – it surprised me that my 32Gb machine could. After reading the great [shared tips][3], switching to creating a NumPy array upfront then populating that by loading feature columns was very efficient, with 33 features it barely even needed any swap space.\n\nI validated on the 2nd half of the last day of training for simplicity, just one validation set, then I’d save predictions in a dataframe and do groupby(‘hour’) and check AUC that way. I did pay more attention to the hours that matched the test set hours, but I thought other hours might lead to useful insights (but nothing noteworthy to report).\n\nMy general process for one iteration was quite lightweight:\n\n1. an hour of feature engineering – save columns individually\n2. a 1.5-2 hour validation run\n3. a 3 hour full model build\n\nAfter setting this up, the only real manual effort is step 1, the validation run just determines # of trees (which was quite stable over feature sets, but sometimes repeated if the results were poor). The full model build is a one click process that builds a model, makes test set predictions with varying numbers of trees (to test assumptions about increased data set size &amp; scaling # of trees), then uploads the submissions via the highly recommended [Kaggle API][4].\n\nI managed 13 iterations, never more than one in a day, so it was a bit of a slow waiting game, and a long way from optimal.  Getting results in fewer iterations is down to some intuition and luck in selecting what works. There are a lot of features I tried that didn’t work, so my remaining features are mostly fairly baseline/simple things… ([Matrix factorization of IP vs app log click counts][5] is the one thing I really wish I’d tried, it will summarize an IP’s affinity for apps and interact very well with the app and IP count features… Well done to @cpmpml &amp; other top teams :)\n\n## Entropy features\n\nWhereas most people used nunique, I thought that might be noisy. With a click count histogram like e.g. counting os’s for an app = [12345, 5, 1, 1] it’s technically 4 unique os’s, but very low entropy (it’s more like one os).\n\nThe original feature encodings made this easier: they were all 0..n, so it was simple to do it all in NumPy using `np.add.at()` to fill a 2D array with counts, then summarize like this:\n\n    from scipy.stats import entropy\n    \n    # e.g. a=’ip’ ; b=’app’\n    cc = np.zeros((int(df[a].max()+1), int(df[b].max()+1)), dtype=np.int32)\n    np.add.at(cc, (df[a].values, df[b].values), 1)\n    \n    (cc&gt;0).sum(0) # this is nunique for a\n    (cc&gt;0).sum(1) # this is nunique for b\n    map(entropy, cc)    # entropy of a over b\n    map(entropy, cc.T)  # entropy of b over a\n\nAll two way interactions were possible with 32Gb of RAM, and took only minutes to compute this way. I also combined device &amp; os into one new categorical (only 6561 unique values), and app &amp; channel (1518 unique values), which enabled some 3-way interactions like entropy of IP over (device, os).\n\n## Sessions\n\nFour new features:\n\n - sort by ip, dev, os\n - define max gap e.g. 60 seconds\n - every time there is a change in ip/device/os, or a gap in the clicks, assign categorical session ID\n - for sessions, summarize:\n    - duration\n    - count\n    - click rate (duration/count)\n    - app entropy\n\nI intended to add entropy over gaps, e.g. lots of similar gaps of 5-6 secs = low entropy = bots. Thought maybe click rate was enough.\n\n## Naive Count Features\n\nNormally, count features group on one or more columns and count the actual occurrences. Inspired by Naive Bayes classifiers: another view on the data is to assume the fields are independent, and do:\n\n - univariate count\n - divide counts by the number of rows to get a probability of seeing the value\n - log() the probability\n - sum different combinations of these columns similarly to normal feature interactions (i.e. -, +, /, *).\n\nMathematically, this is the same as a geometric mean probability of seeing the record, under the assumption the columns are independent. Again, extremely fast to compute using NumPy alone.\n\nThe combinations &amp; loading code I used:\n\n\n    def loadraw(name, dtype):\n        return np.fromfile(name, dtype=dtype)\n    \n    def log_p_feat(t, col):\n        # t is ‘train’ or ‘test’\n        return loadraw('../feats/%s_log_p_%s'%(t,col), np.float32)\n    \n    x[:,a] = log_p_feat(t, 'ip') + log_p_feat(t, 'dev') + log_p_feat(t, 'os')\n    x[:,a+1] = x[:,a] + log_p_feat(t, 'app')\n    x[:,a+2] = x[:,a] + log_p_feat(t, 'chan')\n    x[:,a+3] = x[:,a] + log_p_feat(t, 'app') + log_p_feat(t, 'chan')\n\n\nAdding these features in with normal 2D/3D counts might pick up on interesting things. With these ‘naive’ counts: rare values overlap - a low (log) probability row might be rare IP with common OS ***or*** a rare OS with common IP. (Although IP was much higher cardinality, so probably dominates the log sum. A weighted mix might be better.)\n\nThis can produce a very large number of unique values, so lightgbm value binning might come to the rescue here and avoid over-fitting. I’m not 100% convinced this was a good idea… but it did lead to a good single model (0.9811 public, 0.9825 private) where these naive count features had high feature importances.\n\n## Bit Shifting\n\nSmall tip: by looking at the maximum values for each column, you can work out how many bits are required to store each, and happily, the five categoricals all fit in a single 64 bit integer. You can pack them like this:\n\n    def do_pack(df):\n        a = 0\n        a += df.app.values.astype(np.uint64)\n        a += df.os.values.astype(np.uint64)&lt;&lt;10\n        a += df.channel.values.astype(np.uint64)&lt;&lt;20\n        a += df.ip.values.astype(np.uint64)&lt;&lt;30\n        a += df.device.values.astype(np.uint64)&lt;&lt;50\n        return a\n\nThey can be unpacked (these work on values or NumPy arrays):\n\n    def fapp(v): return v&amp;0x3ff\n    def fos(v): return (v&gt;&gt;10)&amp;0x3ff\n    def fchan(v): return (v&gt;&gt;20)&amp;0x3ff\n    def fip(v): return (v&gt;&gt;30)&amp;0xfffff\n    def fdev(v): return (v&gt;&gt;50)&amp;0xfffff\n\nThis is vectorized, so runs extremely quickly, and the result is like a hash, but without collisions. You can use the result to do groupby() operations - more efficiently in both time &amp; RAM. You can also ignore fields by setting all their bits to one, e.g. to ignore device `a |= 0xfffffl&lt;&lt;50`. (Because 0 was a valid value for each field.) Using `np.save()` (or `a.tofile` and `np.fromfile`) on the 64 bit array you could save the entire 4 days of train &amp; full test into 1.9Gb on disk. Times could be similarly packed into 32 bits, about 900Mb… With these copies of the data I actually managed to do some feature exploration in idle moments (in Java) on an old 8Gb MacBook!\n\n## Categoricals\n\nGrouping on all five combined fields, there were many different series, some with 10k-20k appearances, but a long tail of shorter series. Instead of separate count &amp; cumcount features I tried a categorical feature to mark where the row is in the rarer/shorter series. Using max_bin of 255, there are 8 bits to play with: 4 bits to mark the length of the series (1..15) and 4 bits to mark where in the series the record appeared (0..14).\n\nWhere v is all five fields combined as above, use groupby(v).size() to get df.series_size and groupby(v).cumcount() to get df.series_cumcount, then:\n\n    idx = df.series_size.values&lt;=15\n    cat = np.where(idx, df.series_size.values&lt;&lt;4, 0)\n    cat |= np.where(idx, df.series_cumcount.values, 0)\n\nSo this is one feature that encodes the length of the ‘series’ of [ip,dev,os,app,chan] **and** the position within the series. Tree splits can then address subsets like:\n\n    if cat == 0x32\n    if 3_element_series and this_record_is_last_in_series\n    \n    if cat == 0x20\n    if 2_element_series and this_record_is_first_in_series\n\nTo make the latest versions of lightgbm do these kind of splits you need to set the `max_cat_to_onehot` parameter correctly, in this case to 256 or more – but beware this applies to all categorical features, so for example channel may be used for one hot style splits too – it depends on how many values are seen in each column at dataset construction time (lightgbm samples data to construct the bins). All model parameters/regularizations are trade-offs of some kind.\n\nLimited time and iterations mean I can’t be sure this didn’t overfit, particularly for series that span train &amp; test times, and things like “length 14 series, element 4” which is not likely to mean anything (the hope is, this could uncover some really quirky leak!) Perhaps addressing only shorter series would be better, or chunking the series by time. I relied on the model to tell me which were more important - a quick inspection of models shows low values like 0x20, 0x21 were used the most.\n\n## Leaderboard\n\nFrom https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56524\n\n![JT public LB][6]\n\nFor me, validation scores stayed very consistent, but leaderboard scores were mysteriously lower. It was only eventually removing some features that fixed this: all of the count features that used the corrected Chinese timezone day. (These were: `ip_day_count`, `app_channel_day_count`, `device_os_day_count`.) I’m still not sure what caused that, and it’s puzzling that other’s made these work, but removing them on the penultimate day finally achieved validation / leaderboard correlation, getting me to public #105.\n\nAs #105 slid down into the #200’s (I need finer resolution on that graph!) I gave in and as a backup plan rank averaged my main sub with Dirk’s which was worth 2e-5 of AUC and one spot on the private LB!\n\n\n                  private      public    model\n    single lgbm 0.9825639    0.9811706   v12\n    single lgbm 0.9825211    0.9810671   v13\n    selection 1 0.9828068    0.9813324   rank average v12+v13 (Private #23)\n    selection 2 0.9828279    0.9816611   rank average v12+v13 + Dirk (Private #22)\n\n\n  [1]: https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots/code#310213\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54765\n  [3]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325\n  [4]: https://github.com/Kaggle/kaggle-api\n  [5]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\n  [6]: https://kaggle2.blob.core.windows.net/forum-message-attachments/326995/9398/jt_public_lb.png",
    "327437": "Great writeup, thanks, and congrats on the good result.  Thanks for citing me as well ;)\n\nI like your entropy approach and you naive count.  Will try to reuse them in the future!\n\nWhen I think you wrote you have not much to say ;)",
    "327452": "Thanks... sorry about the wait - these are some pretty lightweight tips, most of which I could probably have safely shared long before the end :)",
    "327584": "jtrotman wow, some advanced coding there - much more efficient than the pandas stuff I do\nthanks for sharing, things like the entropy features and bit shifting are stuff that I could never come up on my own\nvery useful for all the \"big data\" competitions we have at the moment :)",
    "327746": "Thanks Yifan - my background is in embedded software engineering, so I have a habit of trying to cram things into as few bits as possible :)\n\nThe use of entropy is something I discovered in my first Kaggle competition, on a similar theme to this one: [Facebook Recruiting IV: Human or Robot?][1] in the [secret sauce thread][2] post by @smallyellowduck - I remember being blown away by that solution at the time - and thinking that was a lot of data to deal with... in another few years this dataset might start to look small, and if Kaggle keep updating Kernel capability roughly in line with Moore's Law, tips like this could have a short half life!\n\n  [1]: https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot\n  [2]: https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot/discussion/14628"
  },
  "source": "meta"
}