{
  "id": 466504,
  "title": "🥉3rd Place Solution Writeup",
  "url": "/competitions/vpn-classification/writeups/pietro-maldini-3rd-place-solution-writeup",
  "author_name": "",
  "post_date": "2024-01-08T23:04:42.315199200Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<h2>Greetings Kagglers! 🎉</h2>\n<p>I am thrilled to share my journey during this competition that led me to the 🥉third position on both public and private leaderboards.</p>\n<p>In this competion I faced two challenges:</p>\n<ol>\n<li>Limited knowlegde of networks and their security </li>\n<li>Unbalance in the dataset</li>\n</ol>\n<h1>My Journey 🚀</h1>\n<h2>EDA 📊</h2>\n<p>I started this competition from some basic analysis of the dataset provided by the organizers and I've done some basic EDA to visualize patterns in the dataset, in particular due to the unbalance of the dataset this first visual analysis led to practically no insights, maybe my inexperience in the field made me overlook some important insight that could have helped later on.</p>\n<h2>Modelling, Phase 1 🔬</h2>\n<p>After performing EDA I tried creating some basic feature from the dataset, for example I counted the occurences of a specific IP, i found the attack duration from a specific IP, I found the most frequent attacker country for an IP.<br>\nAt the beginning I didn't include shodan features.</p>\n<p>In this first phase I used LGBM Classifier as my model. <br>\nAt the beginning my scores were around 0.5 on the public leaderboard but also locally. <br>\nAfter analyzing the predictions of the model I noticed that the predictions were strongly skewed and the majority was under 0.5, so the threshold for considering a sample a Proxy/VPN could not be kept at 0.5.</p>\n<p>Changing the threshold to values around 0.3 helped me boost the initial scores of my LGBM classfier models to around 0.6.</p>\n<h2>Modelling, Phase 2🧪</h2>\n<p>The big jump from 0.6 to around 0.75 happened when I tried to combine my features with those provided by the organizer in the <a href=\"https://www.kaggle.com/code/alpacads/vpn-proxy-demo-notebook\" target=\"_blank\">Demo Notebook</a>, in particular I used all the featur provided by the organizers and added some new features.</p>\n<ul>\n<li>I added new ports to be counted taking the new values from the most frequent ports seen in the dataset.</li>\n<li>I counted the number of ports open with tcp and udp protocols</li>\n<li>I counted how many values were present (non Null) for jarm, ja3s and headers_hash in the dataset for that IP</li>\n</ul>\n<p>To these features I added frequency and a target encoding of the attacker country and the frequency of attacker name (no target encoding since this would probably lead to overfitting features).</p>\n<p>From the features computed in Phase 1 I kept the count of attacks, the duration of the attack and the standard deviation of the timestamps of the attack.</p>\n<p>Using those features boosted a lot the model leading almost to the final results.</p>\n<h2>Modeling, Phase 3⚗️</h2>\n<p>After finding the features in Phase 2 that seemed to bring lot of information to the model I was using I tried using several models.<br>\nIn particular I tried these models and found those results (F1 Scores):</p>\n<ol>\n<li>SGDClassifier: 0.446</li>\n<li>LogisticRegression: 0.568</li>\n<li>KNeighborsClassifier (K=7,weights=distance): 0.595</li>\n<li>GaussianNB: 0.160</li>\n<li>DecisionTreeClassifier: 0.710</li>\n<li>ExtraTreesClassifier: 0.747</li>\n<li>HistGradientBoostingClassifier: 0.767</li>\n<li>RandomForestClassifier: 0.780</li>\n</ol>\n<p>These results were computed training the models using GridSearchCV to find some good sets of hyperparameters.</p>\n<h2>Inference and final results 📝</h2>\n<p>Using the results found in Phase 3 I decided to use a Random Forest Classifier as my final model, in particular I trained 5 models, using KFold to generate 5 sets of train and validation dataset one for each model.</p>\n<p>I averaged the prediction of the five models and used a custom threshold for binarization (around 0.3) that was computed by averaging the best threshold (rounded to 2 decimals) for each of the models on their corresponding validation set.</p>\n<p>This allowed my final prediction to have a stable result on both public and private leaderboard.</p>\n<p>In particular for the final submission on the leaderboard I took the result from different runs of this process (not seeded) and performed voting ensemble among them.</p>\n<h1>Conclusions ✍</h1>\n<p>This was a great learning experience, I learned a lot and had to automate as much as possible modelling to try out many models and possibilities to get which models performed the best.</p>\n<p>I hope to get to interact with more people during the next challenge.<br>\nI was quite too busy and had too small time for the challenge, I had practically my notebooks running experiments and also submitting for me because I had just that free time to start the notebooks.</p>\n<p>I'm looking forward to learn from other people solution and see what different approaches were tested by other participants.</p>\n<h1>Possible improvements, Next steps👣</h1>\n<p>Better analysis of model results splitting between precision and recall scores to better understand the strength and weaknessed of the models.<br>\nFor example it could be possible to use a high recall model to filter candidate positive samples and after that use a high precision model to get accurate results on the filtered values.</p>\n<p>Ensembles of different models if implemented correctly could also improve the score, the reliability and the stability of the results.</p>\n<h1>Thank you for reading. Happy Kaggling🌟</h1>",
  "messages": [
    {
      "id": "2592918",
      "postDate": "01/08/2024 23:04:42",
      "content": "<h2>Greetings Kagglers! 🎉</h2>\n<p>I am thrilled to share my journey during this competition that led me to the 🥉third position on both public and private leaderboards.</p>\n<p>In this competion I faced two challenges:</p>\n<ol>\n<li>Limited knowlegde of networks and their security </li>\n<li>Unbalance in the dataset</li>\n</ol>\n<h1>My Journey 🚀</h1>\n<h2>EDA 📊</h2>\n<p>I started this competition from some basic analysis of the dataset provided by the organizers and I've done some basic EDA to visualize patterns in the dataset, in particular due to the unbalance of the dataset this first visual analysis led to practically no insights, maybe my inexperience in the field made me overlook some important insight that could have helped later on.</p>\n<h2>Modelling, Phase 1 🔬</h2>\n<p>After performing EDA I tried creating some basic feature from the dataset, for example I counted the occurences of a specific IP, i found the attack duration from a specific IP, I found the most frequent attacker country for an IP.<br>\nAt the beginning I didn't include shodan features.</p>\n<p>In this first phase I used LGBM Classifier as my model. <br>\nAt the beginning my scores were around 0.5 on the public leaderboard but also locally. <br>\nAfter analyzing the predictions of the model I noticed that the predictions were strongly skewed and the majority was under 0.5, so the threshold for considering a sample a Proxy/VPN could not be kept at 0.5.</p>\n<p>Changing the threshold to values around 0.3 helped me boost the initial scores of my LGBM classfier models to around 0.6.</p>\n<h2>Modelling, Phase 2🧪</h2>\n<p>The big jump from 0.6 to around 0.75 happened when I tried to combine my features with those provided by the organizer in the <a href=\"https://www.kaggle.com/code/alpacads/vpn-proxy-demo-notebook\" target=\"_blank\">Demo Notebook</a>, in particular I used all the featur provided by the organizers and added some new features.</p>\n<ul>\n<li>I added new ports to be counted taking the new values from the most frequent ports seen in the dataset.</li>\n<li>I counted the number of ports open with tcp and udp protocols</li>\n<li>I counted how many values were present (non Null) for jarm, ja3s and headers_hash in the dataset for that IP</li>\n</ul>\n<p>To these features I added frequency and a target encoding of the attacker country and the frequency of attacker name (no target encoding since this would probably lead to overfitting features).</p>\n<p>From the features computed in Phase 1 I kept the count of attacks, the duration of the attack and the standard deviation of the timestamps of the attack.</p>\n<p>Using those features boosted a lot the model leading almost to the final results.</p>\n<h2>Modeling, Phase 3⚗️</h2>\n<p>After finding the features in Phase 2 that seemed to bring lot of information to the model I was using I tried using several models.<br>\nIn particular I tried these models and found those results (F1 Scores):</p>\n<ol>\n<li>SGDClassifier: 0.446</li>\n<li>LogisticRegression: 0.568</li>\n<li>KNeighborsClassifier (K=7,weights=distance): 0.595</li>\n<li>GaussianNB: 0.160</li>\n<li>DecisionTreeClassifier: 0.710</li>\n<li>ExtraTreesClassifier: 0.747</li>\n<li>HistGradientBoostingClassifier: 0.767</li>\n<li>RandomForestClassifier: 0.780</li>\n</ol>\n<p>These results were computed training the models using GridSearchCV to find some good sets of hyperparameters.</p>\n<h2>Inference and final results 📝</h2>\n<p>Using the results found in Phase 3 I decided to use a Random Forest Classifier as my final model, in particular I trained 5 models, using KFold to generate 5 sets of train and validation dataset one for each model.</p>\n<p>I averaged the prediction of the five models and used a custom threshold for binarization (around 0.3) that was computed by averaging the best threshold (rounded to 2 decimals) for each of the models on their corresponding validation set.</p>\n<p>This allowed my final prediction to have a stable result on both public and private leaderboard.</p>\n<p>In particular for the final submission on the leaderboard I took the result from different runs of this process (not seeded) and performed voting ensemble among them.</p>\n<h1>Conclusions ✍</h1>\n<p>This was a great learning experience, I learned a lot and had to automate as much as possible modelling to try out many models and possibilities to get which models performed the best.</p>\n<p>I hope to get to interact with more people during the next challenge.<br>\nI was quite too busy and had too small time for the challenge, I had practically my notebooks running experiments and also submitting for me because I had just that free time to start the notebooks.</p>\n<p>I'm looking forward to learn from other people solution and see what different approaches were tested by other participants.</p>\n<h1>Possible improvements, Next steps👣</h1>\n<p>Better analysis of model results splitting between precision and recall scores to better understand the strength and weaknessed of the models.<br>\nFor example it could be possible to use a high recall model to filter candidate positive samples and after that use a high precision model to get accurate results on the filtered values.</p>\n<p>Ensembles of different models if implemented correctly could also improve the score, the reliability and the stability of the results.</p>\n<h1>Thank you for reading. Happy Kaggling🌟</h1>",
      "rawMarkdown": "## Greetings Kagglers! 🎉\n\nI am thrilled to share my journey during this competition that led me to the 🥉third position on both public and private leaderboards.\n\nIn this competion I faced two challenges:\n1. Limited knowlegde of networks and their security \n2. Unbalance in the dataset\n\n# My Journey 🚀\n\n## EDA 📊\n\nI started this competition from some basic analysis of the dataset provided by the organizers and I've done some basic EDA to visualize patterns in the dataset, in particular due to the unbalance of the dataset this first visual analysis led to practically no insights, maybe my inexperience in the field made me overlook some important insight that could have helped later on.\n\n## Modelling, Phase 1 🔬\n\nAfter performing EDA I tried creating some basic feature from the dataset, for example I counted the occurences of a specific IP, i found the attack duration from a specific IP, I found the most frequent attacker country for an IP.\nAt the beginning I didn't include shodan features.\n\nIn this first phase I used LGBM Classifier as my model. \nAt the beginning my scores were around 0.5 on the public leaderboard but also locally. \nAfter analyzing the predictions of the model I noticed that the predictions were strongly skewed and the majority was under 0.5, so the threshold for considering a sample a Proxy/VPN could not be kept at 0.5.\n\nChanging the threshold to values around 0.3 helped me boost the initial scores of my LGBM classfier models to around 0.6.\n\n## Modelling, Phase 2🧪\n\nThe big jump from 0.6 to around 0.75 happened when I tried to combine my features with those provided by the organizer in the [Demo Notebook](https://www.kaggle.com/code/alpacads/vpn-proxy-demo-notebook), in particular I used all the featur provided by the organizers and added some new features.\n\n- I added new ports to be counted taking the new values from the most frequent ports seen in the dataset.\n- I counted the number of ports open with tcp and udp protocols\n- I counted how many values were present (non Null) for jarm, ja3s and headers_hash in the dataset for that IP\n\nTo these features I added frequency and a target encoding of the attacker country and the frequency of attacker name (no target encoding since this would probably lead to overfitting features).\n\nFrom the features computed in Phase 1 I kept the count of attacks, the duration of the attack and the standard deviation of the timestamps of the attack.\n\nUsing those features boosted a lot the model leading almost to the final results.\n\n## Modeling, Phase 3⚗️\n\nAfter finding the features in Phase 2 that seemed to bring lot of information to the model I was using I tried using several models.\nIn particular I tried these models and found those results (F1 Scores):\n1. SGDClassifier: 0.446\n2. LogisticRegression: 0.568\n3. KNeighborsClassifier (K=7,weights=distance): 0.595\n4. GaussianNB: 0.160\n5. DecisionTreeClassifier: 0.710\n6. ExtraTreesClassifier: 0.747\n7. HistGradientBoostingClassifier: 0.767\n8. RandomForestClassifier: 0.780\n\nThese results were computed training the models using GridSearchCV to find some good sets of hyperparameters.\n\n\n## Inference and final results 📝\n\nUsing the results found in Phase 3 I decided to use a Random Forest Classifier as my final model, in particular I trained 5 models, using KFold to generate 5 sets of train and validation dataset one for each model.\n\nI averaged the prediction of the five models and used a custom threshold for binarization (around 0.3) that was computed by averaging the best threshold (rounded to 2 decimals) for each of the models on their corresponding validation set.\n\nThis allowed my final prediction to have a stable result on both public and private leaderboard.\n\nIn particular for the final submission on the leaderboard I took the result from different runs of this process (not seeded) and performed voting ensemble among them.\n\n\n# Conclusions ✍ \n\nThis was a great learning experience, I learned a lot and had to automate as much as possible modelling to try out many models and possibilities to get which models performed the best.\n\nI hope to get to interact with more people during the next challenge.\nI was quite too busy and had too small time for the challenge, I had practically my notebooks running experiments and also submitting for me because I had just that free time to start the notebooks.\n\nI'm looking forward to learn from other people solution and see what different approaches were tested by other participants.\n\n# Possible improvements, Next steps👣\n\nBetter analysis of model results splitting between precision and recall scores to better understand the strength and weaknessed of the models.\nFor example it could be possible to use a high recall model to filter candidate positive samples and after that use a high precision model to get accurate results on the filtered values.\n\nEnsembles of different models if implemented correctly could also improve the score, the reliability and the stability of the results.\n\n# Thank you for reading. Happy Kaggling🌟",
      "votes": null
    },
    {
      "id": "2621784",
      "postDate": "01/27/2024 02:25:11",
      "content": "<p>Gr8 work! I had a few features in common with your work (Mainly ports data feature engineering) but learned a lot from others you engineered. Really interesting work testing different thresholds. Congrats! </p>",
      "rawMarkdown": "Gr8 work! I had a few features in common with your work (Mainly ports data feature engineering) but learned a lot from others you engineered. Really interesting work testing different thresholds. Congrats!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2621784,
      "author_name": "juanmanuelpascual",
      "author_url": "",
      "post_date": "01/27/2024 02:25:11",
      "content": "<p>Gr8 work! I had a few features in common with your work (Mainly ports data feature engineering) but learned a lot from others you engineered. Really interesting work testing different thresholds. Congrats! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2592918": "## Greetings Kagglers! 🎉\n\nI am thrilled to share my journey during this competition that led me to the 🥉third position on both public and private leaderboards.\n\nIn this competion I faced two challenges:\n1. Limited knowlegde of networks and their security \n2. Unbalance in the dataset\n\n# My Journey 🚀\n\n## EDA 📊\n\nI started this competition from some basic analysis of the dataset provided by the organizers and I've done some basic EDA to visualize patterns in the dataset, in particular due to the unbalance of the dataset this first visual analysis led to practically no insights, maybe my inexperience in the field made me overlook some important insight that could have helped later on.\n\n## Modelling, Phase 1 🔬\n\nAfter performing EDA I tried creating some basic feature from the dataset, for example I counted the occurences of a specific IP, i found the attack duration from a specific IP, I found the most frequent attacker country for an IP.\nAt the beginning I didn't include shodan features.\n\nIn this first phase I used LGBM Classifier as my model. \nAt the beginning my scores were around 0.5 on the public leaderboard but also locally. \nAfter analyzing the predictions of the model I noticed that the predictions were strongly skewed and the majority was under 0.5, so the threshold for considering a sample a Proxy/VPN could not be kept at 0.5.\n\nChanging the threshold to values around 0.3 helped me boost the initial scores of my LGBM classfier models to around 0.6.\n\n## Modelling, Phase 2🧪\n\nThe big jump from 0.6 to around 0.75 happened when I tried to combine my features with those provided by the organizer in the [Demo Notebook](https://www.kaggle.com/code/alpacads/vpn-proxy-demo-notebook), in particular I used all the featur provided by the organizers and added some new features.\n\n- I added new ports to be counted taking the new values from the most frequent ports seen in the dataset.\n- I counted the number of ports open with tcp and udp protocols\n- I counted how many values were present (non Null) for jarm, ja3s and headers_hash in the dataset for that IP\n\nTo these features I added frequency and a target encoding of the attacker country and the frequency of attacker name (no target encoding since this would probably lead to overfitting features).\n\nFrom the features computed in Phase 1 I kept the count of attacks, the duration of the attack and the standard deviation of the timestamps of the attack.\n\nUsing those features boosted a lot the model leading almost to the final results.\n\n## Modeling, Phase 3⚗️\n\nAfter finding the features in Phase 2 that seemed to bring lot of information to the model I was using I tried using several models.\nIn particular I tried these models and found those results (F1 Scores):\n1. SGDClassifier: 0.446\n2. LogisticRegression: 0.568\n3. KNeighborsClassifier (K=7,weights=distance): 0.595\n4. GaussianNB: 0.160\n5. DecisionTreeClassifier: 0.710\n6. ExtraTreesClassifier: 0.747\n7. HistGradientBoostingClassifier: 0.767\n8. RandomForestClassifier: 0.780\n\nThese results were computed training the models using GridSearchCV to find some good sets of hyperparameters.\n\n\n## Inference and final results 📝\n\nUsing the results found in Phase 3 I decided to use a Random Forest Classifier as my final model, in particular I trained 5 models, using KFold to generate 5 sets of train and validation dataset one for each model.\n\nI averaged the prediction of the five models and used a custom threshold for binarization (around 0.3) that was computed by averaging the best threshold (rounded to 2 decimals) for each of the models on their corresponding validation set.\n\nThis allowed my final prediction to have a stable result on both public and private leaderboard.\n\nIn particular for the final submission on the leaderboard I took the result from different runs of this process (not seeded) and performed voting ensemble among them.\n\n\n# Conclusions ✍ \n\nThis was a great learning experience, I learned a lot and had to automate as much as possible modelling to try out many models and possibilities to get which models performed the best.\n\nI hope to get to interact with more people during the next challenge.\nI was quite too busy and had too small time for the challenge, I had practically my notebooks running experiments and also submitting for me because I had just that free time to start the notebooks.\n\nI'm looking forward to learn from other people solution and see what different approaches were tested by other participants.\n\n# Possible improvements, Next steps👣\n\nBetter analysis of model results splitting between precision and recall scores to better understand the strength and weaknessed of the models.\nFor example it could be possible to use a high recall model to filter candidate positive samples and after that use a high precision model to get accurate results on the filtered values.\n\nEnsembles of different models if implemented correctly could also improve the score, the reliability and the stability of the results.\n\n# Thank you for reading. Happy Kaggling🌟",
    "2621784": "Gr8 work! I had a few features in common with your work (Mainly ports data feature engineering) but learned a lot from others you engineered. Really interesting work testing different thresholds. Congrats!"
  },
  "source": "meta"
}