Home < Articles < Article Details
Feature Reduction Sensitivity to Evaluation Protocol in Phishing URL Detection
| Researcher(s): | |
| Institution: | Faculty of Information Technology, University of Gharyan, Libya |
| Field: | Computer science, expert systems and information technology |
| Published in: | 39th volume - July 2026 |
الملخص
يُستخدم اختزال السمات على نطاق واسع في اكتشاف روابط التصيد الاحتيالي المعتمد على تعلم الآلة بهدف تقليل تعقيد النموذج مع الحفاظ على أداء الكشف. ومع ذلك، قد تعتمد الاستنتاجات المتعلقة بمدى إمكانية اختزال السمات على الطريقة المستخدمة في تقسيم بيانات التدريب والاختبار. تبحث هذه الدراسة فيما إذا كانت الاستنتاجات المتعلقة باختزال السمات تظل مستقرة عند الانتقال من التحقق المتقاطع العشوائي الطبقي إلى التحقق المتقاطع المنفصل حسب النطاق. أُجريت التجارب على مجموعة بيانات URL-Phish Version 2 باستخدام ترتيب السمات بطريقة Mutual Information داخل كل طية تدريب، مع أربع ميزانيات محددة مسبقًا للسمات تتراوح من 22 إلى 5 سمات. وتم تقييم كل من Logistic Regression وRandom Forest تحت كلا البروتوكولين، مع استخدام PR AUC بوصفه المقياس الأساسي. انخفضت قيمة PR AUC مع تقليل ميزانية السمات لدى كلا المصنفين وتحت كلا بروتوكولي التقييم، ولذلك ظل الاستنتاج النوعي بشأن اختزال السمات متسقًا ضمن هذا الإعداد التجريبي. ومع ذلك، اختلف حجم الفجوة بين بروتوكولي التقييم باختلاف المصنف وميزانية السمات؛ إذ كانت أوضح بالنسبة إلى Random Forest عند أصغر ميزانية للسمات، في حين كانت فجوات Logistic Regression أقل وضوحًا مقارنة بالتباين على مستوى الطيات. كما أدى التقييم المنفصل حسب النطاق إلى تباين أكبر على مستوى الطيات. وتوضح النتائج أن الادعاءات المتعلقة بفعالية التمثيلات المختصرة لروابط التصيد الاحتيالي ينبغي تفسيرها بالاقتران مع بروتوكول التقييم المستخدم. وينبغي أن تبحث الدراسات المستقبلية فيما إذا كان هذا النمط يستمر عبر مجموعات بيانات ومصنفات وقواعد تجميع نطاقات إضافية.....................
الكلمات المفتاحية:............. اكتشاف روابط التصيد الاحتيالي، اختزال السمات، التقييم المنفصل حسب النطاق، تعلم الآلة، المعلومات المتبادلة، التعميم
Abstract
Feature reduction is widely used in machine learning based phishing URL detection to lower model complexity while retaining detection performance. However, conclusions about how much reduction is acceptable may depend on how training and testing data are partitioned. This study examines whether conclusions about feature reduction remain stable when evaluation changes from random stratified to domain disjoint cross validation. Experiments were conducted on URL-Phish Version 2 using Mutual Information ranking within each training fold and four predefined feature budgets, from 22 to 5 features. Logistic Regression and Random Forest were evaluated under both protocols, with PR AUC as the primary metric. PR AUC decreased as the feature budget was reduced for both classifiers under both evaluation protocols, so the qualitative conclusion about feature reduction remained consistent within this experimental setting. However, the size of the protocol gap varied across classifiers and feature budgets: it was clearest for Random Forest under the smallest feature budget, whereas the Logistic Regression gaps were less clearly separated from fold-level variability. Domain disjoint evaluation also produced greater fold-level variability. The results show that claims about compact phishing URL representations should be interpreted together with the evaluation protocol used. Future work should examine whether this pattern persists across additional datasets, classifiers, and domain grouping rules...............
Keywords:............ Phishing URL Detection, Feature Reduction, Domain Disjoint Evaluation, Machine Learning, Mutual Information, Generalization